Evidence / Bounded web model / 5 October 2026
Put the improvement
in context.
A useful result needs a useful comparison. We test the same synthetic workloads against basic first-fit, five fixed heuristic rules and a larger search. Every allocation faces the same capacity checks.
01 / The default company challenge
A stronger reference changes the headline.
48 workloads, 120% CPU demand, 20% CPU and memory headroom, and 512 search candidates. The basic reference is intentionally simple. The five-rule portfolio makes a more useful comparison.
| Method | Cost units | Hosts | Selected checks |
|---|---|---|---|
| Basic input-order first-fit | 454 | 31 | Passed |
| Five-rule heuristic portfolio | 334 | 10 | Passed |
| Bounded search | 319 | 10 | Passed |
4.49% lower cost than the five-rule reference. 29.74% lower than basic first-fit. Both comparisons stay visible.
The five-rule reference comes from the same demonstration and is included in the larger search. This is an internal heuristic comparison, not evidence of superiority over Hexaly, OR-Tools, Gurobi or another commercial solver.
02 / All supported settings
Include the ties.
The declared matrix covers three scenarios, two capacity policies, three CPU demand levels and four search budgets: 72 configurations. These settings share three synthetic scenario families; they are not 72 independent customer problems.
- Better than five rules
- 69
- Tied with five rules
- 3
- Failed selected checks
- 0
All 72 runs completed. The three ties occur in memory pressure at 140% CPU demand and full capacity, with budgets of 64, 128 and 256. Their candidate and five-rule costs both equal 271 units. Ties remain in the report.
Mixed fleet 24 workloads · 24 configurations
| Configuration | Basic | Five rules | Search | Beyond five rules | Search hosts | Outcome |
|---|---|---|---|---|---|---|
| 100% · Full capacity · 64 | 140 | 108 | 102 | 5.56% | 3 | Improved |
| 100% · Full capacity · 128 | 140 | 108 | 102 | 5.56% | 3 | Improved |
| 100% · Full capacity · 256 | 140 | 108 | 102 | 5.56% | 3 | Improved |
| 100% · Full capacity · 512 | 140 | 108 | 102 | 5.56% | 3 | Improved |
| 120% · Full capacity · 64 | 167 | 137 | 126 | 8.03% | 4 | Improved |
| 120% · Full capacity · 128 | 167 | 137 | 125 | 8.76% | 5 | Improved |
| 120% · Full capacity · 256 | 167 | 137 | 125 | 8.76% | 5 | Improved |
| 120% · Full capacity · 512 | 167 | 137 | 125 | 8.76% | 5 | Improved |
| 140% · Full capacity · 64 | 196 | 157 | 139 | 11.46% | 5 | Improved |
| 140% · Full capacity · 128 | 196 | 157 | 139 | 11.46% | 5 | Improved |
| 140% · Full capacity · 256 | 196 | 157 | 139 | 11.46% | 5 | Improved |
| 140% · Full capacity · 512 | 196 | 157 | 139 | 11.46% | 5 | Improved |
| 100% · 20% reserve · 64 | 186 | 144 | 132 | 8.33% | 4 | Improved |
| 100% · 20% reserve · 128 | 186 | 144 | 132 | 8.33% | 4 | Improved |
| 100% · 20% reserve · 256 | 186 | 144 | 131 | 9.03% | 5 | Improved |
| 100% · 20% reserve · 512 | 186 | 144 | 131 | 9.03% | 5 | Improved |
| 120% · 20% reserve · 64 | 215 | 173 | 162 | 6.36% | 5 | Improved |
| 120% · 20% reserve · 128 | 215 | 173 | 162 | 6.36% | 5 | Improved |
| 120% · 20% reserve · 256 | 215 | 173 | 156 | 9.83% | 5 | Improved |
| 120% · 20% reserve · 512 | 215 | 173 | 155 | 10.40% | 6 | Improved |
| 140% · 20% reserve · 64 | 252 | 193 | 180 | 6.74% | 6 | Improved |
| 140% · 20% reserve · 128 | 252 | 193 | 180 | 6.74% | 6 | Improved |
| 140% · 20% reserve · 256 | 252 | 193 | 175 | 9.33% | 6 | Improved |
| 140% · 20% reserve · 512 | 252 | 193 | 175 | 9.33% | 6 | Improved |
Memory pressure 48 workloads · 24 configurations
| Configuration | Basic | Five rules | Search | Beyond five rules | Search hosts | Outcome |
|---|---|---|---|---|---|---|
| 100% · Full capacity · 64 | 337 | 262 | 238 | 9.16% | 8 | Improved |
| 100% · Full capacity · 128 | 337 | 262 | 238 | 9.16% | 8 | Improved |
| 100% · Full capacity · 256 | 337 | 262 | 238 | 9.16% | 8 | Improved |
| 100% · Full capacity · 512 | 337 | 262 | 238 | 9.16% | 8 | Improved |
| 120% · Full capacity · 64 | 364 | 262 | 255 | 2.67% | 8 | Improved |
| 120% · Full capacity · 128 | 364 | 262 | 246 | 6.11% | 8 | Improved |
| 120% · Full capacity · 256 | 364 | 262 | 246 | 6.11% | 8 | Improved |
| 120% · Full capacity · 512 | 364 | 262 | 246 | 6.11% | 8 | Improved |
| 140% · Full capacity · 64 | 383 | 271 | 271 | 0.00% | 8 | Tie |
| 140% · Full capacity · 128 | 383 | 271 | 271 | 0.00% | 8 | Tie |
| 140% · Full capacity · 256 | 383 | 271 | 271 | 0.00% | 8 | Tie |
| 140% · Full capacity · 512 | 383 | 271 | 262 | 3.32% | 8 | Improved |
| 100% · 20% reserve · 64 | 425 | 326 | 300 | 7.98% | 9 | Improved |
| 100% · 20% reserve · 128 | 425 | 326 | 300 | 7.98% | 9 | Improved |
| 100% · 20% reserve · 256 | 425 | 326 | 284 | 12.88% | 9 | Improved |
| 100% · 20% reserve · 512 | 425 | 326 | 284 | 12.88% | 9 | Improved |
| 120% · 20% reserve · 64 | 454 | 334 | 324 | 2.99% | 9 | Improved |
| 120% · 20% reserve · 128 | 454 | 334 | 324 | 2.99% | 9 | Improved |
| 120% · 20% reserve · 256 | 454 | 334 | 324 | 2.99% | 9 | Improved |
| 120% · 20% reserve · 512 | 454 | 334 | 319 | 4.49% | 10 | Improved |
| 140% · 20% reserve · 64 | 482 | 352 | 334 | 5.11% | 10 | Improved |
| 140% · 20% reserve · 128 | 482 | 352 | 334 | 5.11% | 10 | Improved |
| 140% · 20% reserve · 256 | 482 | 352 | 334 | 5.11% | 10 | Improved |
| 140% · 20% reserve · 512 | 482 | 352 | 334 | 5.11% | 10 | Improved |
Scale pressure 64 workloads · 24 configurations
| Configuration | Basic | Five rules | Search | Beyond five rules | Search hosts | Outcome |
|---|---|---|---|---|---|---|
| 100% · Full capacity · 64 | 467 | 343 | 320 | 6.71% | 10 | Improved |
| 100% · Full capacity · 128 | 467 | 343 | 313 | 8.75% | 11 | Improved |
| 100% · Full capacity · 256 | 467 | 343 | 313 | 8.75% | 11 | Improved |
| 100% · Full capacity · 512 | 467 | 343 | 312 | 9.04% | 10 | Improved |
| 120% · Full capacity · 64 | 471 | 343 | 335 | 2.33% | 10 | Improved |
| 120% · Full capacity · 128 | 471 | 343 | 335 | 2.33% | 10 | Improved |
| 120% · Full capacity · 256 | 471 | 343 | 335 | 2.33% | 10 | Improved |
| 120% · Full capacity · 512 | 471 | 343 | 335 | 2.33% | 10 | Improved |
| 140% · Full capacity · 64 | 510 | 362 | 360 | 0.55% | 10 | Improved |
| 140% · Full capacity · 128 | 510 | 362 | 360 | 0.55% | 10 | Improved |
| 140% · Full capacity · 256 | 510 | 362 | 360 | 0.55% | 10 | Improved |
| 140% · Full capacity · 512 | 510 | 362 | 360 | 0.55% | 10 | Improved |
| 100% · 20% reserve · 64 | 539 | 424 | 408 | 3.77% | 12 | Improved |
| 100% · 20% reserve · 128 | 539 | 424 | 408 | 3.77% | 12 | Improved |
| 100% · 20% reserve · 256 | 539 | 424 | 403 | 4.95% | 13 | Improved |
| 100% · 20% reserve · 512 | 539 | 424 | 400 | 5.66% | 12 | Improved |
| 120% · 20% reserve · 64 | 608 | 434 | 426 | 1.84% | 13 | Improved |
| 120% · 20% reserve · 128 | 608 | 434 | 426 | 1.84% | 13 | Improved |
| 120% · 20% reserve · 256 | 608 | 434 | 424 | 2.30% | 12 | Improved |
| 120% · 20% reserve · 512 | 608 | 434 | 424 | 2.30% | 12 | Improved |
| 140% · 20% reserve · 64 | 655 | 487 | 468 | 3.90% | 13 | Improved |
| 140% · 20% reserve · 128 | 655 | 487 | 462 | 5.13% | 13 | Improved |
| 140% · 20% reserve · 256 | 655 | 487 | 462 | 5.13% | 13 | Improved |
| 140% · 20% reserve · 512 | 655 | 487 | 460 | 5.54% | 13 | Improved |
03 / Method and limits
Separate search from checking.
- Fix the inputs and policies. Reuse each prepared scenario at 100%, 120% or 140% CPU demand. Apply either full capacity or a 20% CPU and memory reserve, rounded down to whole usable units.
- Run the references and search. Basic first-fit keeps input order and opens the cheapest feasible host. The stronger reference selects the best of five fixed ordering and packing rules, including cost-aware opening. The larger search evaluates 64, 128, 256 or 512 candidates.
- Reconstruct each answer. A separate checker implementation reads the original requirements and chosen host types. It checks unique assignments, CPU and memory limits, selected headroom and total cost. It does not trust declared costs or the searcher's pass flags.
- Repeat the measurement. Two operator runs produced identical non-timing results for all 72 configurations. A service rerun also agreed with the separate plan reconstruction. This is operator verification, not third-party certification.
One excluded, bounded warm-up precedes each operator benchmark. Each measured configuration runs in an isolated worker. Published timings are one observation per configuration on Apple M4 Pro, arm64, Node v22.23.3. Candidate selection took 3.061–38.283 ms in the published run. Per-method measurements exclude startup, network and file I/O; search-internal checks are included. Whole-configuration measurements include isolated-worker startup, both references, search, independent checks and the service rerun, and took 23.484–71.372 ms. Post-selection verification is timed separately in the JSON.
The candidate-count limit bounds the work examined here. An enforced 5,000 ms ceiling covers each complete configuration, including worker startup, references, search, independent checks and the service rerun. The parent interrupts a timed-out worker and records the failure; later configurations are marked unrun. These timings are not a Cloudflare latency promise or a production service-level agreement. The result does not establish global optimality, full-engine performance, customer savings or production readiness.
04 / Reproduction and review
Check a supplied result yourself.
The public example already exposes its synthetic inputs and allocations. Download its recorded default run and a standalone checker. The checker reconstructs all three allocations, constraints and cost differences without running or importing Istimthal's optimizer.
node verify-demo-record.mjs benchmark-public-example.jsonYou can also run the public example, download a fresh test record and check it with the same tool. A supplied-record check verifies internal arithmetic and constraints; it does not authenticate the inputs, prove optimality or rerun the search.
The company matrix contains aggregate results only. An existing finite invitation lets its holder repeat the displayed settings and compare totals. The ten-run cap and fourteen-day maximum lifetime still apply. Detailed company-showcase inputs and allocations remain private. Independent reconstruction of those results requires a separately scoped evaluation with agreed access.
Source digests and operator reproduction
The operator reruns the current private source with node scripts/run-showcase-benchmark.mjs --output <new-file.json>, then checks the aggregate artifact with node scripts/verify-showcase-benchmark.mjs <new-file.json>. These digests identify the tested revision; they do not disclose the optimizer.
| Operator file | SHA-256 |
|---|---|
| src/demo-model.js | f1d53abaa1039156edd837386d283b6fc7ec2ae575ba95cc3156feac0de62427 |
| src/company-demo-service.js | 67893db887595b0d9e2c0e3623d9d98dc21c25edd70f30977cade9c318a4de7d |
| scripts/run-showcase-benchmark.mjs | eda57b02df5437187c8293b80767ebbf8f4620ec8cb8634fb53a82e508fe77c9 |
| scripts/verify-showcase-benchmark.mjs | afed7e444c18468b738319c08b9ad5e73cda5e6de875168220ebd367c22c854b |
Bring your own baseline.
The next useful evidence is a result on an operator-owned problem. Agree on the workload, objective, constraints, budget and acceptance checks before a paid pilot begins.
Scope a paid pilot