istimthal

Evidence / Bounded web model / 5 October 2026

Put the improvement
in context.

A useful result needs a useful comparison. We test the same synthetic workloads against basic first-fit, five fixed heuristic rules and a larger search. Every allocation faces the same capacity checks.

01 / The default company challenge

A stronger reference changes the headline.

48 workloads, 120% CPU demand, 20% CPU and memory headroom, and 512 search candidates. The basic reference is intentionally simple. The five-rule portfolio makes a more useful comparison.

Same requirements and constraints. Synthetic costs, not provider prices.
MethodCost unitsHostsSelected checks
Basic input-order first-fit45431Passed
Five-rule heuristic portfolio33410Passed
Bounded search31910Passed

4.49% lower cost than the five-rule reference. 29.74% lower than basic first-fit. Both comparisons stay visible.

The five-rule reference comes from the same demonstration and is included in the larger search. This is an internal heuristic comparison, not evidence of superiority over Hexaly, OR-Tools, Gurobi or another commercial solver.

02 / All supported settings

Include the ties.

The declared matrix covers three scenarios, two capacity policies, three CPU demand levels and four search budgets: 72 configurations. These settings share three synthetic scenario families; they are not 72 independent customer problems.

Better than five rules
69
Tied with five rules
3
Failed selected checks
0

All 72 runs completed. The three ties occur in memory pressure at 140% CPU demand and full capacity, with budgets of 64, 128 and 256. Their candidate and five-rule costs both equal 271 units. Ties remain in the report.

Mixed fleet 24 workloads · 24 configurations
CPU demand, capacity policy and candidate budget; all costs in illustrative units. All selected assignment, capacity and headroom checks passed.
ConfigurationBasicFive rulesSearchBeyond five rulesSearch hostsOutcome
100% · Full capacity · 641401081025.56%3Improved
100% · Full capacity · 1281401081025.56%3Improved
100% · Full capacity · 2561401081025.56%3Improved
100% · Full capacity · 5121401081025.56%3Improved
120% · Full capacity · 641671371268.03%4Improved
120% · Full capacity · 1281671371258.76%5Improved
120% · Full capacity · 2561671371258.76%5Improved
120% · Full capacity · 5121671371258.76%5Improved
140% · Full capacity · 6419615713911.46%5Improved
140% · Full capacity · 12819615713911.46%5Improved
140% · Full capacity · 25619615713911.46%5Improved
140% · Full capacity · 51219615713911.46%5Improved
100% · 20% reserve · 641861441328.33%4Improved
100% · 20% reserve · 1281861441328.33%4Improved
100% · 20% reserve · 2561861441319.03%5Improved
100% · 20% reserve · 5121861441319.03%5Improved
120% · 20% reserve · 642151731626.36%5Improved
120% · 20% reserve · 1282151731626.36%5Improved
120% · 20% reserve · 2562151731569.83%5Improved
120% · 20% reserve · 51221517315510.40%6Improved
140% · 20% reserve · 642521931806.74%6Improved
140% · 20% reserve · 1282521931806.74%6Improved
140% · 20% reserve · 2562521931759.33%6Improved
140% · 20% reserve · 5122521931759.33%6Improved
Memory pressure 48 workloads · 24 configurations
CPU demand, capacity policy and candidate budget; all costs in illustrative units. All selected assignment, capacity and headroom checks passed.
ConfigurationBasicFive rulesSearchBeyond five rulesSearch hostsOutcome
100% · Full capacity · 643372622389.16%8Improved
100% · Full capacity · 1283372622389.16%8Improved
100% · Full capacity · 2563372622389.16%8Improved
100% · Full capacity · 5123372622389.16%8Improved
120% · Full capacity · 643642622552.67%8Improved
120% · Full capacity · 1283642622466.11%8Improved
120% · Full capacity · 2563642622466.11%8Improved
120% · Full capacity · 5123642622466.11%8Improved
140% · Full capacity · 643832712710.00%8Tie
140% · Full capacity · 1283832712710.00%8Tie
140% · Full capacity · 2563832712710.00%8Tie
140% · Full capacity · 5123832712623.32%8Improved
100% · 20% reserve · 644253263007.98%9Improved
100% · 20% reserve · 1284253263007.98%9Improved
100% · 20% reserve · 25642532628412.88%9Improved
100% · 20% reserve · 51242532628412.88%9Improved
120% · 20% reserve · 644543343242.99%9Improved
120% · 20% reserve · 1284543343242.99%9Improved
120% · 20% reserve · 2564543343242.99%9Improved
120% · 20% reserve · 5124543343194.49%10Improved
140% · 20% reserve · 644823523345.11%10Improved
140% · 20% reserve · 1284823523345.11%10Improved
140% · 20% reserve · 2564823523345.11%10Improved
140% · 20% reserve · 5124823523345.11%10Improved
Scale pressure 64 workloads · 24 configurations
CPU demand, capacity policy and candidate budget; all costs in illustrative units. All selected assignment, capacity and headroom checks passed.
ConfigurationBasicFive rulesSearchBeyond five rulesSearch hostsOutcome
100% · Full capacity · 644673433206.71%10Improved
100% · Full capacity · 1284673433138.75%11Improved
100% · Full capacity · 2564673433138.75%11Improved
100% · Full capacity · 5124673433129.04%10Improved
120% · Full capacity · 644713433352.33%10Improved
120% · Full capacity · 1284713433352.33%10Improved
120% · Full capacity · 2564713433352.33%10Improved
120% · Full capacity · 5124713433352.33%10Improved
140% · Full capacity · 645103623600.55%10Improved
140% · Full capacity · 1285103623600.55%10Improved
140% · Full capacity · 2565103623600.55%10Improved
140% · Full capacity · 5125103623600.55%10Improved
100% · 20% reserve · 645394244083.77%12Improved
100% · 20% reserve · 1285394244083.77%12Improved
100% · 20% reserve · 2565394244034.95%13Improved
100% · 20% reserve · 5125394244005.66%12Improved
120% · 20% reserve · 646084344261.84%13Improved
120% · 20% reserve · 1286084344261.84%13Improved
120% · 20% reserve · 2566084344242.30%12Improved
120% · 20% reserve · 5126084344242.30%12Improved
140% · 20% reserve · 646554874683.90%13Improved
140% · 20% reserve · 1286554874625.13%13Improved
140% · 20% reserve · 2566554874625.13%13Improved
140% · 20% reserve · 5126554874605.54%13Improved

03 / Method and limits

Separate search from checking.

  1. Fix the inputs and policies. Reuse each prepared scenario at 100%, 120% or 140% CPU demand. Apply either full capacity or a 20% CPU and memory reserve, rounded down to whole usable units.
  2. Run the references and search. Basic first-fit keeps input order and opens the cheapest feasible host. The stronger reference selects the best of five fixed ordering and packing rules, including cost-aware opening. The larger search evaluates 64, 128, 256 or 512 candidates.
  3. Reconstruct each answer. A separate checker implementation reads the original requirements and chosen host types. It checks unique assignments, CPU and memory limits, selected headroom and total cost. It does not trust declared costs or the searcher's pass flags.
  4. Repeat the measurement. Two operator runs produced identical non-timing results for all 72 configurations. A service rerun also agreed with the separate plan reconstruction. This is operator verification, not third-party certification.

One excluded, bounded warm-up precedes each operator benchmark. Each measured configuration runs in an isolated worker. Published timings are one observation per configuration on Apple M4 Pro, arm64, Node v22.23.3. Candidate selection took 3.061–38.283 ms in the published run. Per-method measurements exclude startup, network and file I/O; search-internal checks are included. Whole-configuration measurements include isolated-worker startup, both references, search, independent checks and the service rerun, and took 23.484–71.372 ms. Post-selection verification is timed separately in the JSON.

The candidate-count limit bounds the work examined here. An enforced 5,000 ms ceiling covers each complete configuration, including worker startup, references, search, independent checks and the service rerun. The parent interrupts a timed-out worker and records the failure; later configurations are marked unrun. These timings are not a Cloudflare latency promise or a production service-level agreement. The result does not establish global optimality, full-engine performance, customer savings or production readiness.

04 / Reproduction and review

Check a supplied result yourself.

The public example already exposes its synthetic inputs and allocations. Download its recorded default run and a standalone checker. The checker reconstructs all three allocations, constraints and cost differences without running or importing Istimthal's optimizer.

node verify-demo-record.mjs benchmark-public-example.json

You can also run the public example, download a fresh test record and check it with the same tool. A supplied-record check verifies internal arithmetic and constraints; it does not authenticate the inputs, prove optimality or rerun the search.

The company matrix contains aggregate results only. An existing finite invitation lets its holder repeat the displayed settings and compare totals. The ten-run cap and fourteen-day maximum lifetime still apply. Detailed company-showcase inputs and allocations remain private. Independent reconstruction of those results requires a separately scoped evaluation with agreed access.

Source digests and operator reproduction

The operator reruns the current private source with node scripts/run-showcase-benchmark.mjs --output <new-file.json>, then checks the aggregate artifact with node scripts/verify-showcase-benchmark.mjs <new-file.json>. These digests identify the tested revision; they do not disclose the optimizer.

Operator fileSHA-256
src/demo-model.jsf1d53abaa1039156edd837386d283b6fc7ec2ae575ba95cc3156feac0de62427
src/company-demo-service.js67893db887595b0d9e2c0e3623d9d98dc21c25edd70f30977cade9c318a4de7d
scripts/run-showcase-benchmark.mjseda57b02df5437187c8293b80767ebbf8f4620ec8cb8634fb53a82e508fe77c9
scripts/verify-showcase-benchmark.mjsafed7e444c18468b738319c08b9ad5e73cda5e6de875168220ebd367c22c854b

Bring your own baseline.

The next useful evidence is a result on an operator-owned problem. Agree on the workload, objective, constraints, budget and acceptance checks before a paid pilot begins.

Scope a paid pilot