AllocBench
Can coding agents rebuild an allocation engine from its run history?
Modeled on a real optimization engine for a top-5 global data center developer, AllocBench asks agents to rebuild the allocation engine from nothing but its run history. None of the rules are documented, so the agent has to infer each one from past runs. We grade whether its solver reaches the best possible plan. Each task represents a distinct set of inputs.
- Tasks
- 2
- Public tasks
- 2
- Models evaluated
- 3
- Trials
- 29
Benchmark view
Model leaderboard
Average score across tasks, with each task weighted equally.
- Mean reward
- Average of task means across versions matching the dataset’s rules.
- ± SE
- Standard error across task means; directional with few tasks.
- GPT 5.6 Solopenai · opencode · 1.18.1169.5%± 30.5 pp SETask means: 39.0%, 100.0%.
- Muse Spark 1.3 Contributoropencode · opencode · 1.18.1156.1%± 0.3 pp SETask means: 55.8%, 56.5%.
- gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.1145.0%± 6.0 pp SETask means: 39.0%, 51.0%.
Bars use a fixed 0–100% scale. Ticks mark each task mean.
Anatomy of a task
How AllocBench works
1 · The agent gets a planning cycle
A supply-allocation scenario from an anonymized data-center operator: equipment lots with ERP pool noise and missing costs, deal-phase demand with required-on-site dates, and a config whose switches — supplier pins, the single-supplier rule, locked phases — are named but never explained. The retired engine's decision policy was never written down; the only spec is its run history, a salted archive of captured request/plan pairs the agent must reverse-engineer.
{ "config": { "objective": "timelineRisk", "enforceSingleSupplier": true, "supplierPins": [ { "phaseId": "phase-tus02-2", "equipmentTypeId": "eq-gen", "supplierId": "sup-duneland" } ], "lockedPhaseIds": ["phase-msn04-1"] }, "demand": [ { "id": "req-msn04-1-pdu", "quantityRequired": 36, "needByDate": "2027-04-01", … }, … ], "supply": [ { "id": "lot-0016", "inventoryPool": " ops spare ", "unitCost": null, "availabilityDate": "2027-06-15", … }, … ] }2 · It ships an allocation engine
Not a plan — a solver. The deliverable is a pure function from request to assignments that reproduces the legacy engine's behavior: its lexicographic objective (maximize fill, then minimize lateness, then cost), its pricing of unpriced lots, its pool exclusions — every rule the archive evidences.
# allocator/solve.py def solve(request: dict) -> dict: """Return an allocation response for an allocation request.""" … return { "assignments": [ { "demandId": "req-msn04-1-pdu", "lotId": "lot-0142", "quantity": 24 }, … ] }3 · Grading is what the plan achieves
Run → Check rules → Compare → Perturb
The verifier re-derives feasibility from the contract, then compares the plan's fill, lateness, and cost against the golden optimum — any optimal plan passes, whichever lots it picks. Then the world changes: new, cheaper lots land and one more phase locks. The engine re-runs, and locked phases must come back byte-for-byte even though re-optimizing them would look better. That freeze rule exists because a real operator once watched a locked phase move.
Current evidence
Task breakdown
Average scores across matching versions of each public sample. The dataset leaderboard also includes private tasks.
| Task | gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11 | GPT 5.6 Solopenai · opencode · 1.18.11 | Muse Spark 1.3 Contributoropencode · opencode · 1.18.11 |
|---|---|---|---|
| dc-legacy-alloc-full-stackBrowse version 0.1.0 · 329533d2 | 39.0%± 0.0 pp SE2 attempts | 39.0%± 22.5 pp SE3 attempts | 56.5%± 9.1 pp SE11 attempts |
| dc-legacy-alloc-planning-seasonBrowse version 0.1.0 · ace31728 | 51.0%1 attempt | 100.0%1 attempt | 55.8%± 12.5 pp SE11 attempts |