nplus1 × Zendo

AllocBench

Can coding agents rebuild an allocation engine from its run history?

Modeled on a real optimization engine for a top-5 global data center developer, AllocBench asks agents to rebuild the allocation engine from nothing but its run history. None of the rules are documented, so the agent has to infer each one from past runs. We grade whether its solver reaches the best possible plan. Each task represents a distinct set of inputs.

Optimization
Tasks
2
Public tasks
2
Models evaluated
3
Trials
29

Benchmark view

Model leaderboard

Average score across tasks, with each task weighted equally.

Mean reward
Average of task means across versions matching the dataset’s rules.
± SE
Standard error across task means; directional with few tasks.
  1. GPT 5.6 Solopenai · opencode · 1.18.11
    69.5%± 30.5 pp SE
    2/2 tasks · 4 attempts
    Task means: 39.0%, 100.0%.
  2. Muse Spark 1.3 Contributoropencode · opencode · 1.18.11
    56.1%± 0.3 pp SE
    2/2 tasks · 22 attempts
    Task means: 55.8%, 56.5%.
  3. gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11
    45.0%± 6.0 pp SE
    2/2 tasks · 3 attempts
    Task means: 39.0%, 51.0%.

Bars use a fixed 0–100% scale. Ticks mark each task mean.

Anatomy of a task

How AllocBench works

  1. 1 · The agent gets a planning cycle

    A supply-allocation scenario from an anonymized data-center operator: equipment lots with ERP pool noise and missing costs, deal-phase demand with required-on-site dates, and a config whose switches — supplier pins, the single-supplier rule, locked phases — are named but never explained. The retired engine's decision policy was never written down; the only spec is its run history, a salted archive of captured request/plan pairs the agent must reverse-engineer.

    {
      "config": {
        "objective": "timelineRisk",
        "enforceSingleSupplier": true,
        "supplierPins": [ { "phaseId": "phase-tus02-2",
          "equipmentTypeId": "eq-gen",
          "supplierId": "sup-duneland" } ],
        "lockedPhaseIds": ["phase-msn04-1"]
      },
      "demand": [ { "id": "req-msn04-1-pdu",
        "quantityRequired": 36,
        "needByDate": "2027-04-01", … }, … ],
      "supply": [ { "id": "lot-0016",
        "inventoryPool": " ops  spare ",
        "unitCost": null,
        "availabilityDate": "2027-06-15", … }, … ]
    }
  2. 2 · It ships an allocation engine

    Not a plan — a solver. The deliverable is a pure function from request to assignments that reproduces the legacy engine's behavior: its lexicographic objective (maximize fill, then minimize lateness, then cost), its pricing of unpriced lots, its pool exclusions — every rule the archive evidences.

    # allocator/solve.py
    def solve(request: dict) -> dict:
        """Return an allocation response
        for an allocation request."""
        …
        return { "assignments": [
          { "demandId": "req-msn04-1-pdu",
            "lotId": "lot-0142",
            "quantity": 24 }, … ] }
  3. 3 · Grading is what the plan achieves

    Run → Check rules → Compare → Perturb

    The verifier re-derives feasibility from the contract, then compares the plan's fill, lateness, and cost against the golden optimum — any optimal plan passes, whichever lots it picks. Then the world changes: new, cheaper lots land and one more phase locks. The engine re-runs, and locked phases must come back byte-for-byte even though re-optimizing them would look better. That freeze rule exists because a real operator once watched a locked phase move.

Current evidence

Task breakdown

Average scores across matching versions of each public sample. The dataset leaderboard also includes private tasks.

Average score by public task and model configuration
Taskgemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11GPT 5.6 Solopenai · opencode · 1.18.11Muse Spark 1.3 Contributoropencode · opencode · 1.18.11
dc-legacy-alloc-full-stackBrowse version 0.1.0 · 329533d239.0%± 0.0 pp SE2 attempts39.0%± 22.5 pp SE3 attempts56.5%± 9.1 pp SE11 attempts
dc-legacy-alloc-planning-seasonBrowse version 0.1.0 · ace3172851.0%1 attempt100.0%1 attempt55.8%± 12.5 pp SE11 attempts