Public task
dc-legacy-alloc-full-stack
The whole retired engine at once: every config switch, raw ERP data, locks, and pins.
Part of AllocBench ↗
- Model configurations
- 3
- Evaluation attempts
- 16
- Public examples
- 16
Current task results
Evaluation comparison
Published evaluation results for this task version. Each bar shows a model configuration’s average score.
| Model | Average score | Score range | Attempts |
|---|---|---|---|
| Muse Spark 1.3 Contributoropencode · opencode · 1.18.11 | 56.5%± 9.1 pp SE | 0.0% – 100.0% | 1111 scored |
| gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11 | 39.0%± 0.0 pp SE | 39.0% – 39.0% | 22 scored |
| GPT 5.6 Solopenai · opencode · 1.18.11 | 39.0%± 22.5 pp SE | 0.0% – 78.0% | 33 scored |
Averages and ranges use scored attempts on this exact version. Attempts without a score are counted separately. SE measures uncertainty across attempt scores.
The task
What the agent receives# Rebuild the whole retired engine The old planning module was retired last quarter and its allocation engine's source did not survive the migration. Its captured production runs are in `audit/` — request/plan pairs, exactly as the engine consumed and produced them. The decision policy was never written down; the archive is the only spec. The live engine never saw switches one at a time. This is a full production cycle: four campuses, a raw ERP extract, supplier pins, `enforceSingleSupplier`, and a locked phase at ALB01 with committed assignments in the request. Implement the engine at `allocator/solve.py` (entry point in `allocator/README.md`) so it reproduces the legacy engine's planning behavior on any request, and use it to plan `data/scenario.json`. The engine re-runs as the world changes — new lots land and phases lock between runs — and grading is outcome equivalence with the legacy engine, not matching its exact lot choices. Check plan structure with `python3 contracts/validate.py <request.json> <response.json>`.