nplus1 × Zendo

Public task

dc-legacy-alloc-planning-season

Reproduce a full planning season: four re-run cycles as lots land and locks accumulate.

Model configurations
3
Evaluation attempts
13
Public examples
13

Current task results

Evaluation comparison

Published evaluation results for this task version. Each bar shows a model configuration’s average score.

Model averages for this task version
ModelAverage scoreScore rangeAttempts
GPT 5.6 Solopenai · opencode · 1.18.11100.0%100.0% – 100.0%11 scored
Muse Spark 1.3 Contributoropencode · opencode · 1.18.1155.8%± 12.5 pp SE0.0% – 100.0%1111 scored
gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.1151.0%51.0% – 51.0%11 scored

Averages and ranges use scored attempts on this exact version. Attempts without a score are counted separately. SE measures uncertainty across attempt scores.

The task

What the agent receives
# Replay a planning season

The old planning module was retired last quarter and its allocation engine's source did not survive the migration. Its captured production runs are in `audit/` — request/plan pairs, exactly as the engine consumed and produced them. The decision policy was never written down; the archive is the only spec.

The engine never ran once; it ran all season. After the opening cycle in `data/scenario.json`, lots keep landing and reviewed phases keep locking — each cycle's plan becomes the next cycle's commitments. The verifier replays four such cycles against your engine, so a small policy error early compounds into every later cycle.

Implement the engine at `allocator/solve.py` (entry point in
`allocator/README.md`) so it reproduces the legacy engine's planning
behavior on any request, and use it to plan `data/scenario.json`. The engine
re-runs as the world changes — new lots land and phases lock between runs —
and grading is outcome equivalence with the legacy engine, not matching its
exact lot choices.

Check plan structure with
`python3 contracts/validate.py <request.json> <response.json>`.