What the agent gets
A runnable repository, one engineering instruction, normal coding tools, and a fixed time budget. It does not receive the reference patch or verifier implementation.
Code agent evaluations — framework reasoning · reusable plugins
nplus1 × Zendo
Patterns over patches
These tasks test whether models can reason about n1x, reuse its existing patterns, and build new reusable plugins. Harbor verifies each result from a clean repository.
Benchmark view
Average score across tasks, with each task weighted equally.
Bars use a fixed 0–100 scale. Ticks mark each task mean.
Current evidence
| Task | Model / batch | Full score | Mean | Avg runtime | Attempts |
|---|---|---|---|---|---|
| n1x-openai-provisioninghard | gpt-5.6-lunaopenai · 2026-08-05__12-16-56 | 0/3 | 0.67 | 7m | |
| gpt-5.6-solopenai · 2026-08-05__12-16-56 | 0/3 | 0.72 | 13m | ||
| v9m-rl-learnability-tp8openai · 2026-08-05__12-16-56 | 0/3 | 0.65 | 24m | ||
| n1x-web-ag-compositionhard | gpt-5.6-lunaopenai · 2026-08-05__10-04-35 | 0/2 | 0.58 | 7m | |
| gpt-5.6-solopenai · 2026-08-05__10-04-35 | 0/2 | 0.82 | 15m | ||
| v9m-rl-learnability-tp8openai · 2026-08-05__10-04-35 | 0/2 | 0.80 | 1h 1m | ||
| n1x-web-job-runnerhard | gpt-5.6-lunaopenai · 2026-08-05__12-16-53 | 0/3 | 0.93 | 6m | |
| gpt-5.6-solopenai · 2026-08-05__12-16-53 | 0/3 | 0.84 | 16m | ||
| v9m-rl-learnability-tp8openai · 2026-08-05__12-16-53 | 0/3 | 0.96 | 46m |
A runnable repository, one engineering instruction, normal coding tools, and a fixed time budget. It does not receive the reference patch or verifier implementation.
Harbor runs hidden behavioral checks against the final workspace. Rewards are continuous, so a partial result can expose exactly which contracts held and which did not.
A verified zero differs from an infrastructure or agent failure. Zendo keeps reward and run health separate so timeouts do not look like ordinary model answers.