InfraBench

Code agent evaluations — framework reasoning · reusable plugins

nplus1 × Zendo

Patterns over patches

Can coding agents extend an infra code-generation framework?

These tasks test whether models can reason about n1x, reuse its existing patterns, and build new reusable plugins. Harbor verifies each result from a clean repository.

Tasks3
Models & controls3
Published attempts24
Data as ofAug 5, 2026

Benchmark view

Model leaderboard

Average score across tasks, with each task weighted equally.

Mean reward
Macro-average across the latest published task cohorts.
± SE
Standard error across task means; directional with 3 tasks.
  1. v9m-rl-learnability-tp8openai
    80.3%± 9.0 pp SE
    3/3 tasks · 8 attempts
    Task means: nplus1/n1x-openai-provisioning 65.0%, nplus1/n1x-web-ag-composition 80.0%, nplus1/n1x-web-job-runner 96.0%
  2. gpt-5.6-solopenai
    79.6%± 3.7 pp SE
    3/3 tasks · 8 attempts
    Task means: nplus1/n1x-openai-provisioning 72.3%, nplus1/n1x-web-ag-composition 82.0%, nplus1/n1x-web-job-runner 84.3%
  3. gpt-5.6-lunaopenai
    72.7%± 10.5 pp SE
    3/3 tasks · 8 attempts
    Task means: nplus1/n1x-openai-provisioning 67.0%, nplus1/n1x-web-ag-composition 58.0%, nplus1/n1x-web-job-runner 93.0%

Bars use a fixed 0–100 scale. Ticks mark each task mean.

Current evidence

Task breakdown

Published task evaluation results
TaskModel / batchFull scoreMeanAvg runtimeAttempts
n1x-openai-provisioninghardgpt-5.6-lunaopenai · 2026-08-05__12-16-560/30.677m
  1. 1
  2. 2
  3. 3
gpt-5.6-solopenai · 2026-08-05__12-16-560/30.7213m
  1. 1
  2. 2
  3. 3
v9m-rl-learnability-tp8openai · 2026-08-05__12-16-560/30.6524m
  1. 1
  2. 2
  3. 3
n1x-web-ag-compositionhardgpt-5.6-lunaopenai · 2026-08-05__10-04-350/20.587m
  1. 1
  2. 2
gpt-5.6-solopenai · 2026-08-05__10-04-350/20.8215m
  1. 1
  2. 2
v9m-rl-learnability-tp8openai · 2026-08-05__10-04-350/20.801h 1m
  1. 1
  2. 2
n1x-web-job-runnerhardgpt-5.6-lunaopenai · 2026-08-05__12-16-530/30.936m
  1. 1
  2. 2
  3. 3
gpt-5.6-solopenai · 2026-08-05__12-16-530/30.8416m
  1. 1
  2. 2
  3. 3
v9m-rl-learnability-tp8openai · 2026-08-05__12-16-530/30.9646m
  1. 1
  2. 2
  3. 3
Full reward Partial reward Verified zero Execution errorOutlined cell = score plus warning

What the agent gets

A runnable repository, one engineering instruction, normal coding tools, and a fixed time budget. It does not receive the reference patch or verifier implementation.

What the verifier measures

Harbor runs hidden behavioral checks against the final workspace. Rewards are continuous, so a partial result can expose exactly which contracts held and which did not.

How to read failures

A verified zero differs from an infrastructure or agent failure. Zendo keeps reward and run health separate so timeouts do not look like ordinary model answers.