InfraBench
Can coding agents write high-quality infrastructure code?
InfraBench asks agents to extend the proprietary infra code-generation framework nplus1 uses to build and run infrastructure for its clients, across Terraform, GCP, Vercel and CI. Each task is real work on that framework. We grade it by generating and running real projects from the agent's changes, not by comparing against a reference solution.
- Tasks
- 2
- Public tasks
- 2
- Models evaluated
- 3
- Trials
- 10
Benchmark view
Model leaderboard
Average score across tasks, with each task weighted equally.
- Mean reward
- Average of task means across versions matching the dataset’s rules.
- ± SE
- Standard error across task means; directional with few tasks.
- GPT 5.6 Solopenai · opencode · 1.18.1180.0%SE —Task means: 80.0%.
- Muse Spark 1.3 Contributoropencode · opencode · 1.18.1176.3%± 17.0 pp SETask means: 59.3%, 93.3%.
- gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.1148.0%SE —Task means: 48.0%.
Bars use a fixed 0–100% scale. Ticks mark each task mean.
Anatomy of a task
How InfraBench works
1 · The agent gets a framework
A runnable copy of n1x — the code-generation framework that stamps out client applications from composable plugins — plus one engineering instruction taken from real capability work.
# Add composable AG Grid and AG Charts # Enterprise support to n1x Introduce a `web-ag` capability that requires `web` … registers AllEnterpriseModule exactly once, including across development reloads. … Keep the capability isolated and lifecycle-safe: - `web-ag` must work without `infra`. - Repeated apply operations must be idempotent. - Removing `web-ag` must clean its managed output without removing PostHog …
2 · It ships a plugin
The deliverable is a declarative plugin: a manifest plus managed templates, composed through the framework's existing seams. The core engine and unrelated capabilities stay untouched.
{ "name": "web-ag", "requires": ["web"], "replace": [ { "from": "files/src/_lib/ag.ts", "to": "web/src/_lib/ag.ts" } ], "mergeJson": { "web/package.json": { "dependencies": { "ag-charts-enterprise": "14.0.0", "ag-grid-enterprise": "36.0.0", … } } }, … }3 · Grading is generated behavior
Generate → Execute → Recompose → Remove
The verifier stamps out real projects from the modified framework — alone, combined with other plugins in both orders, reapplied, and removed — and runs the generated code with stubs. Structure is never string-matched against a reference patch.
4% core engine and unrelated capabilities are unchanged 4% web-ag resolves web and renders its managed helper 1% web-ag pins ag-charts-community 1% web-ag pins ag-charts-enterprise + 28 more weighted checks
Current evidence
Task breakdown
Average scores across matching versions of each public sample. The dataset leaderboard also includes private tasks.
| Task | gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11 | GPT 5.6 Solopenai · opencode · 1.18.11 | Muse Spark 1.3 Contributoropencode · opencode · 1.18.11 |
|---|---|---|---|
| OpenAI provisioning + key-policy consolidation for n1xBrowse version 0.2.0 · ff52a388 | 48.0%1 attempt | 80.0%1 attempt | 59.3%± 1.6 pp SE5 attempts · 4 scored |
| OpenAI provisioning for n1xBrowse version 0.2.0 · 83861dd2 | — | — | 93.3%± 3.3 pp SE3 attempts |