nplus1 × Zendo

InfraBench

Can coding agents write high-quality infrastructure code?

InfraBench asks agents to extend the proprietary infra code-generation framework nplus1 uses to build and run infrastructure for its clients, across Terraform, GCP, Vercel and CI. Each task is real work on that framework. We grade it by generating and running real projects from the agent's changes, not by comparing against a reference solution.

InfrastructurePlatform engineering
Tasks
2
Public tasks
2
Models evaluated
3
Trials
10

Benchmark view

Model leaderboard

Average score across tasks, with each task weighted equally.

Mean reward
Average of task means across versions matching the dataset’s rules.
± SE
Standard error across task means; directional with few tasks.
  1. GPT 5.6 Solopenai · opencode · 1.18.11
    80.0%SE —
    1/2 tasks · 1 attempt
    Task means: 80.0%.
  2. Muse Spark 1.3 Contributoropencode · opencode · 1.18.11
    76.3%± 17.0 pp SE
    2/2 tasks · 8 attempts
    Task means: 59.3%, 93.3%.
  3. gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11
    48.0%SE —
    1/2 tasks · 1 attempt
    Task means: 48.0%.

Bars use a fixed 0–100% scale. Ticks mark each task mean.

Anatomy of a task

How InfraBench works

  1. 1 · The agent gets a framework

    A runnable copy of n1x — the code-generation framework that stamps out client applications from composable plugins — plus one engineering instruction taken from real capability work.

    # Add composable AG Grid and AG Charts
    # Enterprise support to n1x
    
    Introduce a `web-ag` capability that requires
    `web` … registers AllEnterpriseModule exactly
    once, including across development reloads. …
    Keep the capability isolated and lifecycle-safe:
    - `web-ag` must work without `infra`.
    - Repeated apply operations must be idempotent.
    - Removing `web-ag` must clean its managed
      output without removing PostHog …
  2. 2 · It ships a plugin

    The deliverable is a declarative plugin: a manifest plus managed templates, composed through the framework's existing seams. The core engine and unrelated capabilities stay untouched.

    {
      "name": "web-ag",
      "requires": ["web"],
      "replace": [ { "from": "files/src/_lib/ag.ts",
        "to": "web/src/_lib/ag.ts" } ],
      "mergeJson": { "web/package.json": {
        "dependencies": {
          "ag-charts-enterprise": "14.0.0",
          "ag-grid-enterprise": "36.0.0", … }
      } }, …
    }
  3. 3 · Grading is generated behavior

    Generate → Execute → Recompose → Remove

    The verifier stamps out real projects from the modified framework — alone, combined with other plugins in both orders, reapplied, and removed — and runs the generated code with stubs. Structure is never string-matched against a reference patch.

    4% core engine and unrelated capabilities are unchanged 4% web-ag resolves web and renders its managed helper 1% web-ag pins ag-charts-community 1% web-ag pins ag-charts-enterprise + 28 more weighted checks

Current evidence

Task breakdown

Average scores across matching versions of each public sample. The dataset leaderboard also includes private tasks.

Average score by public task and model configuration
Taskgemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11GPT 5.6 Solopenai · opencode · 1.18.11Muse Spark 1.3 Contributoropencode · opencode · 1.18.11
OpenAI provisioning + key-policy consolidation for n1xBrowse version 0.2.0 · ff52a38848.0%1 attempt80.0%1 attempt59.3%± 1.6 pp SE5 attempts · 4 scored
OpenAI provisioning for n1xBrowse version 0.2.0 · 83861dd2——93.3%± 3.3 pp SE3 attempts