nplus1 × Zendo

DiffBench

Can coding agents understand a business from its database?

Inspired by a real business need, DiffBench asks agents to work out what a business actually did between two snapshots of its production database. Imports, renames and records split across locations make a raw row-by-row diff misleading, and the rules for what counts as the same record are never written down. To grade it, we make known changes to the database, such as renaming a site or re-running an import that changes nothing, and check that the agent's report describes exactly what happened.

Data analysis
Tasks
2
Public tasks
2
Models evaluated
3
Trials
13

Benchmark view

Model leaderboard

Average score across tasks, with each task weighted equally.

Mean reward
Average of task means across versions matching the dataset’s rules.
± SE
Standard error across task means; directional with few tasks.
  1. GPT 5.6 Solopenai · opencode · 1.18.11
    64.8%± 31.7 pp SE
    2/2 tasks · 3 attempts
    Task means: 33.0%, 96.5%.
  2. Muse Spark 1.3 Contributoropencode · opencode · 1.18.11
    32.0%± 10.0 pp SE
    2/2 tasks · 6 attempts
    Task means: 22.0%, 42.0%.
  3. gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11
    48.0%± 15.0 pp SE
    2/2 tasks · 4 attempts
    Task means: 33.0%, 63.0%.

Bars use a fixed 0–100% scale. Ticks mark each task mean.

Anatomy of a task

How DiffBench works

  1. 1 · Two snapshots, one question

    The agent gets the planning platform — a live working database, its point-in-time snapshots, and the report contract the comparison screen expects — with no engine behind it. The job: given any two moments, report what the business actually did between them. Row differences are not the answer. The weekly import rebuilds rows without changing anything; a renamed site is still the same site; a quantity re-split between storage slots leaves every total an operator reads untouched. The engine and the golden it is graded against are extracted from a shipped client system, and the original client asks from that engagement survive as companion fix tasks.

    -- the grader makes ONE change
    -- to a fresh copy of the database:
    update sites set name = 'Redrock 09B'
     where name = 'Redrock 09';
    
    -- so one correct account exists:
    --   1 changed · Site · name
    --   'Redrock 09' -> 'Redrock 09B'
    
    -- a raw row diff reports instead:
    --   8 records — one site and its three
    --   phases, each deleted-and-recreated.
  2. 2 · The translation is company habit — and habits redraw

    What separates a real change from churn is never in the schema. It is a fact about how the deployment works: whether its imports recreate rows or update them, whether storage numbers name equipment or shelf slots, whether reviewers read totals or the splits beneath them. Each entity needs its own answer, evidenced only by the platform's flows and its past reports. Change the habits and every answer moves — which is why the family regenerates: variant worlds redraw the judgment at the same measured difficulty, and how much the shipped history reveals dials difficulty continuously.

    // world A — the import recreates rows,
    // so a phase's identity is its meaning:
    phases: { keyFields: ["site_customer_name",
                          "site_name", "name"] }
    
    // world B — the import updates in place,
    // so the row id is the identity:
    phases: { keyFields: ["id"] }
    
    // same schema, same ask, different habits.
    // a perfect world-A engine scores 0.4 on
    // world B. sol: 0.33 (A) vs 0.41 (B).
  3. 3 · Graded against known answers

    The verifier builds two fresh copies of the database and makes one known change — renames a site, replays the import, moves quantity between slots, or touches nothing at all — so there is exactly one correct account of what the business did. The model's report and the shipped engine's report are compared record by record; naming and internal id formats are free, the drawn conclusions are not. Each scenario carries a weighted slice of the score, so partial understanding earns partial credit: frontier models land near 0.35 on fresh worlds, fail with consistent signatures, and never sweep.

Current evidence

Task breakdown

Average scores across matching versions of each public sample. The dataset leaderboard also includes private tasks.

Average score by public task and model configuration
Taskgemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11GPT 5.6 Solopenai · opencode · 1.18.11Muse Spark 1.3 Contributoropencode · opencode · 1.18.11
dc-0727-version-diffBrowse version 0.1.0 · 0180407833.0%1 attempt33.0%1 attempt22.0%± 11.0 pp SE3 attempts
dc-0727-version-diff-hBrowse version 0.1.0 · 97de9fdd63.0%± 0.0 pp SE3 attempts · 2 scored96.5%± 3.5 pp SE2 attempts42.0%± 21.0 pp SE3 attempts