DiffBench
Can coding agents understand a business from its database?
Inspired by a real business need, DiffBench asks agents to work out what a business actually did between two snapshots of its production database. Imports, renames and records split across locations make a raw row-by-row diff misleading, and the rules for what counts as the same record are never written down. To grade it, we make known changes to the database, such as renaming a site or re-running an import that changes nothing, and check that the agent's report describes exactly what happened.
- Tasks
- 2
- Public tasks
- 2
- Models evaluated
- 3
- Trials
- 13
Benchmark view
Model leaderboard
Average score across tasks, with each task weighted equally.
- Mean reward
- Average of task means across versions matching the dataset’s rules.
- ± SE
- Standard error across task means; directional with few tasks.
- GPT 5.6 Solopenai · opencode · 1.18.1164.8%± 31.7 pp SETask means: 33.0%, 96.5%.
- Muse Spark 1.3 Contributoropencode · opencode · 1.18.1132.0%± 10.0 pp SETask means: 22.0%, 42.0%.
- gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.1148.0%± 15.0 pp SETask means: 33.0%, 63.0%.
Bars use a fixed 0–100% scale. Ticks mark each task mean.
Anatomy of a task
How DiffBench works
1 · Two snapshots, one question
The agent gets the planning platform — a live working database, its point-in-time snapshots, and the report contract the comparison screen expects — with no engine behind it. The job: given any two moments, report what the business actually did between them. Row differences are not the answer. The weekly import rebuilds rows without changing anything; a renamed site is still the same site; a quantity re-split between storage slots leaves every total an operator reads untouched. The engine and the golden it is graded against are extracted from a shipped client system, and the original client asks from that engagement survive as companion fix tasks.
-- the grader makes ONE change -- to a fresh copy of the database: update sites set name = 'Redrock 09B' where name = 'Redrock 09'; -- so one correct account exists: -- 1 changed · Site · name -- 'Redrock 09' -> 'Redrock 09B' -- a raw row diff reports instead: -- 8 records — one site and its three -- phases, each deleted-and-recreated.
2 · The translation is company habit — and habits redraw
What separates a real change from churn is never in the schema. It is a fact about how the deployment works: whether its imports recreate rows or update them, whether storage numbers name equipment or shelf slots, whether reviewers read totals or the splits beneath them. Each entity needs its own answer, evidenced only by the platform's flows and its past reports. Change the habits and every answer moves — which is why the family regenerates: variant worlds redraw the judgment at the same measured difficulty, and how much the shipped history reveals dials difficulty continuously.
// world A — the import recreates rows, // so a phase's identity is its meaning: phases: { keyFields: ["site_customer_name", "site_name", "name"] } // world B — the import updates in place, // so the row id is the identity: phases: { keyFields: ["id"] } // same schema, same ask, different habits. // a perfect world-A engine scores 0.4 on // world B. sol: 0.33 (A) vs 0.41 (B).3 · Graded against known answers
The verifier builds two fresh copies of the database and makes one known change — renames a site, replays the import, moves quantity between slots, or touches nothing at all — so there is exactly one correct account of what the business did. The model's report and the shipped engine's report are compared record by record; naming and internal id formats are free, the drawn conclusions are not. Each scenario carries a weighted slice of the score, so partial understanding earns partial credit: frontier models land near 0.35 on fresh worlds, fail with consistent signatures, and never sweep.
Current evidence
Task breakdown
Average scores across matching versions of each public sample. The dataset leaderboard also includes private tasks.
| Task | gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11 | GPT 5.6 Solopenai · opencode · 1.18.11 | Muse Spark 1.3 Contributoropencode · opencode · 1.18.11 |
|---|---|---|---|
| dc-0727-version-diffBrowse version 0.1.0 · 01804078 | 33.0%1 attempt | 33.0%1 attempt | 22.0%± 11.0 pp SE3 attempts |
| dc-0727-version-diff-hBrowse version 0.1.0 · 97de9fdd | 63.0%± 0.0 pp SE3 attempts · 2 scored | 96.5%± 3.5 pp SE2 attempts | 42.0%± 21.0 pp SE3 attempts |