Public task
dc-0727-version-diff-h
Build the comparison engine from precedent alone: past reports and the snapshots that produced them are the only specification of what counts as a real change.
Part of DiffBench ↗
- Model configurations
- 3
- Evaluation attempts
- 8
- Public examples
- 8
Current task results
Evaluation comparison
Published evaluation results for this task version. Each bar shows a model configuration’s average score.
| Model | Average score | Score range | Attempts |
|---|---|---|---|
| GPT 5.6 Solopenai · opencode · 1.18.11 | 96.5%± 3.5 pp SE | 93.0% – 100.0% | 22 scored |
| gemini-3.1-pro-previewgoogle-vertex · opencode · 1.18.11 | 63.0%± 0.0 pp SE | 63.0% – 63.0% | 32 scored |
| Muse Spark 1.3 Contributoropencode · opencode · 1.18.11 | 42.0%± 21.0 pp SE | 0.0% – 63.0% | 33 scored |
Averages and ranges use scored attempts on this exact version. Attempts without a score are counted separately. SE measures uncertainty across attempt scores.
The task
What the agent receives# Build the version diff review The platform snapshots the working database on a schedule (see `platform/snapshots/`), and operators review two versions against each other before deciding anything: "pull up today against last week's snapshot and tell me everything that changed, so I can decide whether to roll back." Implement `lib/versionDiff.mjs` exporting `computeVersionDiff` and `computeVersionDiffDetails` as consumed by `bin/diff.mjs`, matching the response contracts in `contracts/versionDiff.schema.mjs`. The review's behavior is specified by precedent: `assets/history/` holds the review exports operators saw for past snapshot pairs, and `platform/snapshots/` holds the snapshots that produced them. New reviews must be consistent with that precedent on databases the history has never seen. Check your work with `node bin/diff.mjs snap-2026-07-19-0600` and `node contracts/validate.mjs <saved-output.json>`.