Fieldwork
Can agents ship a real production feature in a private codebase?
PrivateDeepSWE is built from the recorded sessions of engineers and business teams shipping real features with coding agents in private production codebases. The tasks model DeepSWE in real-world complexity, diversity and verification. Each task starts before a feature shipped, and the agent's brief includes customer context and the engineer's real direction. Wherever the engineer had to step in and correct the agent, that correction informs the verifiers. We grade what the code does, using the same system boundaries the real feature had to satisfy.
- Tasks
- 6
- Public tasks
- 0
- Models evaluated
- 1
- Trials
- 61
Benchmark view
Model leaderboard
Average score across tasks, with each task weighted equally.
- Mean reward
- Average of task means across versions matching the dataset’s rules.
- ± SE
- Standard error across task means; directional with few tasks.
- V9 Learnabilityzendo · grok-build · 1.0.4645.3%± 11.9 pp SETask means: 0.0%, 16.7%, 61.1%, 62.5%, 64.7%, 66.7%.
Bars use a fixed 0–100% scale. Ticks mark each task mean.
Anatomy of a task
How PrivateDeepSWE works
1 · A real feature in a real codebase
Every task comes from a feature that actually shipped in a private production codebase. We start just before the feature landed and give the agent a short, realistic brief. The agent must understand the existing system and build a solution that fits it.
2 · The existing system provides the context
The brief does not explain everything. Agents must learn from the code, tests, data models, and neighboring features already in the repository. We grade through the same public boundaries the production system uses—not by requiring a particular implementation.
3 · Behavior determines the score
Each task has a reference implementation and a suite of real-world scenarios. We first check that the existing system still works, then test the feature, its edge cases, and how it behaves when things change or fail. Any implementation can pass if it produces the right behavior.
Current evidence
Task breakdown
Average scores across matching versions of each public sample. The dataset leaderboard also includes private tasks.
Publish task samples to show their individual results. Dataset averages include private tasks.