Benchmarks built from real client engineering
nplus1 solves bespoke, complex engineering problems for real operating companies. Each benchmark is born from one of these problems, rebuilt as a clean RL environment with a behavioral verifier, and aimed at a weakness we found at the frontier.
Capability benchmarks
Each one targets a capability public suites don't measure, rebuilt from a problem nplus1 solved for a client.
AllocBench
Can coding agents rebuild an allocation engine from its run history?
- Tasks
- 2
- Models
- 3
- Attempts
- 29
DiffBench
Can coding agents understand a business from its database?
- Tasks
- 2
- Models
- 3
- Attempts
- 13
InfraBench
Can coding agents write high-quality infrastructure code?
- Tasks
- 2
- Models
- 3
- Attempts
- 10
PlotterBench
Can coding agents recreate financial widgets from a screenshot?
- Tasks
- 10
- Models
- 4
- Attempts
- 40
Private counterparts to public benchmarks
Familiar public benchmark formats, rebuilt on private codebases no model has trained on.