Versioned evidence
Older results still explain model behavior. They remain separate when a stronger verifier or harness changes what a pass means.
Benchmark 20260815.0.0
7/12 durable repairs, 4:36 median, $0.9276 known cost.
Benchmark 20260812.0.0
122/131 durable repairs, 3:04 median, $6.6771+ known cost.
Benchmark 20260811.0.0
8/12 durable repairs, 5:54 median, $0.4651 known cost.
Benchmark 20260810.0.2
82/88 durable repairs, 3:24 median, $3.1828+ known cost.
Benchmark 20260810.0.1
82/88 durable repairs, 3:24 median, $3.1828+ known cost.
Benchmark 20260810.0.0
39/44 durable repairs, 2:42 median, $2.4652+ known cost.
Benchmark 20260809.0.3
53/60 durable repairs, 1:56 median, $0.7964 known cost.
Benchmark 20260809.0.2
13/15 durable repairs, 2:28 median, $0.2908 known cost.
Benchmark 20260809.0.1
109/120 durable repairs, 2:16 median, $12.2351+ known cost.
Benchmark 20260809.0.0
94/105 durable repairs, 2:34 median, $5.4983+ known cost.
One attempt per model and scenario, August 8, 2026
DeepSeek V4 Pro, Tencent HY3 Preview, and GLM 5.2 each repaired all five scenarios once. The 15/15 result established cost and latency profiles, but the repeated 45-trial baseline exposed reliability differences that a perfect smoke run could not. Total reported cost was $0.2740 and the overall median was 3:26.
Three attempts per scenario, August 8, 2026
| Scenario | Version | Repairs | Median | Cost |
|---|---|---|---|---|
| Nginx 502 | v1 | 3/3 | 0:55 | $0.0093 |
| Sidekiq wrong Redis database | v2 | 3/3 | 2:02 | $0.0146 |
| Missing Rails migration | v2 | 3/3 | 4:24 | $0.0353 |
| Total | 9/9 | 2:02 | $0.0592 |
All nine trials survived service restart and host reboot verification. They used 1,703,055 input tokens and 44,979 output tokens. This was the first baseline with the reference repair and verifier absent from the model VM.
Original V4 Flash and the 0731 revision used the same Replaybook commit, harness, scenarios, Claux release, timeout, concurrency, and attempt count. Only the model ID changed.
| Metric | Original | 0731 | Change |
|---|---|---|---|
| Durable repairs | 8/9 | 9/9 | +1 pass |
| Median trial time | 2:13 | 2:02 | 11s faster |
| Input tokens | 3,488,467 | 1,703,055 | 51% fewer |
| Output tokens | 64,448 | 44,979 | 30% fewer |
| Reported cost | $0.0811 | $0.0592 | 27% lower |
migration_not_applied. The 0731 revision completed the full
migration repair in all three attempts.
Scenario v2, three attempts per model, August 7, 2026
These runs tested backlog preservation, but harness v1 exposed the reference repair inside the VM. They remain useful development records and must not be presented as a controlled model ranking.
| Model | Repairs | Median | Cost |
|---|---|---|---|
| DeepSeek V4 Flash | 3/3 | 1:31 | $0.0187 |
| Poolside Laguna S 2.1 | 3/3 | 0:57 | $0.0133 |
| MiniMax M3 | 3/3 | 1:49 | $0.1348 |
| Qwen3.8 Max | 3/3 | 4:52 | $0.5816 |
| GPT-5.6 Luna | 2/3 | 1:12 | $0.0236 |
Archiving does not mean every agent used the leaked repair. It means the harness did not guarantee isolation, so its scores are not comparable to current harness-v5 results.
A verifier change can make yesterday's pass invalid under today's rules.
Replaybook keeps the old result, its version, and the behavior it exposed.
It does not silently recalculate history or mix incompatible denominators.
The full development record lives in
benchmarks.md.