Versioned evidence
Older results still explain model behavior. They remain separate when a stronger verifier or harness changes what a pass means.
Benchmark 20260906.0.0 · Core tier
125/144 durable repairs, 2:36 median, $22.3951 known cost.
Benchmark 20260905.0.0 · Unclassified tier
56/66 durable repairs, 2:20 median, $29.6143+ known cost.
Benchmark 20260904.0.0 · Core tier
102/115 durable repairs, 3:19 median, n/a known cost.
Benchmark 20260830.0.1 · Core tier
22/24 durable repairs, 1:54 median, n/a known cost.
Benchmark 20260830.0.0 · Core tier
22/24 durable repairs, 1:52 median, $0.8045+ known cost.
Benchmark 20260827.0.1 · Unclassified tier
27/27 durable repairs, 2:07 median, $2.2074 known cost.
Benchmark 20260827.0.0 · Unclassified tier
34/36 durable repairs, 2:30 median, n/a known cost.
Benchmark 20260826.0.1 · Unclassified tier
20/20 durable repairs, 2:16 median, n/a known cost.
Benchmark 20260826.0.0 · Unclassified tier
8/12 durable repairs, 1:58 median, $1.5424 known cost.
Benchmark 20260825.0.2 · Unclassified tier
4/6 durable repairs, 1:34 median, $0.1987 known cost.
Benchmark 20260825.0.1 · Unclassified tier
19/25 durable repairs, 2:25 median, $3.1130+ known cost.
Benchmark 20260825.0.0 · Unclassified tier
16/22 durable repairs, 2:30 median, $2.5716+ known cost.
Benchmark 20260822.0.4 · Unclassified tier
32/36 durable repairs, 2:22 median, $2.3708+ known cost.
Benchmark 20260822.0.3 · Unclassified tier
12/12 durable repairs, 2:20 median, $0.7839 known cost.
Benchmark 20260822.0.2 · Unclassified tier
6/6 durable repairs, 2:00 median, $0.8096+ known cost.
Benchmark 20260822.0.1 · Core tier
22/24 durable repairs, 4:06 median, $0.0000 known cost.
Benchmark 20260822.0.0 · Core tier
110/120 durable repairs, 2:42 median, $12.1587 known cost.
Benchmark 20260821.0.1 · Core tier
47/48 durable repairs, 2:19 median, $7.5341 known cost.
Benchmark 20260821.0.0 · Full tier
42/48 durable repairs, 2:50 median, $0.0000 known cost.
Benchmark 20260820.0.1 · Unclassified tier
80/86 durable repairs, 3:22 median, n/a known cost.
Benchmark 20260820.0.0 · Unclassified tier
14/18 durable repairs, 3:32 median, n/a known cost.
Benchmark 20260819.0.0 · Full tier
214/240 durable repairs, 2:34 median, $22.4110+ known cost.
Benchmark 20260818.0.0 · Full tier
47/48 durable repairs, 1:38 median, $4.8792 known cost.
Benchmark 20260817.0.3 · Unclassified tier
59/65 durable repairs, 2:43 median, $6.2799+ known cost.
Benchmark 20260817.0.2 · Unclassified tier
14/15 durable repairs, 3:46 median, $2.3316+ known cost.
Benchmark 20260817.0.1 · Unclassified tier
15/15 durable repairs, 2:37 median, $1.3191 known cost.
Benchmark 20260817.0.0 · Unclassified tier
12/15 durable repairs, 1:46 median, $0.7299 known cost.
Benchmark 20260816.0.0 · Frontier tier
5/5 durable repairs, 11:30 median, $0.7927 known cost.
Benchmark 20260815.0.1 · Unclassified tier
59/65 durable repairs, 2:31 median, $13.4734+ known cost.
Benchmark 20260815.0.0 · Unclassified tier
7/12 durable repairs, 4:36 median, $0.9276 known cost.
Benchmark 20260812.0.0 · Full tier
122/131 durable repairs, 3:04 median, $6.6771+ known cost.
Benchmark 20260811.0.0 · Unclassified tier
8/12 durable repairs, 5:54 median, $0.4651 known cost.
Benchmark 20260810.0.2 · Unclassified tier
82/88 durable repairs, 3:24 median, $3.1828+ known cost.
Benchmark 20260810.0.1 · Unclassified tier
82/88 durable repairs, 3:24 median, $3.1828+ known cost.
Benchmark 20260810.0.0 · Unclassified tier
39/44 durable repairs, 2:42 median, $2.4652+ known cost.
Benchmark 20260809.0.3 · Unclassified tier
53/60 durable repairs, 1:56 median, $0.7964 known cost.
Benchmark 20260809.0.2 · Unclassified tier
13/15 durable repairs, 2:28 median, $0.2908 known cost.
Benchmark 20260809.0.1 · Unclassified tier
109/120 durable repairs, 2:16 median, $12.2351+ known cost.
Benchmark 20260809.0.0 · Unclassified tier
94/105 durable repairs, 2:34 median, $5.4983+ known cost.
One attempt per model and scenario, August 8, 2026
DeepSeek V4 Pro, Tencent HY3 Preview, and GLM 5.2 each repaired all five scenarios once. The 15/15 result established cost and latency profiles, but the repeated 45-trial baseline exposed reliability differences that a perfect smoke run could not. Total reported cost was $0.2740 and the overall median was 3:26.
Three attempts per scenario, August 8, 2026
| Scenario | Version | Repairs | Median | Cost |
|---|---|---|---|---|
| Nginx 502 | v1 | 3/3 | 0:55 | $0.0093 |
| Sidekiq wrong Redis database | v2 | 3/3 | 2:02 | $0.0146 |
| Missing Rails migration | v2 | 3/3 | 4:24 | $0.0353 |
| Total | 9/9 | 2:02 | $0.0592 |
All nine trials survived service restart and host reboot verification. They used 1,703,055 input tokens and 44,979 output tokens. This was the first baseline with the reference repair and verifier absent from the model VM.
Original V4 Flash and the 0731 revision used the same Replaybook commit, harness, scenarios, Claux release, timeout, concurrency, and attempt count. Only the model ID changed.
| Metric | Original | 0731 | Change |
|---|---|---|---|
| Durable repairs | 8/9 | 9/9 | +1 pass |
| Median trial time | 2:13 | 2:02 | 11s faster |
| Input tokens | 3,488,467 | 1,703,055 | 51% fewer |
| Output tokens | 64,448 | 44,979 | 30% fewer |
| Reported cost | $0.0811 | $0.0592 | 27% lower |
migration_not_applied. The 0731 revision completed the full
migration repair in all three attempts.
Scenario v2, three attempts per model, August 7, 2026
These runs tested backlog preservation, but harness v1 exposed the reference repair inside the VM. They remain useful development records and must not be presented as a controlled model ranking.
| Model | Repairs | Median | Cost |
|---|---|---|---|
| DeepSeek V4 Flash | 3/3 | 1:31 | $0.0187 |
| Poolside Laguna S 2.1 | 3/3 | 0:57 | $0.0133 |
| MiniMax M3 | 3/3 | 1:49 | $0.1348 |
| Qwen3.8 Max | 3/3 | 4:52 | $0.5816 |
| GPT-5.6 Luna | 2/3 | 1:12 | $0.0236 |
Archiving does not mean every agent used the leaked repair. It means the harness did not guarantee isolation, so its scores are not comparable to current harness-v5 results.
A verifier change can make yesterday's pass invalid under today's rules.
Replaybook keeps the old result, its version, and the behavior it exposed.
It does not silently recalculate history or mix incompatible denominators.
The full development record lives in
benchmarks.md.