$ replaybook

Versioned evidence

Benchmark history

Older results still explain model behavior. They remain separate when a stronger verifier or harness changes what a pass means.

OpenRouter core infrastructure baseline

Benchmark 20260906.0.0 · Core tier

Superseded

125/144 durable repairs, 2:36 median, $22.3951 known cost.

OpenRouter Fable 5.1 and Gemini 3.8 mixed infrastructure cohort

Benchmark 20260905.0.0 · Unclassified tier

Superseded

56/66 durable repairs, 2:20 median, $29.6143+ known cost.

OpenCode Go core infrastructure cohort

Benchmark 20260904.0.0 · Core tier

Superseded

102/115 durable repairs, 3:19 median, n/a known cost.

OpenCode Go GPT-5.6 Luna core infrastructure cohort

Benchmark 20260830.0.1 · Core tier

Companion cohort

22/24 durable repairs, 1:54 median, n/a known cost.

OpenRouter GPT-5.6 Luna core infrastructure cohort

Benchmark 20260830.0.0 · Core tier

Superseded

22/24 durable repairs, 1:52 median, $0.8045+ known cost.

Vercel AI Gateway funded infrastructure cohort

Benchmark 20260827.0.1 · Unclassified tier

Superseded

27/27 durable repairs, 2:07 median, $2.2074 known cost.

OpenCode Go visual infrastructure cohort

Benchmark 20260827.0.0 · Unclassified tier

Superseded

34/36 durable repairs, 2:30 median, n/a known cost.

OpenCode Go visual infrastructure cohort

Benchmark 20260826.0.1 · Unclassified tier

Superseded

20/20 durable repairs, 2:16 median, n/a known cost.

Reasoning effort and convergence: three OpenAI models

Benchmark 20260826.0.0 · Unclassified tier

Superseded

8/12 durable repairs, 1:58 median, $1.5424 known cost.

Reasoning effort experiment: Luna on interrupted deploy

Benchmark 20260825.0.2 · Unclassified tier

Superseded

4/6 durable repairs, 1:34 median, $0.1987 known cost.

Visual infrastructure benchmark: deployment timeline

Benchmark 20260825.0.1 · Unclassified tier

Superseded

19/25 durable repairs, 2:25 median, $3.1130+ known cost.

Visual infrastructure benchmark: deployment timeline

Benchmark 20260825.0.0 · Unclassified tier

Superseded

16/22 durable repairs, 2:30 median, $2.5716+ known cost.

Visual infrastructure benchmark: topology and metrics

Benchmark 20260822.0.4 · Unclassified tier

Superseded

32/36 durable repairs, 2:22 median, $2.3708+ known cost.

Four-model visual infrastructure cohort

Benchmark 20260822.0.3 · Unclassified tier

Superseded

12/12 durable repairs, 2:20 median, $0.7839 known cost.

Kimi and Gemini visual infrastructure cohort

Benchmark 20260822.0.2 · Unclassified tier

Superseded

6/6 durable repairs, 2:00 median, $0.8096+ known cost.

Ox Alpha core infrastructure cohort

Benchmark 20260822.0.1 · Core tier

Superseded

22/24 durable repairs, 4:06 median, $0.0000 known cost.

Five-model core infrastructure cohort

Benchmark 20260822.0.0 · Core tier

Superseded

110/120 durable repairs, 2:42 median, $12.1587 known cost.

Qwen 3.8 and GLM 5.3 core infrastructure cohort

Benchmark 20260821.0.1 · Core tier

Superseded

47/48 durable repairs, 2:19 median, $7.5341 known cost.

Ox Alpha full infrastructure cohort

Benchmark 20260821.0.0 · Full tier

Superseded

42/48 durable repairs, 2:50 median, $0.0000 known cost.

OpenCode Go five-scenario infrastructure cohort

Benchmark 20260820.0.1 · Unclassified tier

Superseded

80/86 durable repairs, 3:22 median, n/a known cost.

OpenCode Go infrastructure smoke cohort

Benchmark 20260820.0.0 · Unclassified tier

Companion cohort

14/18 durable repairs, 3:32 median, n/a known cost.

Five models across the full infrastructure suite

Benchmark 20260819.0.0 · Full tier

Superseded

214/240 durable repairs, 2:34 median, $22.4110+ known cost.

GLM 5.3 across the full infrastructure suite

Benchmark 20260818.0.0 · Full tier

Superseded

47/48 durable repairs, 1:38 median, $4.8792 known cost.

Five agents across thirteen infrastructure incidents

Benchmark 20260817.0.3 · Unclassified tier

Superseded

59/65 durable repairs, 2:43 median, $6.2799+ known cost.

Five agents repair Nix store disk pressure

Benchmark 20260817.0.2 · Unclassified tier

Superseded

14/15 durable repairs, 3:46 median, $2.3316+ known cost.

Five agents complete a partial service rollout

Benchmark 20260817.0.1 · Unclassified tier

Superseded

15/15 durable repairs, 2:37 median, $1.3191 known cost.

Five agents face a Discourse plugin boot loop

Benchmark 20260817.0.0 · Unclassified tier

Superseded

12/15 durable repairs, 1:46 median, $0.7299 known cost.

Five agents deploy real Discourse

Benchmark 20260816.0.0 · Frontier tier

Superseded

5/5 durable repairs, 11:30 median, $0.7927 known cost.

Five agents across 13 infrastructure incidents

Benchmark 20260815.0.1 · Unclassified tier

Superseded

59/65 durable repairs, 2:31 median, $13.4734+ known cost.

Four agents repair an interrupted Discourse deployment

Benchmark 20260815.0.0 · Unclassified tier

Superseded

7/12 durable repairs, 4:36 median, $0.9276 known cost.

Eleven incidents across four infrastructure agents

Benchmark 20260812.0.0 · Full tier

Superseded

122/131 durable repairs, 3:04 median, $6.6771+ known cost.

Durable upload recovery on an isolated host

Benchmark 20260811.0.0 · Unclassified tier

Superseded

8/12 durable repairs, 5:54 median, $0.4651 known cost.

Six agent configurations with execution recording

Benchmark 20260810.0.2 · Unclassified tier

Superseded

82/88 durable repairs, 3:24 median, $3.1828+ known cost.

Six agent configurations with execution recording

Benchmark 20260810.0.1 · Unclassified tier

Superseded

82/88 durable repairs, 3:24 median, $3.1828+ known cost.

Three models with execution recording

Benchmark 20260810.0.0 · Unclassified tier

Superseded

39/44 durable repairs, 2:42 median, $2.4652+ known cost.

Four models on the declarative host suite

Benchmark 20260809.0.3 · Unclassified tier

Superseded

53/60 durable repairs, 1:56 median, $0.7964 known cost.

DeepSeek on the declarative host suite

Benchmark 20260809.0.2 · Unclassified tier

Superseded

13/15 durable repairs, 2:28 median, $0.2908 known cost.

Eight models across five stateful incidents

Benchmark 20260809.0.1 · Unclassified tier

Superseded

109/120 durable repairs, 2:16 median, $12.2351+ known cost.

Seven models across five stateful incidents

Benchmark 20260809.0.0 · Unclassified tier

Superseded

94/105 durable repairs, 2:34 median, $5.4983+ known cost.

Harness v5 initial smoke matrix

Harness v5 discounted-model smoke matrix

One attempt per model and scenario, August 8, 2026

Superseded

DeepSeek V4 Pro, Tencent HY3 Preview, and GLM 5.2 each repaired all five scenarios once. The 15/15 result established cost and latency profiles, but the repeated 45-trial baseline exposed reliability differences that a perfect smoke run could not. Total reported cost was $0.2740 and the overall median was 3:26.

Host harness v2 baseline

DeepSeek V4 Flash 0731

Three attempts per scenario, August 8, 2026

Verified
ScenarioVersionRepairsMedianCost
Nginx 502v13/30:55$0.0093
Sidekiq wrong Redis databasev23/32:02$0.0146
Missing Rails migrationv23/34:24$0.0353
Total9/92:02$0.0592

All nine trials survived service restart and host reboot verification. They used 1,703,055 input tokens and 44,979 output tokens. This was the first baseline with the reference repair and verifier absent from the model VM.

Controlled DeepSeek revision comparison

Original V4 Flash and the 0731 revision used the same Replaybook commit, harness, scenarios, Claux release, timeout, concurrency, and attempt count. Only the model ID changed.

MetricOriginal0731Change
Durable repairs8/99/9+1 pass
Median trial time2:132:0211s faster
Input tokens3,488,4671,703,05551% fewer
Output tokens64,44844,97930% fewer
Reported cost$0.0811$0.059227% lower
The failed repair was meaningful. Original V4 Flash added the missing database column manually but did not apply or record the deployed migration. The verifier returned migration_not_applied. The 0731 revision completed the full migration repair in all three attempts.

Host harness v1 archive

Sidekiq wrong Redis database

Scenario v2, three attempts per model, August 7, 2026

Archived

These runs tested backlog preservation, but harness v1 exposed the reference repair inside the VM. They remain useful development records and must not be presented as a controlled model ranking.

ModelRepairsMedianCost
DeepSeek V4 Flash3/31:31$0.0187
Poolside Laguna S 2.13/30:57$0.0133
MiniMax M33/31:49$0.1348
Qwen3.8 Max3/34:52$0.5816
GPT-5.6 Luna2/31:12$0.0236

Archiving does not mean every agent used the leaked repair. It means the harness did not guarantee isolation, so its scores are not comparable to current harness-v5 results.

Why results move here

A verifier change can make yesterday's pass invalid under today's rules. Replaybook keeps the old result, its version, and the behavior it exposed. It does not silently recalculate history or mix incompatible denominators. The full development record lives in benchmarks.md.