$ replaybook

Versioned evidence

Benchmark history

Older results still explain model behavior. They remain separate when a stronger verifier or harness changes what a pass means.

Four agents repair an interrupted Discourse deployment

Benchmark 20260815.0.0

Superseded

7/12 durable repairs, 4:36 median, $0.9276 known cost.

Eleven incidents across four infrastructure agents

Benchmark 20260812.0.0

Superseded

122/131 durable repairs, 3:04 median, $6.6771+ known cost.

Durable upload recovery on an isolated host

Benchmark 20260811.0.0

Superseded

8/12 durable repairs, 5:54 median, $0.4651 known cost.

Six agent configurations with execution recording

Benchmark 20260810.0.2

Superseded

82/88 durable repairs, 3:24 median, $3.1828+ known cost.

Six agent configurations with execution recording

Benchmark 20260810.0.1

Superseded

82/88 durable repairs, 3:24 median, $3.1828+ known cost.

Three models with execution recording

Benchmark 20260810.0.0

Superseded

39/44 durable repairs, 2:42 median, $2.4652+ known cost.

Four models on the declarative host suite

Benchmark 20260809.0.3

Superseded

53/60 durable repairs, 1:56 median, $0.7964 known cost.

DeepSeek on the declarative host suite

Benchmark 20260809.0.2

Superseded

13/15 durable repairs, 2:28 median, $0.2908 known cost.

Eight models across five stateful incidents

Benchmark 20260809.0.1

Superseded

109/120 durable repairs, 2:16 median, $12.2351+ known cost.

Seven models across five stateful incidents

Benchmark 20260809.0.0

Superseded

94/105 durable repairs, 2:34 median, $5.4983+ known cost.

Harness v5 initial smoke matrix

Harness v5 discounted-model smoke matrix

One attempt per model and scenario, August 8, 2026

Superseded

DeepSeek V4 Pro, Tencent HY3 Preview, and GLM 5.2 each repaired all five scenarios once. The 15/15 result established cost and latency profiles, but the repeated 45-trial baseline exposed reliability differences that a perfect smoke run could not. Total reported cost was $0.2740 and the overall median was 3:26.

Host harness v2 baseline

DeepSeek V4 Flash 0731

Three attempts per scenario, August 8, 2026

Verified
ScenarioVersionRepairsMedianCost
Nginx 502v13/30:55$0.0093
Sidekiq wrong Redis databasev23/32:02$0.0146
Missing Rails migrationv23/34:24$0.0353
Total9/92:02$0.0592

All nine trials survived service restart and host reboot verification. They used 1,703,055 input tokens and 44,979 output tokens. This was the first baseline with the reference repair and verifier absent from the model VM.

Controlled DeepSeek revision comparison

Original V4 Flash and the 0731 revision used the same Replaybook commit, harness, scenarios, Claux release, timeout, concurrency, and attempt count. Only the model ID changed.

MetricOriginal0731Change
Durable repairs8/99/9+1 pass
Median trial time2:132:0211s faster
Input tokens3,488,4671,703,05551% fewer
Output tokens64,44844,97930% fewer
Reported cost$0.0811$0.059227% lower
The failed repair was meaningful. Original V4 Flash added the missing database column manually but did not apply or record the deployed migration. The verifier returned migration_not_applied. The 0731 revision completed the full migration repair in all three attempts.

Host harness v1 archive

Sidekiq wrong Redis database

Scenario v2, three attempts per model, August 7, 2026

Archived

These runs tested backlog preservation, but harness v1 exposed the reference repair inside the VM. They remain useful development records and must not be presented as a controlled model ranking.

ModelRepairsMedianCost
DeepSeek V4 Flash3/31:31$0.0187
Poolside Laguna S 2.13/30:57$0.0133
MiniMax M33/31:49$0.1348
Qwen3.8 Max3/34:52$0.5816
GPT-5.6 Luna2/31:12$0.0236

Archiving does not mean every agent used the leaked repair. It means the harness did not guarantee isolation, so its scores are not comparable to current harness-v5 results.

Why results move here

A verifier change can make yesterday's pass invalid under today's rules. Replaybook keeps the old result, its version, and the behavior it exposed. It does not silently recalculate history or mix incompatible denominators. The full development record lives in benchmarks.md.