$ replaybook

Benchmark 20260815.0.1

Five agents across 13 infrastructure incidents

Five infrastructure agents each attempted 13 host-native incidents once under the same frozen harness and scenario pack. The suite spans Nginx, Ruby, Rails, Sidekiq, Node, Rust, Python, deployment, authentication, and Discourse-shaped failures, with every repair verified immediately, after service restart, and after host reboot.

59/65durable repairs
2:31overall median
$13.4734+known total cost
65 trials across 4 controlled matrices. 59 repairs passed durable verification. 6 evaluated attempts failed and 0 trials were unavailable.

Model summary

ModelRepairsPass rateMedianInput tokensKnown costCost / repair
GLM 5.2 (high)13/13100%2:312,907,050$0.5196$0.0400
DeepSeek V4 Flash 0731 (high)12/1392%4:334,293,355$0.1029$0.0086
GPT-5.6 Luna (high)12/1392%1:567,161,480$0.1929+$0.0161+
Claude Sonnet 5 (high)12/1392%2:034,839,113$10.8662$0.9055
Gemini 3.7 Flash (high)10/1377%3:0915,163,587$1.7917$0.1792

Execution recording

Medians across trials with transcript schema v2 recording. First non-read is time before the first potentially mutating tool call; after non-read is the remaining agent time. Model and tool time can overlap.

ModelRecordedRoundsModel timeToolsTool timeFirst non-readAfter non-read
GLM 5.2 (high)13/13162:04250:100:032:26
DeepSeek V4 Flash 0731 (high)13/13173:52230:130:074:26
GPT-5.6 Luna (high)12/13171:20340:130:061:42
Claude Sonnet 5 (high)13/13171:58220:110:051:58
Gemini 3.7 Flash (high)13/13402:36390:050:053:03

GLM 5.2 made 13 durable repairs in 13 attempts with a 2:31 median. It was the only model to sweep the suite and spent $0.5196, about $0.0400 per durable repair.

DeepSeek V4 Flash 0731 made 12 durable repairs in 13 attempts and spent $0.1029, about $0.0086 per repair. It was the least expensive repair agent in this cohort, though its 4:33 median was the slowest among the four non-Sonnet fleet models.

GPT-5.6 Luna made 12 durable repairs in 13 attempts with the fastest median at 1:56. Its known spend was $0.1929, at least $0.0161 per durable repair; the missing timeout usage means the true value is higher.

Claude Sonnet 5 made 12 durable repairs in 13 attempts with a 2:03 median. It spent $10.8662, about $0.9055 per durable repair, making its observed repairs roughly 105 times as expensive as DeepSeek's in this cohort.

Gemini 3.7 Flash made 10 durable repairs in 13 attempts with a 3:09 median and spent $1.7917, about $0.1792 per durable repair. Its three failures were later provider rejections after meaningful inference: one corrupted thought signature and two content-policy responses. Replaybook counts them as evaluated failures while preserving that distinction from verifier failures.

The remaining failures separated cleanly: DeepSeek did not converge the interrupted deployment, Luna timed out with exports still failing, and Sonnet did not converge the interrupted deployment.

Scenario breakdown

ScenarioVersionGLM 5.2 (high)DeepSeek V4 Flash 0731 (high)GPT-5.6 Luna (high)Claude Sonnet 5 (high)Gemini 3.7 Flash (high)
001-nginx-502-hostv11/1 · 0:571/1 · 1:571/1 · 3:051/1 · 1:041/1 · 1:34
013-sidekiq-wrong-redisv21/1 · 1:121/1 · 5:391/1 · 1:241/1 · 1:511/1 · 2:02
014-missing-rails-migrationv21/1 · 2:381/1 · 5:011/1 · 2:061/1 · 1:521/1 · 3:40
015-sidekiq-poison-pillv11/1 · 3:441/1 · 5:141/1 · 3:291/1 · 4:010/1 · 1:00
016-rails-pool-exhaustionv11/1 · 1:261/1 · 2:491/1 · 1:231/1 · 1:511/1 · 3:17
017-partial-rails-rolloutv11/1 · 2:311/1 · 4:331/1 · 1:451/1 · 3:530/1 · 1:25
018-node-event-loop-blockingv11/1 · 2:521/1 · 3:481/1 · 2:161/1 · 2:031/1 · 3:09
019-rust-fd-leakv11/1 · 2:141/1 · 4:081/1 · 1:331/1 · 2:141/1 · 1:59
020-python-gunicorn-saturationv11/1 · 4:201/1 · 10:370/1 · 15:311/1 · 2:091/1 · 4:12
021-discourse-shared-uploadsv11/1 · 2:421/1 · 11:191/1 · 2:251/1 · 4:131/1 · 3:56
022-discourse-multisite-migrationv11/1 · 1:191/1 · 3:461/1 · 1:081/1 · 1:510/1 · 1:39
023-auth-secret-rolloutv11/1 · 2:491/1 · 3:171/1 · 1:311/1 · 2:031/1 · 4:39
024-discourse-interrupted-deployv11/1 · 1:570/1 · 5:041/1 · 1:560/1 · 5:191/1 · 3:15

Failure categories

Run notes

Constituent matrices

The publisher recorded harness provenance and validated matching scenario pack revisions, scenario versions, attempts, timeout, agent adapter, and Claux release before combining these summaries.

Scenario packs: ducks/replaybook-infra@20260814.0.0.

MatrixModelsReplaybook commit
host-matrix-2026-08-15__21-56-08.20e83fDeepSeek V4 Flash 0731, Gemini 3.7 Flash, GPT-5.6 Luna, GLM 5.2 · reasoning high7f4b117a
host-matrix-2026-08-15__21-07-50.e717e8Claude Sonnet 5 · reasoning high7f4b117a
host-matrix-2026-08-15__21-28-32.898d2aClaude Sonnet 5 · reasoning high7f4b117a
host-matrix-2026-08-15__20-18-41.a9a4e6Claude Sonnet 5 · reasoning high7f4b117a

Run the matrix

python integrations/host/run_host_matrix.py \
  --scenario 001-nginx-502-host \
  --scenario 013-sidekiq-wrong-redis \
  --scenario 014-missing-rails-migration \
  --scenario 015-sidekiq-poison-pill \
  --scenario 016-rails-pool-exhaustion \
  --scenario 017-partial-rails-rollout \
  --scenario 018-node-event-loop-blocking \
  --scenario 019-rust-fd-leak \
  --scenario 020-python-gunicorn-saturation \
  --scenario 021-discourse-shared-uploads \
  --scenario 022-discourse-multisite-migration \
  --scenario 023-auth-secret-rollout \
  --scenario 024-discourse-interrupted-deploy \
  --models \
    z-ai/glm-5.2 \
    deepseek/deepseek-v4-flash-0731 \
    openai/gpt-5.6-luna \
    anthropic/claude-sonnet-5 \
    google/gemini-3.7-flash \
  --reasoning-efforts high \
  --attempts 1 \
  --concurrency 2

Host harness v19, Claux v20260815.0.0, 900-second agent timeout. Usage was reported for 64 of 65 trials.

Read the methodology, browse the versioned history, or inspect the complete benchmark record.