Benchmark 20260815.0.1
Five infrastructure agents each attempted 13 host-native incidents once under the same frozen harness and scenario pack. The suite spans Nginx, Ruby, Rails, Sidekiq, Node, Rust, Python, deployment, authentication, and Discourse-shaped failures, with every repair verified immediately, after service restart, and after host reboot.
| Model | Repairs | Pass rate | Median | Input tokens | Known cost | Cost / repair |
|---|---|---|---|---|---|---|
| GLM 5.2 (high) | 13/13 | 100% | 2:31 | 2,907,050 | $0.5196 | $0.0400 |
| DeepSeek V4 Flash 0731 (high) | 12/13 | 92% | 4:33 | 4,293,355 | $0.1029 | $0.0086 |
| GPT-5.6 Luna (high) | 12/13 | 92% | 1:56 | 7,161,480 | $0.1929+ | $0.0161+ |
| Claude Sonnet 5 (high) | 12/13 | 92% | 2:03 | 4,839,113 | $10.8662 | $0.9055 |
| Gemini 3.7 Flash (high) | 10/13 | 77% | 3:09 | 15,163,587 | $1.7917 | $0.1792 |
Medians across trials with transcript schema v2 recording. First non-read is time before the first potentially mutating tool call; after non-read is the remaining agent time. Model and tool time can overlap.
| Model | Recorded | Rounds | Model time | Tools | Tool time | First non-read | After non-read |
|---|---|---|---|---|---|---|---|
| GLM 5.2 (high) | 13/13 | 16 | 2:04 | 25 | 0:10 | 0:03 | 2:26 |
| DeepSeek V4 Flash 0731 (high) | 13/13 | 17 | 3:52 | 23 | 0:13 | 0:07 | 4:26 |
| GPT-5.6 Luna (high) | 12/13 | 17 | 1:20 | 34 | 0:13 | 0:06 | 1:42 |
| Claude Sonnet 5 (high) | 13/13 | 17 | 1:58 | 22 | 0:11 | 0:05 | 1:58 |
| Gemini 3.7 Flash (high) | 13/13 | 40 | 2:36 | 39 | 0:05 | 0:05 | 3:03 |
GLM 5.2 made 13 durable repairs in 13 attempts with a 2:31 median. It was the only model to sweep the suite and spent $0.5196, about $0.0400 per durable repair.
DeepSeek V4 Flash 0731 made 12 durable repairs in 13 attempts and spent $0.1029, about $0.0086 per repair. It was the least expensive repair agent in this cohort, though its 4:33 median was the slowest among the four non-Sonnet fleet models.
GPT-5.6 Luna made 12 durable repairs in 13 attempts with the fastest median at 1:56. Its known spend was $0.1929, at least $0.0161 per durable repair; the missing timeout usage means the true value is higher.
Claude Sonnet 5 made 12 durable repairs in 13 attempts with a 2:03 median. It spent $10.8662, about $0.9055 per durable repair, making its observed repairs roughly 105 times as expensive as DeepSeek's in this cohort.
Gemini 3.7 Flash made 10 durable repairs in 13 attempts with a 3:09 median and spent $1.7917, about $0.1792 per durable repair. Its three failures were later provider rejections after meaningful inference: one corrupted thought signature and two content-policy responses. Replaybook counts them as evaluated failures while preserving that distinction from verifier failures.
The remaining failures separated cleanly: DeepSeek did not converge the interrupted deployment, Luna timed out with exports still failing, and Sonnet did not converge the interrupted deployment.
| Scenario | Version | GLM 5.2 (high) | DeepSeek V4 Flash 0731 (high) | GPT-5.6 Luna (high) | Claude Sonnet 5 (high) | Gemini 3.7 Flash (high) |
|---|---|---|---|---|---|---|
| 001-nginx-502-host | v1 | 1/1 · 0:57 | 1/1 · 1:57 | 1/1 · 3:05 | 1/1 · 1:04 | 1/1 · 1:34 |
| 013-sidekiq-wrong-redis | v2 | 1/1 · 1:12 | 1/1 · 5:39 | 1/1 · 1:24 | 1/1 · 1:51 | 1/1 · 2:02 |
| 014-missing-rails-migration | v2 | 1/1 · 2:38 | 1/1 · 5:01 | 1/1 · 2:06 | 1/1 · 1:52 | 1/1 · 3:40 |
| 015-sidekiq-poison-pill | v1 | 1/1 · 3:44 | 1/1 · 5:14 | 1/1 · 3:29 | 1/1 · 4:01 | 0/1 · 1:00 |
| 016-rails-pool-exhaustion | v1 | 1/1 · 1:26 | 1/1 · 2:49 | 1/1 · 1:23 | 1/1 · 1:51 | 1/1 · 3:17 |
| 017-partial-rails-rollout | v1 | 1/1 · 2:31 | 1/1 · 4:33 | 1/1 · 1:45 | 1/1 · 3:53 | 0/1 · 1:25 |
| 018-node-event-loop-blocking | v1 | 1/1 · 2:52 | 1/1 · 3:48 | 1/1 · 2:16 | 1/1 · 2:03 | 1/1 · 3:09 |
| 019-rust-fd-leak | v1 | 1/1 · 2:14 | 1/1 · 4:08 | 1/1 · 1:33 | 1/1 · 2:14 | 1/1 · 1:59 |
| 020-python-gunicorn-saturation | v1 | 1/1 · 4:20 | 1/1 · 10:37 | 0/1 · 15:31 | 1/1 · 2:09 | 1/1 · 4:12 |
| 021-discourse-shared-uploads | v1 | 1/1 · 2:42 | 1/1 · 11:19 | 1/1 · 2:25 | 1/1 · 4:13 | 1/1 · 3:56 |
| 022-discourse-multisite-migration | v1 | 1/1 · 1:19 | 1/1 · 3:46 | 1/1 · 1:08 | 1/1 · 1:51 | 0/1 · 1:39 |
| 023-auth-secret-rollout | v1 | 1/1 · 2:49 | 1/1 · 3:17 | 1/1 · 1:31 | 1/1 · 2:03 | 1/1 · 4:39 |
| 024-discourse-interrupted-deploy | v1 | 1/1 · 1:57 | 0/1 · 5:04 | 1/1 · 1:56 | 0/1 · 5:19 | 1/1 · 3:15 |
agent_runtime_error: 3agent_timeout: 1release_not_converged: 2The publisher recorded harness provenance and validated matching scenario pack revisions, scenario versions, attempts, timeout, agent adapter, and Claux release before combining these summaries.
Scenario packs: ducks/replaybook-infra@20260814.0.0.
| Matrix | Models | Replaybook commit |
|---|---|---|
host-matrix-2026-08-15__21-56-08.20e83f | DeepSeek V4 Flash 0731, Gemini 3.7 Flash, GPT-5.6 Luna, GLM 5.2 · reasoning high | 7f4b117a |
host-matrix-2026-08-15__21-07-50.e717e8 | Claude Sonnet 5 · reasoning high | 7f4b117a |
host-matrix-2026-08-15__21-28-32.898d2a | Claude Sonnet 5 · reasoning high | 7f4b117a |
host-matrix-2026-08-15__20-18-41.a9a4e6 | Claude Sonnet 5 · reasoning high | 7f4b117a |
python integrations/host/run_host_matrix.py \
--scenario 001-nginx-502-host \
--scenario 013-sidekiq-wrong-redis \
--scenario 014-missing-rails-migration \
--scenario 015-sidekiq-poison-pill \
--scenario 016-rails-pool-exhaustion \
--scenario 017-partial-rails-rollout \
--scenario 018-node-event-loop-blocking \
--scenario 019-rust-fd-leak \
--scenario 020-python-gunicorn-saturation \
--scenario 021-discourse-shared-uploads \
--scenario 022-discourse-multisite-migration \
--scenario 023-auth-secret-rollout \
--scenario 024-discourse-interrupted-deploy \
--models \
z-ai/glm-5.2 \
deepseek/deepseek-v4-flash-0731 \
openai/gpt-5.6-luna \
anthropic/claude-sonnet-5 \
google/gemini-3.7-flash \
--reasoning-efforts high \
--attempts 1 \
--concurrency 2
Host harness v19, Claux v20260815.0.0, 900-second agent timeout. Usage was reported for 64 of 65 trials.
Read the methodology, browse the versioned history, or inspect the complete benchmark record.