Operational evidence
Incident reports, shell output, logs, and configuration describe the failure.
Real systems. Observable repairs.
Replaybook is a benchmark for agents operating real infrastructure. It puts an agent inside a broken, disposable NixOS VM and checks the system it leaves behind—not the answer it writes.
Explore the evidence01 / The task
The agent receives an incident report and tools to inspect logs, services, configuration, and application state. It must find the fault and repair the existing host within a fixed deadline.
Incident reports, shell output, logs, and configuration describe the failure.
Visual scenarios require the agent to interpret a diagram or image as operational evidence, then repair the system.
These input modes stay separate in the evidence. Browse the scenario catalog →
02 / The scoring contract
The deployed service must do the work the scenario requires.
A repair must retain the data and queued work protected by the scenario.
A temporary process or in-memory workaround is not enough.
The repaired system must return to working order after reboot.
The host controller owns the verifier, restart, reboot, and deadline. The agent harness drives the model inside the guest. The scoring design keeps the verifier and expected repair outside the agent-visible environment.
03 / Reading the evidence
A pass meets the scenario’s verification checks. An evaluated failure does not. An unavailable trial records a provider or harness problem and is excluded from the published repair-rate denominator; it remains visible beside the results.
Each matrix cell shows the individual attempts. For three fully evaluated attempts, red means zero or one repaired, yellow means two, and green means all three. Incomplete evidence is gray; a missing cell means that scenario version was not run for that configuration.
The interval beside a repair rate is a 95% Wilson interval. Three successes out of three are an observation—not proof that the model will always succeed. Median duration covers evaluated attempts, not only successful repairs. Unreported API cost is unknown, not free.
04 / Comparison limits
A result belongs to a full execution configuration: the exact model ID, reasoning setting, provider, agent harness, and benchmark release. It is not a universal score for a model name.
Smoke runs qualify a setup cheaply. Core runs support routine comparison. Full runs broaden scenario coverage. Frontier runs exercise more demanding systems. The exact scenario set, attempt count, and deadline belong to each published release; tier names alone do not establish compatibility.
Replaybook measures these particular tasks under these particular conditions. It does not establish general coding ability, production safety, or reliability across every infrastructure incident.
Filter by model, provider, harness, and reasoning. Compare selected configurations, inspect failures, and follow the tracked source.
Open Evidence