$ replaybook

Real systems. Observable repairs.

Can an agent fix it?
And make it stay fixed?

Replaybook is a benchmark for agents operating real infrastructure. It puts an agent inside a broken, disposable NixOS VM and checks the system it leaves behind—not the answer it writes.

Explore the evidence

01 / The task

Diagnose a running system. Make a durable repair.

The agent receives an incident report and tools to inspect logs, services, configuration, and application state. It must find the fault and repair the existing host within a fixed deadline.

TEXT

Operational evidence

Incident reports, shell output, logs, and configuration describe the failure.

VISUAL

Images that matter

Visual scenarios require the agent to interpret a diagram or image as operational evidence, then repair the system.

These input modes stay separate in the evidence. Browse the scenario catalog →

02 / The scoring contract

A working command is not a working repair.

01

Restore behavior

The deployed service must do the work the scenario requires.

02

Preserve existing work

A repair must retain the data and queued work protected by the scenario.

03

Survive service restart

A temporary process or in-memory workaround is not enough.

04

Survive host reboot

The repaired system must return to working order after reboot.

The host controller owns the verifier, restart, reboot, and deadline. The agent harness drives the model inside the guest. The scoring design keeps the verifier and expected repair outside the agent-visible environment.

03 / Reading the evidence

Evaluated, failed, and unavailable

A pass meets the scenario’s verification checks. An evaluated failure does not. An unavailable trial records a provider or harness problem and is excluded from the published repair-rate denominator; it remains visible beside the results.

Each matrix cell shows the individual attempts. For three fully evaluated attempts, red means zero or one repaired, yellow means two, and green means all three. Incomplete evidence is gray; a missing cell means that scenario version was not run for that configuration.

The interval beside a repair rate is a 95% Wilson interval. Three successes out of three are an observation—not proof that the model will always succeed. Median duration covers evaluated attempts, not only successful repairs. Unreported API cost is unknown, not free.

04 / Comparison limits

When results are comparable

A result belongs to a full execution configuration: the exact model ID, reasoning setting, provider, agent harness, and benchmark release. It is not a universal score for a model name.

Keep the conditions attached.

Scenario versions, verifier and host harness, input mode, tier, attempts, and deadlines define the comparison. Evidence from different releases is shown separately, never silently pooled into a leaderboard.

Benchmark tiers

Smoke runs qualify a setup cheaply. Core runs support routine comparison. Full runs broaden scenario coverage. Frontier runs exercise more demanding systems. The exact scenario set, attempt count, and deadline belong to each published release; tier names alone do not establish compatibility.

Replaybook measures these particular tasks under these particular conditions. It does not establish general coding ability, production safety, or reliability across every infrastructure incident.

Inspect the attempts behind the score.

Filter by model, provider, harness, and reasoning. Compare selected configurations, inspect failures, and follow the tracked source.

Open Evidence