$ replaybook

Real systems. Observable repairs.

Where does the agent break?

Inspect the attempts behind the score. Keep the model, its execution settings,
and the benchmark conditions visible as separate things.

A repair counts only when it survives
verification → restart → reboot.
Select model rows to compare. Selected view ignores all filters except scenario.
· select a cell to inspect
✓Repaired×Failed–Unavailable— Not run
Whole-cohort summary · includes all scenarios, independent of the scenario filter
ConfigurationRepairs95% intervalMedian timeReported tokensKnown spendCost / repairCost coverage
Tracked evidence only · no live inference · no inferred model aliases