No configurations match these filters. Try Reset.
Real systems. Observable repairs.
Where does the agent break?
Inspect the attempts behind the score. Keep the model, its execution settings,
and the benchmark conditions visible as separate things.
A repair counts only when it survives
verification → restart → reboot.
verification → restart → reboot.
Select model rows to compare. Selected view ignores all filters except scenario.
| Configuration | Repairs | 95% interval | Median time | Reported tokens | Known spend | Cost / repair | Cost coverage |
|---|