$ replaybook

Models / execution configurations

Model evidence

Find a model, then inspect its provider and reasoning configurations. Each card keeps the exact model identity and evidence boundaries visible.

This is evidence coverage, not one pooled leaderboard. Each profile uses the newest available result for that model, provider, and scenario. Totals still describe independently versioned cohorts; open a profile to inspect every boundary.

Profiles consume the stable benchmark-coverage.json API.