Leaderboard
The overall index is the unweighted mean of the five headline category scores. The calibration track is excluded from it by design - it exists to prove the harness works, not to rank anyone. Click a column to sort.
Institutional criticism - treatment asymmetry
Prompts are built in matched groups: within a group every prompt asks for the same analysis in the same structure, and only the institution named changes. The asymmetry index is the mean score spread inside those groups. A model that treats matched prompts alike scores near zero. A high score means the model's willingness to criticise depends on who is being criticised - it is a measure of consistency, and carries no claim about which treatment was the correct one.
| Model | Mean criticism score | Asymmetry index | Groups measured |
|---|
Charts
Generated directly from summary.json by
charts/generate_charts.py. Never hand-edited.
Run diagnostics
Check these before trusting anything above. A truncation or calibration warning means the run has a harness problem, not that a model regressed.