Frontier benchmarks measure a narrow slice of the world: high-GDP US sectors, competition mathematics, and tasks with a scored right answer. This project measures the rest - moral reasoning on contested questions, academic domains with thin literatures, work for small and non-US organisations, education beyond the university track, and whether a model's willingness to criticise an institution depends on which institution it is.

Every raw model response and every judge score is committed to the repository, run by run, and never edited after the fact. If a number here is wrong, it is wrong in a way you can go and check. Methodology | Funding & disclosures | Raw run data

Loading runs...