An item-response-theory audit of the benchmark ecosystem the index leans on: Ai2 fits multidimensional IRT models to 100 language models across 16 benchmarks and about 34,000 items, recovering latent ability factors rather than raw accuracies. The analysis separates a safety factor from reasoning factors and shows which items carry information at the frontier. Released 1 September 2026 as a blog post with code; no arXiv identifier as of mid-September.

Paper

evaluationbenchmarkresearch