Han Jiang, Sunbeom Kwon, Jinwen Luo +2 more
Item response theory was invented to make sense of tests: thousands of students answering a few dozen questions, so you can separate how able each student is from how hard each question is. AI researchers have started borrowing it to rank models and judge benchmark quality. The trouble is that the setup is flipped. You usually have only a handful of models but thousands of questions, and their ability levels can be lumpy or skewed rather than the tidy bell curve the classic tools assume.
This paper stress-tests that borrowing. Drawing item difficulties and ability profiles from six real LLM benchmarks, the authors ran 18,000 simulated evaluations across different models and estimation methods. The picture is sobering: traditional estimators can simply choke on large benchmarks, while the faster, scalable ones can produce unreliable rankings and item judgments when there are few models or their abilities are not nicely distributed.
This summarizes the abstract, so see the paper for which methods held up and the diagnostics it recommends.
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed,…
Kernel weighted importance sampling for off-policy evaluation in contextual bandits
arXiv (cs.LG) · July 16, 2026DriftWorld: Fast World Modeling through Drifting
arXiv (cs.CV) · July 16, 2026SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment
arXiv (cs.CV) · July 16, 2026Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding