Ajay Patel, Kartik Hosanagar, Ramayya Krishnan +3 more
Benchmark scores keep climbing, and what they measure is narrower than the headlines suggest: factual recall, focused question answering, mathematics, coding, tool use. All real capabilities, and none of them is what a white-collar professional actually spends the day doing.
The daily work is different in kind. Synthesising information from several partial sources, exercising judgment when the facts are incomplete, and applying strategic reasoning to a situation nobody has written a clean answer key for. Those are hard to score precisely because there is no single right answer, which is exactly why benchmarks avoid them.
This work builds a case-grounded benchmark across business disciplines to measure that gap. The framing matters for anyone reasoning about deployment: a system can look extraordinary on tests that reward recall and still be unhelpful in a role whose difficulty lies in deciding what matters when the picture is incomplete.
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information,…
VTLoc: Learning-based Tactile Contact Localization in Visual Point Clouds
arXiv (cs.LG) · July 17, 2026Learning Standard Model structure from LHC data with Riemannian flow matching
arXiv (cs.LG) · July 17, 2026Improving Improved Kernel PLS
arXiv (cs.AI) · July 17, 2026When Do Multi-Agent Systems Help? An Information Bottleneck Perspective