Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru +3 more
An evaluation that only produces a score has told you where you stand and nothing about what to do next. Most pipelines get as far as identifying weak examples, topics or categories, which sounds diagnostic but is not: it tells you where the model failed while leaving why implicit.
CRAFT takes a rubric-based evaluation set and converts it into a model-specific diagnosis of weak capabilities, treating each grading criterion as a capability in its own right and clustering from there. The shift is from a map of failures to a statement about which underlying ability is missing.
The practical payoff is what follows. Once a weakness is named as a capability rather than a list of bad examples, you can generate post-training data aimed at that capability. It closes the loop between evaluating and improving, which most evaluation work leaves as an exercise for the reader.
Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or categories, but they leave the underlying capability failure implicit: they say where a model fails, not why. We introduce CRAFT, a method that converts any rubric based evaluation dataset into a model specific diagnosis of weak…
Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search
arXiv (cs.AI) · July 16, 2026AutoSynthesis: An agentic system for automated meta-analysis
arXiv (cs.CV) · July 16, 2026ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
arXiv (cs.LG) · July 16, 2026Mutable Low-Rank Sketches for Retrain-Free Recommendation