Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.AI) · July 17, 2026

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru +3 more

An evaluation that only produces a score has told you where you stand and nothing about what to do next. Most pipelines get as far as identifying weak examples, topics or categories, which sounds diagnostic but is not: it tells you where the model failed while leaving why implicit.

CRAFT takes a rubric-based evaluation set and converts it into a model-specific diagnosis of weak capabilities, treating each grading criterion as a capability in its own right and clustering from there. The shift is from a map of failures to a statement about which underlying ability is missing.

The practical payoff is what follows. Once a weakness is named as a capability rather than a list of bad examples, you can generate post-training data aimed at that capability. It closes the loop between evaluating and improving, which most evaluation work leaves as an exercise for the reader.

From the arXiv (cs.AI) abstract

Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or categories, but they leave the underlying capability failure implicit: they say where a model fails, not why. We introduce CRAFT, a method that converts any rubric based evaluation dataset into a model specific diagnosis of weak…


More Artificial Intelligence papers