Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.CL) · July 16, 2026

Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

Haocheng Yang, Licheng Pan, Xiaoxi Li +5 more

Rubrics, checklists of what a good answer should contain, are handy for training and grading language models, but writing a reliable one for a specific question is hard. You can lean on human-written rubrics or preference data, or you can have the model generate a rubric straight from the query, which is cheap but risky: nothing checks whether that rubric actually separates good answers from bad, rewards mere style, or unfairly punishes a valid alternative approach.

Rubrics on Trial starts from an empty set and grows one, generating synthetic pairs of responses conditioned on each candidate rubric and keeping only the rubrics that genuinely distinguish quality. No human labels and no model training are needed. Across five preference benchmark suites it reports the best average accuracy.

That summary is from the abstract, so read the paper for the validation criteria and benchmark breakdown.

From the arXiv (cs.CL) abstract

Rubrics provide structured, fine-grained signals for training and evaluating large language models (LLMs). Yet reliable query-specific rubrics are difficult to construct. Existing approaches often derive supervision from human-written rubrics, preference data, or sampled responses. Direct query-to-rubric generation avoids these resources, but provides no explicit check that a plausible rubric is useful. Such a rubric may fail to distinguish answer quality, reward an optional…


More Artificial Intelligence papers