Haocheng Yang, Licheng Pan, Xiaoxi Li +5 more
Rubrics, checklists of what a good answer should contain, are handy for training and grading language models, but writing a reliable one for a specific question is hard. You can lean on human-written rubrics or preference data, or you can have the model generate a rubric straight from the query, which is cheap but risky: nothing checks whether that rubric actually separates good answers from bad, rewards mere style, or unfairly punishes a valid alternative approach.
Rubrics on Trial starts from an empty set and grows one, generating synthetic pairs of responses conditioned on each candidate rubric and keeping only the rubrics that genuinely distinguish quality. No human labels and no model training are needed. Across five preference benchmark suites it reports the best average accuracy.
That summary is from the abstract, so read the paper for the validation criteria and benchmark breakdown.
Rubrics provide structured, fine-grained signals for training and evaluating large language models (LLMs). Yet reliable query-specific rubrics are difficult to construct. Existing approaches often derive supervision from human-written rubrics, preference data, or sampled responses. Direct query-to-rubric generation avoids these resources, but provides no explicit check that a plausible rubric is useful. Such a rubric may fail to distinguish answer quality, reward an optional…
AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation
arXiv (cs.CV) · July 17, 2026Beyond Unfolding: 60x Faster One-Stage Unmixing for Closely-Spaced Infrared Small Targets
arXiv (cs.AI) · July 17, 2026Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids
arXiv (cs.CL) · July 17, 2026BayesPO: Bayesian Prompt Optimization via Parallel-Tempered Gradient-Guided Discrete MCMC