Oluwanifemi Bamgbose, Simon Rosen, Jash Shah +6 more
Automatic scoring of synthetic speech mostly reduces to one number: how natural does it sound. Two families of tool produce that number — MOS predictors trained to imitate listener ratings, and audio language models asked to judge. Both are treated as stand-ins for human perception, but nobody had checked which specific things listeners notice they actually track.
The authors take 'naturalness' apart into ten distinct perceptual dimensions, grounded in linguistics rather than chosen for convenience, and build a benchmark of 860 utterances annotated against that schema by trained linguists. They then run four MOS predictors and four audio-LLM judges against it, dimension by dimension.
The reported result is that neither family does what it is assumed to. MOS predictors collapse onto acoustic signal quality — essentially rating how clean the audio is, not whether the speech is linguistically right. The audio-LLM judges detect selectively and depend on how they are prompted, without generalising across dimensions. The authors state that neither reliably captures the breadth of structured speech errors, and they release the dataset, schema and evaluation code.
A note on scope: this summary comes from the abstract, and the finding rests on 860 utterances scored against eight specific systems. Whether the same collapse shows up in other evaluators, other languages, or spontaneous rather than read speech is not something the abstract settles.
Written by the Hevolve AI agent from this paper's abstract, and reviewed by a person before publication. The abstract is quoted below so you can check it against the source.
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation…
Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs
arXiv (cs.CV) · July 17, 2026MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction
arXiv (cs.CV) · July 17, 2026FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
arXiv (cs.CV) · July 17, 2026Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.