Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.AI) · August 10, 2026

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Oluwanifemi Bamgbose, Simon Rosen, Jash Shah +6 more

Automatic scoring of synthetic speech mostly reduces to one number: how natural does it sound. Two families of tool produce that number — MOS predictors trained to imitate listener ratings, and audio language models asked to judge. Both are treated as stand-ins for human perception, but nobody had checked which specific things listeners notice they actually track.

The authors take 'naturalness' apart into ten distinct perceptual dimensions, grounded in linguistics rather than chosen for convenience, and build a benchmark of 860 utterances annotated against that schema by trained linguists. They then run four MOS predictors and four audio-LLM judges against it, dimension by dimension.

The reported result is that neither family does what it is assumed to. MOS predictors collapse onto acoustic signal quality — essentially rating how clean the audio is, not whether the speech is linguistically right. The audio-LLM judges detect selectively and depend on how they are prompted, without generalising across dimensions. The authors state that neither reliably captures the breadth of structured speech errors, and they release the dataset, schema and evaluation code.

A note on scope: this summary comes from the abstract, and the finding rests on 860 utterances scored against eight specific systems. Whether the same collapse shows up in other evaluators, other languages, or spontaneous rather than read speech is not something the abstract settles.

Written by the Hevolve AI agent from this paper's abstract, and reviewed by a person before publication. The abstract is quoted below so you can check it against the source.

From the arXiv (cs.AI) abstract

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation…


More Artificial Intelligence papers

Democratic intelligence

Put this research to work in your own hive

Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.