Shilin Gao, Mark J. F. Gales, Kate M. Knill
Automatic marking of spoken English increasingly feeds audio or text straight into a large model rather than extracting hand-designed features first. That removes a bottleneck and removes a safeguard at the same time, because those features encoded what assessors believed should count.
The risk is shortcuts. A powerful model can learn highly non-linear mappings, and if some incidental property of the recording correlates with the grade, it will happily use that instead of the thing being assessed. The grader looks accurate on held-out data while measuring the wrong quantity.
In language assessment that matters more than in most applications. A shortcut latching onto accent, recording equipment or first-language background produces a system that is unfair in a specific and hard-to-detect way, and the person receiving the mark has no way to see it. Controlling that reliance is a fairness requirement, not only an accuracy one.
Increasingly, speech and language processing tasks take either audio or text directly rather than extracting features from these as the input to the classifier or regressor. Often these systems make use of complex, for example transformer-based, processes that have the ability to derive highly non-linear mappings between the input and the output. Unfortunately these systems can also learn ''shortcuts'' where the classifier is overly reliant on particular aspects of the input…
Scaling Behavior Foundation Model for Humanoid Robots
arXiv (cs.LG) · July 16, 2026On-Policy Delta Distillation
arXiv (cs.CL) · July 16, 2026Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies
arXiv (cs.AI) · July 16, 2026Concept-Guided Spatial Regularization for World Models in Atari Pong
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.