Shilin Gao, Mark J. F. Gales, Kate M. Knill
Automatic marking of spoken English increasingly feeds audio or text straight into a large model rather than extracting hand-designed features first. That removes a bottleneck and removes a safeguard at the same time, because those features encoded what assessors believed should count.
The risk is shortcuts. A powerful model can learn highly non-linear mappings, and if some incidental property of the recording correlates with the grade, it will happily use that instead of the thing being assessed. The grader looks accurate on held-out data while measuring the wrong quantity.
In language assessment that matters more than in most applications. A shortcut latching onto accent, recording equipment or first-language background produces a system that is unfair in a specific and hard-to-detect way, and the person receiving the mark has no way to see it. Controlling that reliance is a fairness requirement, not only an accuracy one.
Increasingly, speech and language processing tasks take either audio or text directly rather than extracting features from these as the input to the classifier or regressor. Often these systems make use of complex, for example transformer-based, processes that have the ability to derive highly non-linear mappings between the input and the output. Unfortunately these systems can also learn ''shortcuts'' where the classifier is overly reliant on particular aspects of the input…
QuReC: All-in-One Image Restoration with Query-Specific Guidance and Local-Global Response Calibration
arXiv (cs.AI) · July 16, 2026Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
arXiv (cs.LG) · July 16, 2026AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning
arXiv (cs.CL) · July 16, 2026Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence