Navya Gupta, Bingjie Xu, Avinash Anand +2 more
Asking a model "is the red mug to the left of the laptop" is not one question but several: find the mug, confirm it is red, locate the laptop, work out the spatial relation. Vision-language models post respectable aggregate scores on this and nobody has looked closely at how the failures are distributed.
Aggregate accuracy hides exactly the thing you would want to know. A model failing 20% of the time could be slightly unreliable everywhere, or completely incapable of one specific operation such as spatial relations while being flawless at the rest. Those are different diagnoses with different fixes, and a single number cannot distinguish them.
This work examines which reasoning operations the failures attach to and what happens internally when they occur, calling the pattern vision-operation misalignment. Naming the failure mode is the prerequisite for fixing it, since you cannot repair a capability you have not separated from the average.
Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attribute verification. Despite strong aggregate performance, the mechanistic basis of VLM failures on this task remains underexplored. To address this gap, we analyze vision-operation misalignment in VLMs by examining how failures relate to specific reasoning operations and the internal computational…
MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization
arXiv (cs.AI) · July 16, 2026Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation
arXiv (cs.AI) · July 16, 2026Mask-Aware Policy Gradients for Diffusion Language Models
arXiv (cs.AI) · July 16, 2026Subjective Risk Decomposition: A New View for Uncertainty Quantification
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.