Maeve Hutchinson, Abderrahmane Wassim Mehdaoui, Pranava Madhyastha
Vision-language models are increasingly asked to read charts and draw conclusions, which makes it worth knowing how they arrive at an answer. A model that reads the correct bar and one that guesses from the caption produce identical output and deserve very different amounts of trust.
This paper offers a lightweight diagnostic saliency method built for the actual setting, text generation over images with transformers, rather than adapted from classification where most interpretability tooling originated. That distinction matters: a chart answer is generated token by token, and which part of the image mattered can change from one token to the next.
The word literacy in the title is apt. The question is not only whether the model gets chart questions right, but whether it is reading the visualisation at all, and those come apart precisely on the questions where being right matters most.
Understanding how vision-language models (VLMs) interpret data visualizations remains an open problem, and is increasingly important as these models are used for analytical tasks where reliable reasoning is essential. We introduce a lightweight, diagnostic saliency map method tailored for text generation over images using transformer models, the current state-of-the-art models in visualization interpretation. Our approach aggregates the language model's attention over the…
Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation
arXiv (cs.AI) · July 16, 2026Mask-Aware Policy Gradients for Diffusion Language Models
arXiv (cs.AI) · July 16, 2026Subjective Risk Decomposition: A New View for Uncertainty Quantification
arXiv (cs.AI) · July 16, 2026Plover: Steering GUI Agents through Plan-Centric Interaction