Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.CV) · July 17, 2026

Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs

Maeve Hutchinson, Abderrahmane Wassim Mehdaoui, Pranava Madhyastha

Vision-language models are increasingly asked to read charts and draw conclusions, which makes it worth knowing how they arrive at an answer. A model that reads the correct bar and one that guesses from the caption produce identical output and deserve very different amounts of trust.

This paper offers a lightweight diagnostic saliency method built for the actual setting, text generation over images with transformers, rather than adapted from classification where most interpretability tooling originated. That distinction matters: a chart answer is generated token by token, and which part of the image mattered can change from one token to the next.

The word literacy in the title is apt. The question is not only whether the model gets chart questions right, but whether it is reading the visualisation at all, and those come apart precisely on the questions where being right matters most.

From the arXiv (cs.CV) abstract

Understanding how vision-language models (VLMs) interpret data visualizations remains an open problem, and is increasingly important as these models are used for analytical tasks where reliable reasoning is essential. We introduce a lightweight, diagnostic saliency map method tailored for text generation over images using transformer models, the current state-of-the-art models in visualization interpretation. Our approach aggregates the language model's attention over the…


More Artificial Intelligence papers