Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.AI) · July 16, 2026

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

Patrick Phuoc Do, Chau M. Ta, Chaoli Wang

Testing whether AI can read charts usually stops at bar charts and line plots. This work looks at something harder: scientific visualizations, the sort of flow fields, volume renders, and technical illustrations you see in a physics or biology paper. The team ran six multimodal models through a standardized literacy test of 49 questions and compared them against 485 human participants.

The picture is uneven. Gemini came out strongest and actually beat the human average on the parts tested, while the open-source models stayed below the human baseline. All the models did fine on illustrations, search, and spatial questions, but stumbled on texture-based and integration-heavy visuals and on reading off exact quantities. The recurring failures were fine-grained estimation, judging flow direction, and tying a value back to how it was encoded.

This summarizes the abstract only, so read the paper for the per-model and per-task details.

From the arXiv (cs.AI) abstract

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We…


More Artificial Intelligence papers