Patrick Phuoc Do, Chau M. Ta, Chaoli Wang
Testing whether AI can read charts usually stops at bar charts and line plots. This work looks at something harder: scientific visualizations, the sort of flow fields, volume renders, and technical illustrations you see in a physics or biology paper. The team ran six multimodal models through a standardized literacy test of 49 questions and compared them against 485 human participants.
The picture is uneven. Gemini came out strongest and actually beat the human average on the parts tested, while the open-source models stayed below the human baseline. All the models did fine on illustrations, search, and spatial questions, but stumbled on texture-based and integration-heavy visuals and on reading off exact quantities. The recurring failures were fine-grained estimation, judging flow direction, and tying a value back to how it was encoded.
This summarizes the abstract only, so read the paper for the per-model and per-task details.
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We…
Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence
arXiv (cs.CV) · July 16, 2026Quantifying Training Membership Information in the Hyperspherical Embedding Geometry of Face Recognition Models
arXiv (cs.AI) · July 16, 2026Towards Hierarchical Structure Understanding of Newspaper Images
arXiv (cs.LG) · July 16, 2026Evaluating covariate balance for long time horizon Markov decision processes
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.