Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.AI) · July 16, 2026

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier +2 more

When a multimodal model writes captions for images, some of its mistakes are not random. It might reliably get something wrong whenever a particular visual feature shows up, say describing the wrong object every time a certain background appears. Symbal is built to catch that kind of patterned error, which the authors call a systematic misalignment.

It runs in two stages using off-the-shelf foundation models, and needs no access to the model that wrote the captions, so you can audit a dataset from the outside. It even writes up what it found in plain language. To measure progress, the team assembled SymbalBench: 1.7 million image-caption pairs across natural and medical images, sorted into 420 datasets with known planted errors. On it, Symbal flagged the systematic error in about 64% of datasets, close to four times the nearest baseline.

That summary rests on the abstract, so check the paper for how the benchmark was built and where the method stumbles.

From the arXiv (cs.AI) abstract

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such…


More Artificial Intelligence papers