Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier +2 more
When a multimodal model writes captions for images, some of its mistakes are not random. It might reliably get something wrong whenever a particular visual feature shows up, say describing the wrong object every time a certain background appears. Symbal is built to catch that kind of patterned error, which the authors call a systematic misalignment.
It runs in two stages using off-the-shelf foundation models, and needs no access to the model that wrote the captions, so you can audit a dataset from the outside. It even writes up what it found in plain language. To measure progress, the team assembled SymbalBench: 1.7 million image-caption pairs across natural and medical images, sorted into 420 datasets with known planted errors. On it, Symbal flagged the systematic error in about 64% of datasets, close to four times the nearest baseline.
That summary rests on the abstract, so check the paper for how the benchmark was built and where the method stumbles.
Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such…
Data Driven Block Replacement Scheduling
arXiv (cs.CV) · July 16, 2026Divergent Gaze Patterns in Artistic Viewing: Spatial and Temporal Signatures of Attention Across Autistic Individuals, Artists, and Neurotypical Observers
arXiv (cs.CV) · July 16, 2026Structural-Semantic Reciprocal Learning for Unsupervised Visible-Infrared Person Re-Identification
arXiv (cs.AI) · July 16, 2026When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space