Sreyan Ghosh, Arushi Goel, Kaousheik Jayakumar +9 more
Most audio-visual language models are built and tested on short clips, which quietly avoids the thing that makes real video hard. In a long recording the sound and the picture drift in and out of agreement, evidence for an answer can be minutes apart, and what you heard early on changes how you should read what you see later.
AV-Flamingo is aimed at that longer form, jointly reasoning over audio, images and long videos, and the authors make a point of it being fully open. That matters for a class of model where the strongest systems are usually closed and therefore impossible to inspect or build on.
Their first contribution is a large collection of real-world audio-visual skills data. That ordering is telling: for long-form understanding the shortage was never architectures, it was training material that actually contains the long-range dependencies a model is supposed to learn.
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of…
Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya
arXiv (cs.LG) · July 16, 2026Delocalization of bias in unadjusted Hamiltonian Monte Carlo and underdamped Langevin
arXiv (cs.LG) · July 16, 2026BadWAM: When World-Action Models Dream Right but Act Wrong
arXiv (cs.AI) · July 16, 2026MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization