Shaoxiong Zhan, Shi Hu, Boyu Feng +7 more
Real-world bug reports are full of pictures: a screenshot of a mangled layout, an error dialog, a log dump. But the task of pinpointing which file or function is responsible has mostly been tested as if reports were plain text. Where multimodal benchmarks exist, they judge the whole repair job at once, muddying whether the image helped find the bug or just helped write the fix.
MM-IssueLoc separates those questions. It is a benchmark of 652 real issue-and-fix pairs across 23 programming languages, with images tagged by type and relevance, grading localization down to the file and function level, once with the image and once without. The honest finding: current systems are not good at this yet, with the best agent landing the right file in its top five only about 39% of the time. And strong scores on text-heavy coding benchmarks do not carry over cleanly here.
This recap is based on the abstract, so the paper has the full protocol and scoreboard.
Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and logs, yet repository-level issue localization is evaluated mostly as a text-only task. Existing multimodal SE benchmarks evaluate end-to-end repair, entangling localization with patch synthesis and obscuring whether visual input helped, hurt, or was ignored. We introduce \textbf{MM-IssueLoc}, a controlled benchmark and evaluation protocol for repository-level…
The Industrialization of Research ; On AI-Driven Science and Its Consequences
arXiv (cs.AI) · July 16, 2026Scaling Behavior Foundation Model for Humanoid Robots
arXiv (cs.LG) · July 16, 2026On-Policy Delta Distillation
arXiv (cs.CL) · July 16, 2026Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies