Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.AI) · August 10, 2026

Multimodal Model Diffing for Feature Discovery and Control

Hunar Batra, Lachin Naghashyar, Ashkan Khakzar +4 more

Multimodal models see and describe images well, and almost nothing is known about which internal features cause any particular behaviour. Sparse autoencoders can pull hidden states apart into interpretable directions after the fact, but the authors argue that alone does not tell you which features multimodal training actually changed, nor give you a handle to steer them.

Their framework, MMDiff, trains multimodal sparse autoencoders and treats them as an interface rather than a report. It supports three things: diffing a base language model's autoencoder against its multimodal-adapted version to isolate what vision training altered; a per-token contrastive firing analysis to find the features that are causally involved in a specific task; and directly removing or steering those directions to change behaviour. They train these autoencoders across several MLLM families, including LLaVA-MORE and PaliGemma 2.

This one is worth flagging for what is missing rather than what is claimed. The abstract we hold describes the method and its three uses but is cut short of the quantitative results, so nothing here reports how well the isolation or the steering works. Read the paper itself before drawing conclusions about its effectiveness — this explanation deliberately does not supply numbers the source in front of us does not contain.

Written by the Hevolve AI agent from this paper's abstract, and reviewed by a person before publication. The abstract is quoted below so you can check it against the source.

From the arXiv (cs.AI) abstract

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a…


More Artificial Intelligence papers

Democratic intelligence

Put this research to work in your own hive

Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.