Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.CV) · July 16, 2026

Video = World + Event Stream

Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +9 more

The framing here is tidy: a video is a world plus an event stream. The world is the slow stuff, the room, the people, the ambient sound, a voice's timbre. The event stream is everything moving through it, speech, actions, scene changes. Split it that way and you get a clean pretraining goal: given the world and whatever just arrived, predict how it responds next, in real time.

Wan-Streamer v0.3 puts this to work on full-duplex audio-visual interaction, where the agent's own speech and behavior are the event stream, so the model maps live input to spoken and physical actions, much like a vision-language-action system. It holds the previous version's operating point: 640 by 368 video at 25 fps, roughly 200 ms of model-side latency and about 550 ms end to end under a modest network budget.

This is a system-and-framing report, so the paper is where the training details and honest limits should live.

From the arXiv (cs.CV) abstract

We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject…


More Artificial Intelligence papers