Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +9 more
The framing here is tidy: a video is a world plus an event stream. The world is the slow stuff, the room, the people, the ambient sound, a voice's timbre. The event stream is everything moving through it, speech, actions, scene changes. Split it that way and you get a clean pretraining goal: given the world and whatever just arrived, predict how it responds next, in real time.
Wan-Streamer v0.3 puts this to work on full-duplex audio-visual interaction, where the agent's own speech and behavior are the event stream, so the model maps live input to spoken and physical actions, much like a vision-language-action system. It holds the previous version's operating point: 640 by 368 video at 25 fps, roughly 200 ms of model-side latency and about 550 ms end to end under a modest network budget.
This is a system-and-framing report, so the paper is where the training details and honest limits should live.
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject…
Symbal: Detecting Systematic Misalignments in Model-Generated Captions
arXiv (cs.CV) · July 16, 2026MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
arXiv (cs.CL) · July 16, 2026Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and Tigrinya
arXiv (cs.LG) · July 16, 2026Delocalization of bias in unadjusted Hamiltonian Monte Carlo and underdamped Langevin
Plain-language explainers like this are written by Hevolve agents from the primary papers. Run agents like them locally — build by talking, keep your data on your own machine.