Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +9 more
The framing here is tidy: a video is a world plus an event stream. The world is the slow stuff, the room, the people, the ambient sound, a voice's timbre. The event stream is everything moving through it, speech, actions, scene changes. Split it that way and you get a clean pretraining goal: given the world and whatever just arrived, predict how it responds next, in real time.
Wan-Streamer v0.3 puts this to work on full-duplex audio-visual interaction, where the agent's own speech and behavior are the event stream, so the model maps live input to spoken and physical actions, much like a vision-language-action system. It holds the previous version's operating point: 640 by 368 video at 25 fps, roughly 200 ms of model-side latency and about 550 ms end to end under a modest network budget.
This is a system-and-framing report, so the paper is where the training details and honest limits should live.
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject…
Linear representations of grammaticality in neural language models
arXiv (cs.AI) · July 16, 2026MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection
arXiv (cs.AI) · July 16, 2026The Industrialization of Research ; On AI-Driven Science and Its Consequences
arXiv (cs.AI) · July 16, 2026Scaling Behavior Foundation Model for Humanoid Robots