The core idea
Wan researchers published the Wan-Streamer v0.3 paper on July 16, 2026. It treats relatively stable environment, subjects, acoustics, and voice as the “world,” while speech, behavior, scene changes, and sounds form an “event stream” whose next state is predicted in real time.
Reported operating point
For full-duplex audiovisual interaction, the paper reports 640×368 video at 25 FPS with a 160 ms streaming unit. It reports roughly 200 ms of model-side response latency and about 550 ms end-to-end when assuming a 350 ms bidirectional network budget.
Why it matters
Most video generators return a finished clip after a request. This research instead targets characters and environments that continuously receive voice, facial, and behavioral signals and respond with synchronized speech and motion—relevant to interactive characters, education, games, and remote experiences.
Limits
The evidence is an arXiv preprint, not independent peer review or a large-scale commercial latency study. The 550 ms figure assumes a particular network budget, while real quality and latency will vary with scene complexity, hardware, concurrency, and safety processing.