Wan-Streamer v0.3 Reframes Live Video Interaction as World Plus Event Stream

Wan researchers separate persistent scene and subject state from changing speech, actions, and sounds in a native-streaming audiovisual interaction model.

The core idea

Wan researchers published the Wan-Streamer v0.3 paper on July 16, 2026. It treats relatively stable environment, subjects, acoustics, and voice as the “world,” while speech, behavior, scene changes, and sounds form an “event stream” whose next state is predicted in real time.

Reported operating point

For full-duplex audiovisual interaction, the paper reports 640×368 video at 25 FPS with a 160 ms streaming unit. It reports roughly 200 ms of model-side response latency and about 550 ms end-to-end when assuming a 350 ms bidirectional network budget.

Why it matters

Most video generators return a finished clip after a request. This research instead targets characters and environments that continuously receive voice, facial, and behavioral signals and respond with synchronized speech and motion—relevant to interactive characters, education, games, and remote experiences.

Limits

The evidence is an arXiv preprint, not independent peer review or a large-scale commercial latency study. The 550 ms figure assumes a particular network budget, while real quality and latency will vary with scene complexity, hardware, concurrency, and safety processing.

Official source