A conversational model gains a live face and background tools
Google introduced Gemini 3.8 Live with Live Avatar on September 24. It processes visual and audio input together and generates streaming video with expressions and lip movements synchronized to spoken responses. Asynchronous tool calls can run in the background without stopping the conversation, targeting enterprise flows such as guest check-in and guided support.
Google says the model can switch across 97 languages in one conversation while adapting lip sync and expression. Preset avatars are available, while custom brand avatars can be created from a reference image only for allowlisted enterprise customers. Generated audio and video include SynthID watermarking.
A polished demo is not an operating SLA
The announcement does not publish latency distributions, concurrency cost, per-language lip-sync accuracy or failure rates. Support for 97 languages does not guarantee equal quality across dialects and accents. SynthID helps identify generated media; it does not prove rights to an avatar or factual accuracy of its speech.
Enterprises need consent for faces and voices, human handoff, approval gates for consequential tool actions, recording notices and retention controls. Latency, recognition and tool failures should be evaluated separately for every supported market.