Google Gives Gemini a Face: Live Avatar Speaks 97 Languages in Real Time

News Summary
Google has given its live conversational AI a face. On September 24, 2026, the company introduced Gemini 3.8 Live with Live Avatar, a feature that adds near real-time generated video to the speech of its native dialogue models, and Google's Cloud blog lists it as generally available on September 25 in Gemini Enterprise, with endpoints in the United States and the European Union. An agent can now listen, see and speak through a responsive on-screen persona instead of a voice alone. The Verge page that first flagged the story could not be retrieved for this report, so the details below come from Google's own announcements and independent coverage.
What Live Avatar Does
Live Avatar pairs Google's native speech-to-speech dialogue model with low-latency streaming video. The generated face performs lip-syncing and facial expressions that follow the synthesized voice. The system also handles turn-taking, so a user can interrupt and the agent recovers naturally.
The engine takes in voice and visual input at the same time. An avatar can therefore look at a user's camera feed or shared screen while it talks. Google says this lets businesses extend virtual offerings such as customer service and interactive walkthroughs.
Languages and Tool Use
Google says the feature supports native multilingual speech-to-speech synchronization across 97 languages, with automatic language detection. When a conversation moves from one language to another, lip-sync and expressions adapt dynamically. One reviewer notes that the multilingual claims have not been independently verified.
Asynchronous tool calling lets the agent call tools and fetch data in the background while the conversation continues. Complex tasks, such as searching inventory or looking up account details, need not interrupt the dialogue.
Preset and Custom Avatars
Organizations can pick from a library of preset avatars. Custom avatars are limited to selected enterprise customers through an allowlist. According to coverage of the terms, custom avatars require verified consent for the reference image and prohibit images of minors, celebrities or offensive content. Businesses must also secure the rights they need. An Extended Thinking variant of the model remains in private preview.
Early Customers
Google's announcement names several early adopters. Cox Automotive, the company behind Autotrader, is building an AI shopping assistant that highlights items on screen for vehicle search and financing. Equal AI runs a personal assistant that handles more than one million calls a day across nine Indian languages. Salesforce is integrating the technology with Agentforce, and Specs reports better voice activity detection and lower latency.
Pricing and Practical Limits
There is no consumer subscription. Pricing is per token for enterprise API users. According to an independent analysis, text input costs $0.75 per million tokens, audio input $3 and image or video input $1. Text output costs $4.50, audio output $12 and avatar video output $1 per million tokens. Avatar video runs at 24 frames per second and converts to about 6,192 tokens per second of speaking, which the analysis puts at roughly $0.37 per minute of avatar speech, plus about $0.018 per minute for audio output.
The same analysis says continuous sessions currently last only a few minutes. That suits bounded tasks such as hotel check-ins or short support conversations, but not long tutoring or consulting sessions. Session context is reprocessed and billed on each turn.
Safety and Transparency
All audio and video generated by Google's AI products carries an imperceptible SynthID watermark, which helps identify AI-generated content. The watermark is invisible, however, so it does not by itself guarantee that users are told they are talking to an AI. Enterprises deploying avatars will need their own clear disclosure. Engadget's headline called the avatars "creepy", a reminder that reactions to lifelike synthetic faces are mixed.
Why It Matters for Learners
Faces make voice assistants feel more like a conversation, which could help in language practice, guided software walkthroughs and interactive training. The technology also shows how quickly speech, vision and video generation are being combined in a single real-time model. Provisioned throughput is available for high-volume deployments, and Google Cloud sales handles allowlisting for custom avatars. Developers can try the feature in the Gemini Enterprise console and read the Gemini Live API documentation.
Sources
Reporting draws on Google's Cloud and product blogs, Unite.AI, Engadget, Fone Arena, Android Headlines, TestingCatalog, TechEBlog and Choosely. All dates refer to Pacific Time announcements unless noted, and Google's blog dates GA as September 25, 2026.