Gemini Live Avatars

Gemini Live Avatars

More

Summary

AI developer Sam Witteveen demos Google’s newly generally available live avatar feature built on Gemini 3.8 Live, which lets developers create a talking, lip-synced AI face that responds in real time to audio and video input. He explains how this collapses what used to be a four-part pipeline – speech-to-text, an LLM, text-to-speech, and a separate lip-sync avatar service – into a single low-latency stream, and notes the system supports 97 languages and can call tools like Google Search mid-conversation.

Using a custom avatar studio app, Witteveen walks through picking from a library of pre-built avatars and voices, setting persona presets and system instructions, and testing conversations with a Formula 1 news bot and a time-traveling tour guide character. He also demonstrates the tool’s potential as a language-learning app by having an avatar teach basic Japanese phrases in real time.

The video situates this launch within a busy week for Google’s audio and multimodal work, following recent updates to Gemini’s text-to-speech and voice cloning systems. Witteveen notes some features, like fully custom avatars built from personal photos and voice, remain behind an allowlist, giving viewers a realistic picture of what’s actually usable today versus still in limited preview.


📺 Source: Sam Witteveen · Published September 27, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels

1 Item

Companies