Google’s Gemini Omni Just Changed AI Video

Google’s Gemini Omni Just Changed AI Video

More

Summary

Youri van Hofwegen walks through Google’s Gemini Omni Flash model as accessed via the Open Art platform, focusing on its ability to generate video with fully synchronized, scene-appropriate audio in a single pass — no separate sound design step required. Rather than pulling from a sound library, Gemini Omni builds ambient audio, environmental effects, and voice simultaneously with the visual content from the same text description, demonstrated across a rain-streaked café, a campfire with crickets, an empty concrete skatepark (where the model correctly generates echo), and a fictional robot power-on sound that has no real-world recording to sample from.

The video covers lip-synced speech generation in detail: by embedding exact dialogue in quotation marks directly inside the scene prompt, users can produce on-screen characters whose mouth movements match the words, with delivery tone influenced by the surrounding scene description — calm delivery in a quiet room, projection over a noisy market. Van Hofwegen notes an early mistake of writing overly long lines that felt rushed in a 10-second clip.

The second half covers Omni’s video editing mode, where a completed clip can be modified using a two-part prompt structure: the first half explicitly locks what stays the same (characters, lighting, camera angle), the second half names the single element to change. This approach preserves visual continuity while making targeted changes like swapping wall color, altering weather, or replacing props — a significant improvement over the standard approach of regenerating from a new prompt and losing everything that worked in the original.


📺 Source: Youri van Hofwegen · Published August 18, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels

1 Item

Companies