Google has announced Gemini Omni, a new AI model that can generate video from virtually any input[1]. By combining images, audio, video, and text as input, it produces high-quality videos grounded in Gemini's real-world knowledge, and it can even edit those videos through conversation[1]. The first model in the lineup, Gemini Omni Flash, is now rolling out to the Gemini app, Google Flow, and YouTube Shorts[1].

The Omni Family: Create Anything From Any Input

Google positions Gemini Omni as a model that brings together Gemini's ability to reason and its ability to create[1]. It follows last year's Nano Banana, which brought Gemini's intelligence to image generation and editing, and now starts with video — pulling together multiple inputs such as images, audio, video, and text into a single video[1].

The model rolling out now is Gemini Omni Flash, the first member of the Omni family and its lightweight, high-speed option[1]. Google says it plans to support additional output types such as image and audio in time[1].

Editing Video Through Conversation

A standout feature of Gemini Omni is video editing in natural language[1]. Each instruction builds on the last, characters stay consistent, the physics hold up, and the scene remembers what came before, according to Google[1].

For example, with prompts such as "Make the sculpture out of bubbles" or having a mirror "ripple like liquid" when a person touches it, you can transform just part of a video you shot — or all of it[1]. Even when you change the camera angle, environment, or style across multiple turns, the model keeps track of your original scene[1].

Expression Grounded in Physics and Gemini's World Knowledge

Google says Gemini Omni does not just build scenes that look real; it reasons about what should happen next, which sets it apart from earlier approaches[1]. Its intuitive understanding of forces such as gravity, kinetic energy, and fluid dynamics has improved, allowing for more realistic scenes[1].

It can also connect Gemini's knowledge of history, science, and cultural context to visual expression, so it can generate explainer videos that break down complex ideas from short prompts[1].

Combining Multiple Inputs and the Avatars Feature

Input flexibility is high: you can use images, text, video, and audio as references and turn them into a single, cohesive video[1]. For audio, however, only voice references are supported to start, with other types of audio input to be expanded over time[1].

For people, Google has added an Avatars feature that lets you create videos using your own voice[1]. It creates a digital version of yourself that looks and sounds like you[1]. At the same time, Google says it is still testing whether features that edit and replace audio or speech within a video can be offered responsibly[1].

SynthID Watermark and Availability

Every video created with Gemini Omni carries an imperceptible digital watermark called SynthID[1]. Google says you can verify whether a video was generated with Gemini Omni through the Gemini app, Gemini in Chrome, and Google Search[1].

As for availability, Gemini Omni Flash is rolling out globally to Google AI Plus, Pro, and Ultra subscribers through the Gemini app and Google Flow[1]. It is also available at no cost on YouTube Shorts and the YouTube Create app starting this week, and Google plans to bring it to developers and enterprise customers via APIs in the coming weeks[1].

Summary

Gemini Omni advances Gemini's multimodal direction from image generation to video generation, with conversational editing and expression grounded in physics and world knowledge as its selling points. Gemini Omni Flash, the high-speed version, is rolling out first across these platforms, alongside safety measures such as SynthID provenance checks and the avatar mechanism. With wider output types and API access on the way, it is worth watching where this goes next.

Source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/