Google has released two new generative media models for developers at the same time[1]. One is the image generation model Nano Banana 2 Lite, billed as the fastest and most cost-efficient model in the Nano Banana family. The other is Gemini Omni Flash, which handles video generation and conversational editing and, after being shown at Google I/O, is now available to developers for the first time. Both can be used through Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform, making it possible to build an end-to-end pipeline from image creation to video.

Two new models, image and video, released together

The announcement centers on two models that raise the bar for speed and ease of iteration in generative media work[1]. Nano Banana 2 Lite is an image generation model built for high throughput and low cost, available from launch in Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform. It is also rolling out to Google consumer surfaces such as AI Mode in Search and the Gemini app.

The other model, Gemini Omni Flash, handles video generation and conversational editing. Previously shown at Google I/O, it is now open to developers for the first time via the Gemini API and Google AI Studio. Google emphasizes that by connecting fast image generation with video creation and editing, developers can build comprehensive, end-to-end multimedia experiences. Whether a workflow calls for generating thousands of images or editing multi-turn video sequences, the two models can be used in combination as needed.

Nano Banana 2 Lite — an image model built for speed and low cost

Nano Banana 2 Lite (model name gemini-3.1-flash-lite-image) is designed for situations where speed and cost are the top priorities[1]. Google recommends it as the replacement for developers currently using the first version of Nano Banana (gemini-2.5-flash-image), saying that simply swapping it in delivers immediate benefits across key performance dimensions.

There are two standout characteristics. The first is low latency: it returns text-to-image outputs in 4 seconds, a speed suited to quickly iterating on prototypes and rough drafts. The second is cost efficiency: it generates images at 0.034 USD per 1,000 images (about 5 yen), a price well suited to drafting in volume or keeping operational budgets in check.※1 USD = 162 JPY (as of June 30, 2026)

Despite prioritizing speed, Google says the model retains quality such as prompt adherence, character consistency, and legible in-image text rendering.

The Nano Banana family is made up of four models for different use cases. The fastest, Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image), is built for high-volume processing that demands ultra-low latency. The general-purpose workhorse, Nano Banana 2 (Gemini 3.1 Flash Image), offers the best balance of quality, speed, and cost. Nano Banana Pro (Gemini 3 Pro Image) is the top tier for complex, professional tasks where accuracy matters more than speed. And the original Nano Banana (Gemini 2.5 Flash Image) is positioned as the legacy model, with an upgrade to Nano Banana 2 Lite recommended.

In addition to developer platforms, Nano Banana 2 Lite is coming to consumer surfaces including AI Mode in Search, the Gemini app, NotebookLM, Google Photos, Stitch, Google Flow, and Google Ads.

Gemini Omni Flash — a video model you can edit by conversation

Gemini Omni Flash (model name gemini-omni-flash-preview) is a model that brings together Gemini's multimodal reasoning with video generation and editing[1]. It supports high-quality video generation and conversational editing from a combination of text, image, and video inputs. It is priced at 0.10 USD per second of video output (about 16 yen), the same level as Veo 3.1 Fast.

Omni Flash's strengths include conversational editing that lets you refine and edit video using natural language, multimodal referencing that combines images, text, and video to keep a scene consistent, the use of Gemini's real-world knowledge such as history, biology, and narrative logic to build coherent videos, and synchronization that ties text and graphics directly to actions in the video.

At the same time, there are several limitations at launch. Video generation is currently limited to 10 seconds, with longer durations to come. Uploading audio references and scene extension are not yet supported in the Gemini API. Video references up to 3 seconds are accepted by the API schema but are not correctly processed at this time. Character consistency during scene changes or panning movements also has limitations, which Google says it is working to improve. Gemini Omni is available in public preview starting today in Google AI Studio and the Gemini API.

Using the two models together

Google sees the real value in chaining the two models[1]. You can generate an image quickly with Nano Banana 2 Lite, then pass that image as a reference to Gemini Omni Flash to animate it into a video. Using the Interactions API, you can also maintain session history and context across multiple turns, letting users stack up to 3 sequential edits.

Google has also prepared demo apps that let you experience the two models together. "Anywhere" takes a selfie or uploaded photo and uses Nano Banana 2 Lite to instantly transport you to landmarks around the world; tapping an image then has Omni Flash turn it into an animated clip of that location. "Space Lift" is an interior design app that generates a range of design concepts from a photo of a room and shows a favorite as a cinematic video. "Omni product studio" is a demo that converts still images into e-commerce videos, illustrating a creative flow that merges multimodal inputs.

For safety and transparency, both Gemini Omni and Nano Banana 2 Lite use SynthID watermarking. Generated content can be verified through the Gemini app, Gemini in Chrome, and Search.

Summary

Google has released Nano Banana 2 Lite, a fast and cost-efficient image model, and Gemini Omni Flash, which handles video generation and conversational editing, for developers. The workflow assumes you create images quickly and then animate them into video, which should broaden the range of services that can be built with generative media. With clear figures for pricing and latency, the models are easier to choose between for everything from prototyping to volume production.

出典:https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/