Google has made its music generation model Lyria 3.5 available in the Gemini app and the Gemini API. The model, which first appeared in Google Flow Music in late July, now reaches Gemini on the web and on mobile, along with Google AI Studio and Google Vids. What sets it apart from earlier music AI is that it does not simply loop a short phrase. It writes a song several minutes long with an intro, choruses and a bridge in a single pass.

In the Gemini App, You Start by Picking a Genre

Inside the Gemini app, the flow begins with choosing a genre from a list or describing one in your own words. From there you decide whether you want vocals or an instrumental, and whether you want a short clip or a longer track. Templates have been added for anyone unsure where to start.

The obvious use cases are background music for video, a jingle for a shop or a personal brand, and a custom ringtone. The point of pushing this kind of model down into the app layer is that you can shape a result by listing conditions in plain language, without knowing anything about composition or arrangement. The rollout covers both web and mobile and is aimed at Gemini users broadly.

Two Models Split the Work: 30-Second Clips and Multi-Minute Songs

On the developer side, the Gemini API splits the job across 2 models.

The first is Lyria 3 Clip (model ID lyria-3-clip-preview), which always returns a 30-second clip. It is the lighter option, suited to iterating on prompts, building loop material, or generating previews.

The second is Lyria 3.5 itself (model ID lyria-3.5), which understands the distinction between verse, chorus and bridge and assembles a full song several minutes long. You control the length either by stating it in the prompt, as in create a 2-minute song, or by defining the structure with the timestamps described below.

Both models output 44.1 kHz stereo audio, with MP3 as the default. For Lyria 3.5 you can switch the response format and receive WAV instead. Calls go through the new Interactions API, and the response carries the generated lyrics and song structure as text alongside the audio.

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="lyria-3.5",
    input="A beautiful piano melody.",
)

You Can Write the Lyrics and the Timeline Yourself

What makes Lyria 3.5 interesting is that it leaves room for precise control rather than forcing you to hand everything over.

When you supply your own lyrics in the prompt, you separate them with section tags such as [Verse], [Chorus] and [Bridge]. If you want to govern how the song progresses, you can specify timing directly.

[0:00 - 0:10] Intro: Begin with a soft lo-fi beat and muffled vinyl crackle.
[0:10 - 0:30] Verse 1: Add a warm Fender Rhodes piano melody and gentle vocals.
[0:30 - 0:50] Chorus: Full band with upbeat drums and soaring synth leads.
[0:50 - 1:00] Outro: Fade out with the piano melody alone.

Input is not limited to text. Lyria 3.5 accepts multimodal input and lets you pass up to 10 images alongside your prompt, so you can build a track from the mood and colors of a photograph.

The language of the lyrics follows the language of the prompt. Write the instruction in French and you get French lyrics, with pronunciation and phrasing adapted to match. The same logic applies to Japanese.

The model also reasons through the song structure internally before any audio is produced. Those intermediate thoughts, however, are not exposed to the user.

Pricing Is 0.08 USD per Song, With No Free Tier

Pricing through the Gemini API is straightforward: a full song from Lyria 3.5 costs 0.08 USD (about 12 yen) per request. A 30-second clip from Lyria 3 Clip costs 0.04 USD (about 6 yen). Neither model has a free tier, so both are paid only.

※1 USD = 156 JPY

Several limitations are stated explicitly. Every prompt passes through safety filters, and requests for a specific artist's voice or for copyrighted lyrics are blocked. All generated audio carries a SynthID watermark, which is inaudible to the human ear and does not affect playback.

Music generation is also a single-turn process. Refining a finished track through back-and-forth editing is not supported in the current version. Results vary between calls even with an identical prompt, and Google documents that as expected behavior. It is better treated as a tool you run several times with tightened conditions than one you expect to hit exactly once.

Summary

With Lyria 3.5 landing in the Gemini app and the API, the standard output of music generation AI moves from 30-second loop material to structured songs several minutes long. The app side asks only for a genre and a vocal preference, while the API side lets you tighten things with section tags and timestamps. Whether it holds up in real work depends on hands-on use, once you factor in the 0.08 USD per song price, the SynthID watermark and the lack of multi-turn editing.