Lyria 3.5 Brings 3-Minute AI Song Generation to Flow Music
Google DeepMind's Lyria 3.5 model introduces variable three-minute track generation, granular creative controls, and multi-language vocals to Flow Music.
Google DeepMind has launched its Lyria 3.5 generative audio model within the Flow Music workspace. The latent diffusion architecture introduces support for variable song lengths up to three minutes and brings granular tempo and duration controls to the interface. For developers handling audio synthesis, the release establishes a new baseline for structural coherence and multi-language vocal generation in production environments.
Generation Capabilities
Lyria 3.5 utilizes a time-audio latent space to construct complete tracks with defined musical structures. The model now adheres strictly to standard compositional formats, generating distinct intro, verse, chorus, bridge, and outro sections while strictly following prompt constraints.
Vocal rendering features precise pronunciation and natural transitions across languages. The model processes mixed-language lyrics, such as combined English-Chinese or Japanese phrasing, without breaking melodic continuity or audio fidelity.
Platform Integration
Flow Music, acquired by Google in early 2026 under the name Producer AI, serves as a comprehensive vibe coding workspace powered entirely by Lyria 3.5. A July 24 update introduced Spaces, a collaborative environment for building custom instruments and music applications.
The workspace integrates direct stem splitting to isolate vocals from instrumentals and supports real-time remixing. Users can control rhythm patterns manually using Drum Machine Pro. The platform also connects with Google’s Veo video model to synthesize matching visuals for completed audio tracks.
API Access and Pricing
Beyond the consumer workspace, developers can route workloads to Lyria models through the Gemini API and Vertex AI. Consumer access is handled via Google AI Subscriptions, featuring the AI Pro plan at $19.99 per month for 1,000 credits and AI Ultra at $200 per month for 25,000 credits.
| Model | Output Duration | API Cost (Approximate) |
|---|---|---|
| Lyria 3 | 30-second clips | $0.04 per clip |
| Lyria 3 Pro / 3.5 | 3-minute songs | $0.08 per song |
Audio Provenance
Every track generated by Lyria 3.5 contains a SynthID watermark. The digital signal is embedded directly into the audio waveform at an inaudible frequency. The provenance marker remains detectable by scanning tools even after users apply compression, adjust playback speed, or introduce background noise.
If you integrate audio generation into production applications, the three-minute threshold and strict structural adherence alter the baseline for standalone music services. Evaluate your compliance posture regarding the ongoing March 2026 copyright litigation before deploying Lyria 3.5 endpoints in commercial environments.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How to Run TPU Workloads on Google Cloud with Ray 2.55
Learn how to provision Google Cloud TPUs, handle slice topologies, and deploy machine learning models using Ray 2.55 and the KubeRay Operator.
Fish Audio Ships Dual-AR Speech Model Alongside $52M Seed
Fish Audio has raised $52 million in seed funding and released Fish Speech S2.1 Pro, a 4.4-billion-parameter Dual-AR voice model supporting 83 languages.
DeepMind Adapts SynthID for DNA in Bioresilience Framework
Google DeepMind and Isomorphic Labs have deployed a three-pillar bioresilience program featuring biological watermarking and AI-driven pathogen surveillance.
PixVerse R1 World Model Powers Game Engine Following $439M Round
PixVerse secured a $439 million Series C extension to scale its real-time generative game engine, pushing the video AI startup's valuation over $2 billion.
Meta's Muse Image Transformer Sparks 15B-Image Opt-Out Backlash
Meta deployed its Masked Generative Transformer, Muse Image, across its social platforms while facing backlash over its 15-billion-image training set.