Gemini 3.8 Flash TTS

Generate studio-grade voiceovers with Gemini 3.8 Flash TTS — shape emotion line by line, cast two speakers and cover 130 languages.

Gemini 3.8 Flash TTS
Coach every line for maximum expression, or drop to Flash-Lite when cost and speed matter more
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

Two Tiers, One Voice Engine: Inside the Gemini TTS Lineup

Launched on September 23, 2026, the Gemini TTS family pairs an expressive flagship model with a lean, high-volume sibling built for bulk narration.

  • A Single Release, Two Very Different Jobs
    The flagship tier handles nuanced, character-driven reads, while its lighter sibling is tuned for fast, budget-friendly output at scale.
  • You Direct the Delivery Instead of Choosing Presets
    Style notes attached to each turn, plus structured metadata and inline vocal tags, give you control over tempo, mood, accent and emphasis.
  • Design a Voice or Replicate One With Consent
    Sketch a persona in plain words, or clone a real speaker using a clean reference clip together with a matching consent recording.

Prompting Rules That Keep the Audio Clean

Four habits that stop stray stage directions from leaking into the finished voice track.

What Gemini 3.8 Flash TTS Can Do: Feature Breakdown

From hands-on performance control to 130-language coverage, these are the strengths the flagship tier brings to production.

Hands-On Performance Control

Per-turn style cues plus inline laughs, sighs, coughs and pauses make it feel like coaching an actor rather than picking from a menu.

Voices Designed in Plain Language

Describe age, personality, accent, vocal texture and role in a prompt, with 2,000+ ready-made voices available through the Voices endpoint.

Cloning Behind a Consent Gate

Replication needs a clean reference clip and a matching consent recording from the same adult, plus SynthID watermarking and C2PA credentials.

Dialogue Scenes for Two Speakers

Write the exchange once and the model paces podcasts, teaching dialogues, product walkthroughs and game scenes — no manual line stitching.

Stable Long-Form Narration

Google documents consistent voice identity, timbre, loudness and room tone across multi-minute narration and extended conversations.

130 Languages and Regional Accents

The flagship tier reaches 130 languages while its lighter sibling stops at 101, with regional accents, minority dialects and IPA overrides.

FAQ

Gemini 3.8 Flash TTS: Frequently Asked Questions

Quick answers on cost per minute, tier selection, benchmark standing and the safety rules around voice cloning.

1

What does it cost to generate a minute of audio?

Roughly 1.35 cents per audio minute, based on launch pricing of $0.50 for input and $9 for output per million tokens.

2

Which tier should I pick, Flash or Flash-Lite?

Reach for the flagship when acting nuance and long-form audio matter; choose Flash-Lite for bulk output and low latency.

3

How does it stack up against rival voice models?

Google cites 71.4 on Hume's Voice Design Benchmark, and Voice Arena places it second with 1,260 Elo.

4

What happened to the older Gemini 3.1 Flash TTS preview?

Flash-Lite TTS supersedes that preview and drops audio output pricing from $20 to $6 per million tokens.

5

Can I clone a voice, and what guardrails apply?

Yes, with a reference clip and a matching consent recording from the same adult speaker; SynthID and C2PA metadata travel with the audio.

6

Why does the model read my stage directions aloud?

Everything in the input is spoken verbatim, so move any lasting direction into the speech metadata fields instead.

Hear Gemini 3.8 Flash TTS Perform Your Own Script

Try both tiers in Google AI Studio or the Gemini API — moving from the flagship to Flash-Lite is nothing more than swapping a model identifier. Weigh batch against priority inference before you lock in a production budget.