Gemini 3.8 adds more control over generated voices
Google’s new speech tools could help small teams produce repeat audio. Start with one approved script and check how the voice sounds before scaling.

Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS, according to TLDR AI’s summary of Google’s announcement. Developers can describe a voice and direct its pacing, dialect and delivery line by line. For a business that regularly makes spoken content, that gives you more control over how a script is read.
TLDR says Flash-Lite is aimed at high-volume work, including dubbing and voice agents. Flash can replicate an authorised voice from a 30-second sample. The summary does not establish what either model costs, how well a particular voice will perform, or which business tools it connects to. Those are questions to answer before committing to a workflow.
Where a team of five to fifty might use it
Think about audio your team makes repeatedly: a short product explanation, an onboarding message or a spoken answer to a common enquiry. A directed synthetic voice could help keep those recordings consistent when the wording changes. That is a possible use of the controls TLDR describes, rather than a measured result from the announcement.
The line-by-line direction is useful to test because spoken meaning depends on delivery. You might want a slower explanation of a complicated step, then a warmer closing sentence. TLDR reports that the models accept direction on pacing and delivery; it does not say they will get every instruction right. Listen to the complete recording before anyone hears it outside your team.
Voice agents need particular care. TLDR identifies them as a high-volume use for Flash-Lite, but a voice that sounds natural is only one part of a useful response. If you are considering an agent for customer enquiries, decide which questions it may answer, when it should hand over to a person and who will review its replies.
Flash’s voice replication may appeal to an owner who already records business messages. TLDR describes replication of an authorised voice from a 30-second sample. Make permission explicit, keep track of where that voice is used and review each script. A familiar voice should still say only what its owner has approved.
Try one approved script before making more audio
Choose one recurring task with a short script and a clear audience. A product walkthrough or an internal training message gives you something concrete to assess. Write the words first, including any pronunciation that matters to your customers, then decide who must approve both the script and the finished recording.
- Generate a sample with Flash or Flash-Lite and note which voice directions you used. TLDR describes voice design and line-by-line control for the new models.
- Ask someone who knows the message to listen for clarity, pronunciation and tone. Check every line against the approved script.
- If you test voice replication, use an authorised sample and have the voice owner approve the result before use.
Compare the sample with the recording process you use now. Count the revisions needed and check whether the result is clear to someone hearing it for the first time. That small trial will tell you more about its fit for your business than a polished demonstration.
Google’s release gives small firms another option for voice production, as described by TLDR. The practical decision is narrower: can this tool deliver one approved message clearly, with a review process your team can repeat? Answer that before using it across more scripts or customer conversations.
Questions
Can Gemini make a voice that sounds like our owner?
TLDR says Gemini 3.8 Flash TTS can replicate an authorised voice from a 30-second sample. The summary does not describe how closely it matches every speaker or how it handles every script. Get the voice owner’s permission, test a short message and have them approve the recording before you use it.
Which Gemini speech model should we try for lots of voiceovers?
TLDR says Flash-Lite TTS targets high-volume work such as dubbing and voice agents. Flash TTS is the model TLDR identifies for replicating an authorised voice from a sample. The summary gives no prices or performance comparison, so test your own scripts before choosing a model for regular production.
Can we control the accent and pace of a Gemini voice?
TLDR reports that developers can direct dialect, pacing and delivery line by line, and create a voice from a description. That makes an Australian dialect worth testing, but the summary does not guarantee a particular accent or pronunciation. Listen to a full sample with the words your customers actually hear.
https://aismith.com.au/blog/gemini-3-8-adds-more-control-over-generated-voices