All posts

Noël Vranckx • • 7 min read

Gemini 3.8 TTS is for the people who never see your dashboard

Much of supply chain works with busy hands, busy eyes or a different first language. Google’s new text-to-speech models make a spoken briefing in every language on the line cost cents.

Illustration of a factory shift-change corner: a saffron horn loudspeaker on a brick wall above a clocking-in time clock, a rack of blank time cards, work jackets and ear defenders on hooks, and a bench by a tall window onto the assembly hall

*We keep building better screens for supply chain. The people who move the goods are mostly the ones who can’t look at them.*

Here’s a claim some colleagues won’t like: the dashboard is the wrong format for half the people in supply chain.

Think about who does the physical work. The line operator has both hands on a frame. The forklift driver must keep their eyes on the aisle. The agency temp who started on Monday reads Polish far better than the language on the whiteboard.

We give all of them the same thing: text. A shift log, a laminated instruction, a Teams message they’ll see at the break. The information is right. The format is wrong.

That’s why I paid attention when Google released two new voice models this week. They turn text into speech that sounds like a person talking, in more than 100 languages, for a fraction of a cent per minute. They are Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.

What Gemini 3.8 TTS is

Google launched both models on 23 September 2026. TTS stands for text-to-speech: you send text, you get audio back. What’s new is how far you can direct the delivery, and how cheap the high-volume version is.

  • Flash TTS is built for quality and creative control: audiobooks, games, long narration.
  • Flash-Lite TTS is built for volume and cost: dubbing, voice agents and translation pipelines.

You don’t record anything. You write a script, pick a voice and describe how it should sound, in plain words such as “calm and clear, speaking slowly”. Both models detect the language of your text by themselves.

The key capabilities, according to Google:

  • More than 2,000 ready-made voices across 100+ languages and dialects, up from 30 voices before.
  • Two speakers in one clip, so a script can be a short conversation instead of a monologue.
  • Line-by-line direction, including pauses and small sounds such as a sigh or a “mhm”.
  • Voice design from a written description, and voice replication from a 30-second sample of a consenting speaker.
  • Long-form audio that stays consistent over hours.
  • A SynthID watermark in every clip, so the audio can be identified as generated.

Google says Flash TTS ranked first and Flash-Lite second on Hume AI’s quality index. Treat that as self-reported until you’ve heard it in your own languages.

The price is the part that matters for operations. Audio counts as 25 tokens per second. On Google’s paid tier, Flash-Lite costs $6 per million audio tokens until 31 December 2026, which is under one cent per minute. From 1 January 2027 the price doubles to $12. Flash costs 50% more than Flash-Lite. Both are also listed on Vercel’s AI Gateway.

What it means for supply chain

I use a simple test. A voice beats a screen when at least one of three gaps is open.

GapWhat it looks likeWhy voice helps
Hands busyAssembly, picking, loadingYou can listen while you work
Eyes busyDriving, forklifts, inspectionsYour eyes stay on the job
LanguageAgency staff, drivers from abroadEveryone hears the same message in their own language

Where two or three gaps are open at once, the case gets strong. A few places to look:

  • Shift handovers as a 90-second spoken briefing, in each language spoken on the line.
  • Driver briefings at your gate: site rules, dock number and unloading steps, in the driver’s own language.
  • Peak-season onboarding, turning your standard work instructions into short audio lessons for temporary staff.
  • Safety toolbox talks that sound the same in every shift, instead of depending on who happens to read them out.
  • Order status and slot booking by phone, where Flash-Lite is built to be the voice of an agent. Keep a human on the line for anything that changes a commitment.
  • An S&OP digest, where two voices talk through the key decisions of this month’s pack in four minutes for the commute.

The shift handover is the one I’d start with. The text already exists, it’s short and it’s written every eight hours.

Worked example: a handover in every language on the line

A fictional bicycle maker runs its e-gravel assembly line in three shifts. The night supervisor writes a handover log at 05:45. Here is the process, with illustrative data.

Step 1: start from the log you already have

Nothing new to write. This is the log, as the night supervisor typed it:

Line 2, e-gravel. Night shift 22:00-06:00.
Output 212 of 240 planned. Torque station 4 down 01:40-02:35, sensor replaced.
SAFETY: oil on floor at station 6, cleaned, cones stay until 08:00 inspection.
QUALITY: 6 frames on hold, paint blisters, batch P-0917, red cage. No rework until QA decides.
MATERIAL: motor brackets MB-22 for about 5 hours. Truck due 10:30, gate 3.
PEOPLE: 2 agency starters on station 2 today, induction before 07:00.
PLAN: 250. Order 4471 for the Lyon hub first, it ships at 14:00.

Step 2: turn it into a spoken script

Text written for the eye doesn’t work for the ear. “MB-22” and “01:40-02:35” are fine on paper and confusing out loud. So a language model first rewrites the log as a script. Any good chat model will do. Paste this:

You are writing a spoken shift handover for a bicycle assembly line. It will be turned into audio and played at the start-of-shift huddle.

Task: turn the shift log below into a two-voice script of about 90 seconds (220 to 250 words).
Speakers: "Night" (the outgoing supervisor) and "Early" (the incoming supervisor, who asks short questions).

Rules:
- Order: safety first, then quality holds, then material, then people, then the plan.
- Write every number the way a person says it ("two hundred and twelve", "half past ten").
- Spell out each code the first time ("batch P, zero nine one seven").
- Short sentences. No jargon a new agency worker wouldn't know.
- Add nothing that is not in the log. If something is unclear, write [CHECK] instead of guessing.
- End with the one thing everyone must remember today.

Format: one line per turn, starting with "Night:" or "Early:".
At the end, list any assumptions you made.

Shift log:
[paste the log here]

Read the script before you do anything else. This is where a wrong number would slip in, not in the audio.

Step 3: generate the audio

Open the speech playground in Google AI Studio and choose Flash-Lite TTS. Use two speakers, give “Night” and “Early” each a voice, and paste the script. Add one line of direction, for example: “Calm, clear, factory-floor pace, a short pause between topics.”

You’ll get a 90-second clip of two supervisors talking the shift through. Listen to it once, all the way through.

Step 4: add the languages on the line

Ask the chat model to translate the script into the languages your teams speak, say Polish, Portuguese and Romanian. Keep the speaker names. Paste each version into the playground. The model hears that the text is Polish and speaks Polish, with no setting to change.

Let a native speaker from the line listen to the first versions. They’ll catch a clumsy phrase in ten seconds.

Step 5: put it where people already are

Play it at the huddle board, or put a QR code on the whiteboard that opens the clip. Someone who arrives late can still hear it.

Now the cost, with the illustrative numbers above. Ninety seconds is 2,250 audio tokens. Four versions, three shifts a day, 250 working days: about 6.75 million tokens a year. On Flash-Lite that is about $40 a year at today’s price, or $81 at next year’s.

What to notice: the model is the cheap part. The value sits in the rules of the prompt, such as safety first and [CHECK] instead of guessing. It also sits in the supervisor who listens once before playing it. For the business, it means the new starter on station 2 hears about the oil at station 6 in their own language. Before they walk past it.

What it can’t do, and the traps

  • It says what you give it, confidently. A wrong figure in the script becomes a wrong figure spoken in a calm, trustworthy voice. Keep a human listen before anything is played, and use [CHECK] markers.
  • Codes and numbers are the weak spot. Test how your part numbers, order references and times come out, and write them out in the script if needed.
  • Translation is the language model’s job, not the voice model’s. For safety-critical instructions, start from your officially translated procedures, not a fresh translation.
  • Mind the free tier and the price change. Google’s pricing page says free-tier data is used to improve its products, so use the paid tier for real shift logs. Plan for the doubled price from 1 January 2027.
  • Leave voice cloning alone. Cloning the plant manager’s voice sounds fun. It needs verified consent, and Google doesn’t offer it in the EEA, the UK or Switzerland. Tell people the voice is synthetic.

Press play at the next handover

This week, take one real handover log (anonymised), run it through steps 2 and 3, and play the clip to one supervisor. Ask them one question: would the new starter on your line understand this better than the whiteboard? It takes 20 minutes.

For years, digital supply chain meant more screens for people at desks. Voice models like these can finally bring the same information to the people doing the physical work, in the language they think in.

Which message in your operation is written down every day but should really be heard?

Sources

Pass it on

Know a colleague who should read this? Post it where they will see it, or send it to them directly.

More posts