Research index
Model releases/

Gemini 3.8 Flash TTS API: Pricing and Setup

Gemini 3.8 Flash TTS is $0.50 text in and $9 audio out per 1M tokens through 31 Dec 2026. Id gemini-3.8-flash-tts. Not on RouterPlex yet.

Written byRouterPlex
Reading time5 min
Gemini 3.8 Flash TTS API: Pricing and Setup
Gemini 3.8 Flash TTS API: Pricing and Setup

Gemini 3.8 Flash TTS is Google's flagship Gemini 3.8 speech model, published 23 September 2026. The API id is gemini-3.8-flash-tts. Standard paid pricing through 31 December 2026 is $0.50 per 1M text input tokens and $9.00 per 1M audio output tokens.

Gemini 3.8 Flash TTS is not on RouterPlex yet. Register for the current 60+ models. The live chat model Gemini 3.8 Flash is gemini-3.8-flash at $0.75 / $3.75 and returns text. The speech id can be used here once the marketplace lists it. Prepaid, 0% markup on the vendor list price, $0 top-up fee, hard spend limit.

Google launch card: Introducing Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS.
Google launch card: Introducing Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS.

Sources: Google's Gemini 3.8 text-to-speech post (23 September 2026), the Flash TTS model page, and the Gemini API pricing page, checked 28 September 2026. Hume scores are the figures on Google's published chart.

Gemini 3.8 Flash TTS API pricing #

Paid rates, USD per 1M tokens. Audio tokens are 25 tokens per second. Google prints the 10-second equivalent next to the output price: $9.00 × 250 / 1,000,000 = $0.00225.

TierThrough 31 Dec 2026From 1 Jan 2027
Standard input, text$0.50$1.00
Standard output, audio$9.00$18.00
Standard, per 10 seconds of audio$0.00225$0.0045
Standard cache, input$0.125$0.25
Batch or Flex input$0.25$0.50
Batch or Flex output$4.50$9.00
Priority input$0.90$1.80
Priority output$16.20$32.40

Storage on the caching rows is $0.50 per 1M tokens per hour through 31 December 2026, then $1.00. Flex cache input is cheaper than batch cache: $0.025 then $0.05, versus batch cache at $0.0625 then $0.125. Priority cache input is $0.225 then $0.45. Priority's 10-second equivalent is $0.00405 through 31 December 2026, then $0.0081.

Google's standard free tier lists input and output as free of charge, and marks Used to improve our products: Yes. The paid column marks that flag No. Batch free tier is not available. That free tier is Google's. It is not a RouterPlex credit.

What the model is #

Gemini 3.8 Flash TTS
Model idgemini-3.8-flash-tts
Input / outputText in, audio out
Languages130
Input / output token limits8,192 / 16,384
Default unary audioWAV (audio/wav) with a RIFF header
Caching, batch, flex, prioritySupported
Function calling, thinking, Live API, searchNot supported
Best fit, per GoogleCreative work: audiobooks, narration, multi-speaker dialogue
On RouterPlexWhen listed

Older Gemini TTS models returned headerless PCM by default. If your code wraps raw PCM in a WAV header, Google says to stop wrapping and write the bytes, or set the response format explicitly if you still need headerless PCM.

Both 3.8 TTS models share a schema with Flash-Lite TTS. Put sustained style and the speaker name in speech_metadata. Leave only point events in the transcript. A laugh, a sigh, and a short pause are the examples Google gives. A line such as "Say cheerfully: Hello" can be spoken aloud.

The 23 September post also describes 2,000+ production voices, including Mexican Spanish, Quebec French, and Scots English, plus voice replication from a 30-second sample with consent verification, SynthID, and C2PA. Every Gemini Audio clip is watermarked with SynthID. The model card is Gemini 3.8 Audio. Voice remixing was marked coming soon on that post.

Hume Voice Design chart #

Google's chart, methodology at deepmind.google/models/evals-methodology/gemini-3-8-tts:

Gemini 3.8 Flash TTSElevenLabs Voice Design v3Inworld Voice Design
Overall, English71.470.869.8
Multilingual3.823.653.57
Accents60.845.435.8
Voice Qualities74.676.676.3

ElevenLabs leads the Voice Qualities row. Google leads the other three rows on this chart. The blog also says Flash and Flash-Lite placed first and second on Hume's Overall Quality Index. That index is not in the table above, and Google did not print its scores.

Hume AI text-to-speech Voice Design leaderboard published with Google's Gemini 3.8 TTS post. Overall English: Gemini 71.4, ElevenLabs Voice Design v3 70.8, Inworld 69.8. Voice Qualities: ElevenLabs 76.6, Inworld 76.3, Gemini 74.6.
Hume AI text-to-speech Voice Design leaderboard published with Google's Gemini 3.8 TTS post. Overall English: Gemini 71.4, ElevenLabs Voice Design v3 70.8, Inworld 69.8. Voice Qualities: ElevenLabs 76.6, Inworld 76.3, Gemini 74.6.

Character pricing for ElevenLabs is on the Eleven v4 page. Different unit, different bill.

Call the live chat model #

gemini-3.8-flash at $0.75 / $3.75 is the chat route. Gemini 3.8 Live is a different product and is already covered.

bash
curl https://api.routerplex.com/v1/chat/completions \
-H "Authorization: Bearer $ROUTERPLEX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.8-flash",
"messages": [
{"role": "user", "content": "Write a two-sentence studio prompt for a narrator voice."}
]
}'

On RouterPlex, once Flash TTS is listed #

The speech route gets a name when the row exists. It will not be silently aliased to gemini-3.8-flash. Billing will be the vendor list price, 0% markup, prepaid, hard per-key budget. Create an account for the 60+ models live now. The catalog and the OpenAI-compatible API are the contract those rows already use.

Common questions

Frequently asked questions

How much does Gemini 3.8 Flash TTS cost?

On Google's Gemini API pricing page, standard paid rates through 31 December 2026 are $0.50 per 1M text input tokens and $9.00 per 1M audio output tokens. From 1 January 2027 those rates step up to $1.00 and $18.00. Google equates the 2026 audio output to $0.00225 per 10 seconds. RouterPlex does not bill this model.

What is the Gemini 3.8 Flash TTS model ID?

gemini-3.8-flash-tts. It takes text in and returns audio. The Gemini API serving limits on the model page are 8,192 input tokens and 16,384 output tokens. There is no RouterPlex id until the catalog lists it.

Is Gemini 3.8 Flash TTS on RouterPlex?

Not yet. Register for the current 60+ models. gemini-3.8-flash is live as a chat model at $0.75/$3.75 and does not speak. Flash TTS can be used here once it is on the marketplace.

Is Google's free tier a RouterPlex credit?

No. Google's standard free tier lists input and output as free of charge, with Used to improve our products set to Yes. Paid tier sets that flag to No. RouterPlex does not offer that free tier and does not grant credits for this model.

How do I keep stage directions from being spoken?

Google says Gemini 3.8 TTS treats input text as a verbatim transcript. Move style and speaker into speech_metadata. Angle-bracket tags are only for point events such as a laugh, a sigh, or a short pause. Inline directions like Say cheerfully may be spoken aloud.

Run the smallest paid test.

Add $5, cap the key, and verify the result with your own workload. No subscription, and credit never expires — a first top-up of $25+ is matched with $25 extra.