Gemini 3.8 TTS reads your instructions aloud? Move delivery into metadata

Migrate to Gemini 3.8 Flash TTS with literal transcript text, speech_metadata and the Interactions API. Check audio output without mixing older request formats.

A stencil sheet drawn in black ink on a pale card, a sage background and the title Gemini 3.8 TTS.

For Gemini 3.8 Flash TTS, put the words to be spoken in the transcript and sustained delivery instructions in speech metadata. If you paste “speak calmly” into text that the model treats as a literal transcript, that instruction can become part of the speech you asked it to produce. The Gemini 3.8 Flash TTS migration guidance explicitly separates transcript text from delivery metadata.

Google announced the new TTS models in its September 22, 2026 API release notes. This article gives a documentation-aligned migration example for the Interactions REST API. The request structure and decoding logic can be checked locally; this article does not claim a live listening test, improved voice quality or current Ofox support for this endpoint.

Separate what is said from how it is said

InformationWhere it belongs in this example
Words the listener should hearThe text content’s text field
Sustained delivery, such as calm and clearA speech_metadata annotation’s style
Selected voicegeneration_config.speech_config
Requested audio outputresponse_format

Keep that separation even when the instruction is short. It makes the request easier to inspect and avoids confusing an instruction with the script itself. A transcript that includes a character saying “speak calmly” is different: those words belong in the transcript because the listener is supposed to hear them.

The model documentation also describes point-in-time vocal events. Do not assume that every older prompt tag or every arbitrary annotation is supported. Use the current reference for the model and API family you selected.

Convert an old narration request field by field

Start with the approved spoken script, not the old prompt as a single string. Suppose the old input was “Read warmly as a station announcer: The next train leaves at noon.” The intended spoken text is only “The next train leaves at noon.” Put “warm station announcer delivery” in style; leave the transcript free of that direction. This is an illustrative edit, not an audio result.

For the first request, keep the selected voice and sentence constant. Save the old request in a separate file, then create the Interactions request shown in the next section. Do not copy contents, parts, speech_metadata placement or response parsing from GenerateContent into this request family. A field can have a familiar name while belonging at a different level of the JSON.

The current model reference distinguishes sustained style from brief events. Whispering throughout a sentence belongs in metadata; a momentary <sigh> can remain in the transcript when you actually want that event. Do not convert every instruction into an invented tag. If a word is part of the approved dialogue, keep it as text even if it happens to sound like a direction.

For two speakers, first write a turn table: turn number, exact transcript, speaker identifier, style and configured voice. Every turn must explicitly name a speaker matching the configuration. A visible prefix such as Speaker 1: in the transcript is not the same as the structured speaker field and may be spoken aloud. Establish that mapping before expanding to a long dialogue; copying a single-speaker voice array is not a complete multi-speaker setup.

Use one API family from request to response

The following example uses the Interactions API. Do not mix its input array and annotations with a GenerateContent request body or with fields from an older SDK example. The official speech-generation guide is the source for this request family.

Save this as request.json:

{
  "model": "gemini-3.8-flash-tts",
  "input": [{
    "type": "user_input",
    "content": [{
      "type": "text",
      "text": "The next train leaves at noon.",
      "annotations": [{
        "type": "speech_metadata",
        "style": "calm and clear"
      }]
    }]
  }],
  "response_format": {"type": "audio"},
  "generation_config": {
    "speech_config": [{"voice": "Kore"}]
  }
}

For an authorized direct Google account, the request is:

curl --fail-with-body --silent --show-error \
  'https://generativelanguage.googleapis.com/v1beta/interactions' \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H 'Content-Type: application/json' \
  --data-binary @request.json > response.json

Supply your own key securely through the environment. Running this command makes a provider API request and can incur usage charges. A successful local JSON parse does not prove that your account has model access or that a provider has accepted the request.

This is a single-speaker example. The official multi-speaker configuration has a different structure; do not turn the speech_config array above into an improvised dialogue schema. Each dialogue turn must use a speaker that matches the configured speakers.

Decode audio only after checking the response

An HTTP error saved into response.json is not audio. Inspect the HTTP result and the response structure before decoding base64 data. For Interactions REST responses, inspect model_output steps and their content blocks. Do not assume that an SDK convenience property such as output_audio exists in the raw JSON.

A defensive decoder should:

  1. Reject an error response instead of writing it as an audio file.
  2. Select audio content from model_output steps, ignoring text and tool content.
  3. Verify the returned MIME type before deciding on a file extension.
  4. Decode the base64 payload and preserve its native container.

The migration guide says unary output defaults to WAV. Do not automatically add a WAV header copied from an older raw-PCM example: a second header can corrupt an already valid WAV file. If you intentionally request another format, use its actual MIME type and container rather than renaming its bytes.

Save a verified WAV with a small decoder

For this unary example, request WAV explicitly by changing response_format to:

{"type":"audio","mime_type":"audio/wav"}

The Interactions response reference defines the step/content structure. Save the following as decode_tts.py in the same folder as response.json, and run python3 decode_tts.py. It uses only Python’s standard library and makes no network request.

import base64
import io
import json
import wave
from pathlib import Path

body = json.loads(Path('response.json').read_text())
if body.get('error') or body.get('status') != 'completed':
    raise SystemExit('Interaction failed or is not completed; inspect response.json')
blocks = [
    item
    for step in body.get('steps', [])
    if step.get('type') == 'model_output'
    for item in step.get('content', [])
    if item.get('type') == 'audio'
]
if len(blocks) != 1:
    raise SystemExit('Expected one audio block; inspect the response before combining audio')
audio = blocks[0]
if audio.get('mime_type') != 'audio/wav' or not isinstance(audio.get('data'), str):
    raise SystemExit('Expected inline audio/wav data; inspect MIME type and delivery mode')
payload = base64.b64decode(audio['data'], validate=True)
if payload[:4] != b'RIFF' or payload[8:12] != b'WAVE':
    raise SystemExit('Payload is not a RIFF/WAVE file')
with wave.open(io.BytesIO(payload), 'rb') as wav:
    frames = wav.getnframes()
    rate = wav.getframerate()
    if frames == 0 or rate == 0:
        raise SystemExit('Audio contains no usable frames')
    print({'channels': wav.getnchannels(), 'sample_rate': rate,
           'duration_seconds': frames / rate})
Path('speech.wav').write_bytes(payload)
print('Saved speech.wav')

Expected local output is metadata followed by Saved speech.wav. The duration and sample rate come from the returned file; no particular value is promised. The decoder deliberately stops on multiple audio blocks, a pending interaction, URI delivery or a different encoding. It is a narrow single-result WAV decoder, not a general streaming audio client. A malformed base64 string or unsupported WAV encoding also raises an error instead of silently writing a misleading file.

We checked this decoder with locally generated PCM WAV fixtures, including malformed and error cases. That proves local parsing and rejection behavior only. It does not verify account access, a live Google response, pronunciation or voice quality.

Separate container errors from speech errors

If speech.wav will not open, investigate HTTP status, MIME type and bytes before changing the narration prompt. Adding another WAV header to WAV bytes cannot fix a bad transcript. Conversely, valid metadata and a playable file do not prove the model spoke every approved word.

FailureSpecific check and correction
Delivery instruction is spokenCompare the literal transcript with the approved script; move the unwanted direction into style
Speaker label is spokenRemove the label from text and use the documented speaker metadata with a matching configuration
JSON error saved as audioStop at the HTTP/interaction error; do not base64-decode an error object
WAV is corrupt or sounds like noiseConfirm container and encoding; remove old raw-PCM wrapping when the response is already WAV
Empty or partial narrationInspect interaction completion status and finish details; check text size and model limits before splitting
Dialogue changes voice unexpectedlyVerify the turn-to-speaker mapping and test each configured voice in a short turn
No data payloadCheck whether URI delivery or streaming was requested; use that delivery mode’s documented handling

Streaming is a separate implementation task. Chunks are not necessarily independent WAV files that can be joined byte-for-byte. Likewise, concatenating several complete WAV files preserves extra headers in the middle. Keep this first migration unary; if you later need streaming or concatenation, handle the documented PCM format and timing explicitly.

Make the first migration test small

Begin with a single speaker, a short script and one delivery instruction. Listen for omitted words, extra instruction text, pronunciation and unexpected changes of voice. Save the request, model ID, timestamp and output file together. These are proposed acceptance checks, not results we obtained for this article.

Only then extend the script or add speakers. Change one variable at a time: moving an instruction into metadata while also changing voices, splitting the script and switching API families makes a failure difficult to diagnose.

For a voiceover workflow, compare the generated audio against the exact text before editing it into a video. The faceless video workflow guide covers the larger production pipeline. TTS migration is one step in that pipeline; it does not guarantee timing, pronunciation or consent for a particular voice.

Accept the narration before attaching it to video

Write the acceptance criteria before generating a long script. For the train sentence, the listener should hear exactly the approved sentence, without “calm and clear,” a speaker prefix, an omitted time or a repeated ending. Mark pronunciation, audible clipping, unwanted silence and voice changes separately. A valid file can still fail any of these checks.

For a practical migration set, use one short neutral sentence, one sentence containing a proper name or number from your project, and a two-turn dialogue only if your product needs multiple speakers. Keep transcripts and voices fixed while moving instructions into metadata. Save request/output pairs and listening notes; do not report a success percentage unless you actually score all runs using a defined denominator.

Once the speech passes, measure the real audio duration and align scenes to it. A requested 15-second video does not force a TTS clip to last 15 seconds. Edit copy or timeline intentionally and listen again after any time adjustment. Keep the original audio so that an editing defect is distinguishable from the generated speech. This completes the migration from a correctly placed instruction to a usable, checked asset without claiming a model-quality benchmark.

What to verify before routing through another provider

An OpenAI-compatible text endpoint does not imply support for Google’s Interactions API or its speech metadata fields. Check the provider’s specific audio route, supported model ID and output format. This article’s direct Google example is not an Ofox endpoint recipe.

Keep model selection separate from API-format migration. Gemini 3.8 Flash-Lite TTS is a related offering, but support and behavior must be checked against its own reference before changing the model string. Our multimodal API overview provides broader context without replacing the current provider documentation.

Frequently Asked Questions

Why might the model read a delivery instruction aloud?
The new model treats input text as a literal transcript. Put sustained delivery guidance in the documented speech metadata location rather than inserting it into the script.
Can I copy this JSON into GenerateContent?
No. This example uses Interactions input and annotation fields. GenerateContent uses a different request structure; use that API's documented example end to end.
Is the output always raw PCM?
No. The migration guidance says unary output defaults to WAV. Inspect the actual response format and do not add a second WAV header.