Gemini 3.5 Transcribe API Tutorial: Build a Transcriber
Learn how to build a speech-to-text app with Gemini 3.5 Transcribe, including Python setup, custom vocabulary, speaker diarization, timestamps, smart transcription, and pricing.
On this page
Googleβs Gemini 3.5 Transcribe is now generally available through the Gemini API, giving developers a dedicated speech-to-text model instead of relying on a general-purpose Gemini model. It can automatically identify languages, separate speakers, add word-level timestamps, and recognize technical terms supplied through custom vocabulary. This tutorial shows how to upload an audio file, transcribe it with Python, and then configure the model for the cases where a basic transcript is not enough.
Why Gemini 3.5 Transcribe changes the transcription workflow
The useful distinction is that Gemini 3.5 Transcribe is built specifically for turning speech into text. Google says the model supports automatic detection across more than 85 locales, handles code-switching between languages, and provides dedicated controls for speaker diarization, timestamps, and vocabulary biasing. That means you can treat transcription as a structured API task rather than asking a general AI model to understand an audio recording and hoping the resulting text is usable. Google also released a separate live version, gemini-3.5-transcribe-live, for low-latency streaming over WebSockets, so the model covered here is best suited to uploaded or otherwise non-streaming audio.
What you need before making the first request
You need a Google AI Studio API key and the Google GenAI SDK for your chosen programming language. The examples below use Python because it keeps the first working request small and makes the configuration easier to see. Install the SDK with pip install google-genai, then make the API key available through the GEMINI_API_KEY environment variable rather than putting the secret directly into source code. Gemini 3.5 Transcribe accepts common audio formats including MP3, WAV, FLAC, AAC, OGG, M4A, WebM, and Opus.
Upload an audio file and get your first transcript
The simplest workflow has two operations: upload the audio through the Files API, then pass the returned file URI to the Interactions API. The model identifier is gemini-3.5-transcribe. Once the request finishes, the transcript is available as the interaction's output text.
from google import genai
client = genai.Client()
audio_file = client.files.upload(
file="sample.mp3"
)
interaction = client.interactions.create(
model="gemini-3.5-transcribe",
input=[
{
"type": "audio",
"uri": audio_file.uri,
"mime_type": audio_file.mime_type,
}
],
)
print(interaction.output_text)Run the script against a short recording first. If the setup is correct, the terminal should contain the spoken content as text. Google also documents a Generate Content route for the same model, but the Interactions API example is useful when you want to build on the returned interaction in a larger application.
Choose the language when automatic detection is not enough
Gemini 3.5 Transcribe detects the spoken language automatically when you leave language_codes empty or omit it. That is convenient for recordings where the language is unknown or speakers switch languages during a conversation. If you already know the language, supplying a BCP-47 language code can give the recognizer a clearer signal about what it should expect.
interaction = client.interactions.create(
model="gemini-3.5-transcribe",
input=[
{
"type": "audio",
"uri": audio_file.uri,
"mime_type": audio_file.mime_type,
}
],
generation_config={
"transcription_config": {
"language_codes": ["en-US"],
}
},
)
print(interaction.output_text)The practical rule is simple: use automatic detection for mixed or unknown recordings, and provide a known language when your application already has that information. You do not need to hard-code a language merely because the application normally handles one language.
Add technical terms with custom vocabulary
Normal speech recognition can struggle with product names, acronyms, unusual names, and specialist terminology. Gemini 3.5 Transcribe lets you provide up to 1,000 custom vocabulary phrases to bias recognition toward those terms, although Google says the best results are typically achieved with up to 100 terms. This is particularly useful for technical interviews, medical terminology, software demonstrations, and company-specific names.
interaction = client.interactions.create(
model="gemini-3.5-transcribe",
input=[
{
"type": "audio",
"uri": audio_file.uri,
"mime_type": audio_file.mime_type,
}
],
generation_config={
"transcription_config": {
"custom_vocabulary": [
"Kubernetes",
"BigQuery",
"Gemini",
"Cloud Run",
]
}
},
)
print(interaction.output_text)The important point is that vocabulary biasing is not a replacement for good audio. It gives the recognizer useful hints about terms that matter to your application, so it works best when the supplied words are genuinely relevant to the recording.
Separate speakers when the recording is a conversation
A plain transcript tells you what was said but can lose who said it. Speaker diarization solves that by assigning recognized speech to speaker labels such as spk_1 and spk_2. Gemini 3.5 Transcribe supports diarization for multiple speakers, with attribution for three or more speakers currently described as experimental in Google's documentation.
interaction = client.interactions.create(
model="gemini-3.5-transcribe",
input=[
{
"type": "audio",
"uri": audio_file.uri,
"mime_type": audio_file.mime_type,
}
],
generation_config={
"transcription_config": {
"mode": {
"type": "verbatim",
"diarization_mode": "speaker",
}
}
},
)
print(interaction.output_text)Use this mode for interviews, meetings, podcasts, and support calls where speaker identity is part of the useful output. Do not use smart transcription at the same time: Google's documentation says the smart mode is incompatible with speaker diarization and word-level timestamps.
Use word timestamps when the application needs timing
Word-level timestamps attach precise start and end offsets to recognized words. That makes them useful for subtitles, searchable recordings, audio editors, and applications that need to jump from text back to the corresponding point in an audio file. Enable them through the transcription mode configuration.
interaction = client.interactions.create(
model="gemini-3.5-transcribe",
input=[
{
"type": "audio",
"uri": audio_file.uri,
"mime_type": audio_file.mime_type,
}
],
generation_config={
"transcription_config": {
"mode": {
"type": "verbatim",
"timestamp_granularities": ["word"],
}
}
},
)
print(interaction.output_text)There is a trade-off worth knowing before enabling this automatically. Google notes that requesting word-level timestamps may reduce overall transcription accuracy. If your application only needs readable text, leave timestamps off. If synchronization is the primary requirement, accept the trade-off and test the resulting output with your own audio.
Use smart transcription for cleaner readable text
Verbatim transcription preserves the way people actually speak, including filler words, repetitions, false starts, and pauses. Smart transcription instead cleans the result for reading: it can remove conversational fillers, resolve spoken self-corrections, add punctuation, and turn spoken lists into structured text. For example, a speaker correcting βTuesdayβ to βWednesdayβ can produce a cleaned sentence rather than preserving the entire false start.
interaction = client.interactions.create(
model="gemini-3.5-transcribe",
input=[
{
"type": "audio",
"uri": audio_file.uri,
"mime_type": audio_file.mime_type,
}
],
generation_config={
"transcription_config": {
"mode": "smart"
}
},
)
print(interaction.output_text)This creates an important design choice for your application. Use smart mode when the transcript is going straight to a reader, such as meeting notes or a cleaned interview draft. Use verbatim mode when preserving the exact spoken record matters, especially if you also need timestamps or speaker labels.
How much Gemini 3.5 Transcribe costs
Google's current standard paid pricing lists Gemini 3.5 Transcribe at $2 per million input audio tokens and $12 per million output text tokens. Google also provides an estimated blended rate of about $0.005 per minute, based on its token assumptions. That makes the dedicated model relatively inexpensive for applications processing large amounts of recorded speech, although your actual bill depends on the audio and generated text token consumption.
| Item | Current standard price | What it means |
|---|---|---|
| Audio input | $2 per 1M tokens | The audio sent for transcription is billed as input. |
| Text output | $12 per 1M tokens | The generated transcript is billed as output. |
| Estimated blended rate | About $0.005 per minute | Google's estimate based on its stated token assumptions. |
There is also a free tier, subject to Google's current limits and model availability. For a prototype, that gives developers a way to test the transcription pipeline before moving to paid usage. Production applications should still measure their own audio lengths and output sizes instead of treating the estimated per-minute figure as a guaranteed invoice amount.
Know when to use Transcribe Live instead
If the application needs text while somebody is still speaking, the non-streaming model is the wrong starting point. Google provides gemini-3.5-transcribe-live for bidirectional streaming speech recognition over WebSockets, with interim and finalized transcription events and low-latency processing. A call assistant, live captioning system, or voice interface therefore has a different architecture from an application that processes uploaded recordings after they finish.
The distinction also prevents a common design mistake: trying to simulate real-time transcription by repeatedly uploading small audio files to the non-streaming endpoint. Use the dedicated live model when latency is part of the product requirement, and keep gemini-3.5-transcribe for recorded or non-streaming workflows.
Build the configuration around the job, not the model name
The biggest advantage of Gemini 3.5 Transcribe is not simply that it produces text. It gives you controls that map directly to different transcription jobs: automatic language detection for unknown recordings, custom vocabulary for specialized terminology, diarization when speaker identity matters, timestamps when synchronization matters, and smart mode when readability matters. Those settings should be selected according to the output your application needs rather than enabled all at once.
Google reports that Artificial Analysis measured an average word error rate of 2.6% for its non-streaming use case and 4.0% for streaming in its evaluation. Those figures are useful context, but they are not a guarantee for your recordings; accents, noise, terminology, microphone quality, overlapping speech, and configuration can change results. The safest production workflow is to assemble a representative test set from the audio your application will actually receive, then compare the transcript against human-checked text before deciding which configuration to ship.
For a first implementation, start with the basic upload-and-transcribe request, add a known language only when you have reliable metadata, and introduce custom vocabulary when domain terminology causes mistakes. Add diarization or timestamps only when the product genuinely needs them, and choose smart mode when clean prose matters more than preserving the exact speech. That keeps the API simple while giving you a clear path from a working prototype to a transcription system designed around the real job.
Written by


