Google

Gemini 3.5 Transcribe

Google's speech-to-text model for precise transcription of recorded audio, with speaker attribution, word-level timestamps, and a 96K-token context window.

ActiveReleased August 26, 2026Knowledge cutoff January 1, 2025
Tool callingCan invoke external tools and APIs (function calling) mid-response.Structured outputCan return JSON or schema-constrained responses instead of free text.ReasoningUses extended thinking / chain-of-thought for harder multi-step problems.Effort controlLets you dial thinking depth (low / medium / high) to trade cost for quality.Fine-tuningSupports training a custom variant on your own data (vendor fine-tune API or self-hosted open weights).

Source: Gemini 3.5 Transcribe announcement, Gemini 3.5 Audio model card, Gemini API pricing, Gemini Enterprise Agent Platform pricing

Price per 1M tokens
$2 in / $12 out
Context window98,304
Max output32,768
ModalitiesAudio, Text → Text

Pricing

Every billing leg from the live catalog, in USD per 1M tokens.

Billing legRate
Input$2 / M
Output$12 / M

Cost calculator

Billing-grade math over the rate card above, including the thinking tokens most estimates miss.

Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.

Estimated monthly cost
$1.1K
at current list prices
Fresh input$400
Output incl. thinking$720
Track this automatically →

Provider

Where this model runs, and what it charges.

PlatformInput / MOutput / MBatchAccess
Google
$2$12-First-party API

Price history

Has this model gotten cheaper or more expensive?

No changePriced at $2 / $12 per 1M since August 27, 2026.

More from Google

Other models from the same provider.

Stop estimating. Track it.

SuperPenguin meters every Gemini 3.5 Transcribe call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.

Get started free

Rates are generated from SuperPenguin's live pricing catalog.