Gemini 3.5 Flash-Lite
Google's most cost-efficient Gemini 3.5-class model for high-volume agentic tasks, translation, and simple data processing on a 1M-token context window.
Source: Gemini API models, Gemini API pricing
Pricing
Every billing leg from the live catalog, in USD per 1M tokens.
| Billing leg | Rate | Batch |
|---|---|---|
| Input | $0.3 / M | $0.15 / M |
| Output | $2.5 / M | $1.25 / M |
| Cache read | $0.03 / M | $0.02 / M |
Cost calculator
Billing-grade math over the rate card above, including the thinking tokens most estimates miss.
Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.
Provider
Where this model runs, and what it charges.
| Platform | Input / M | Output / M | Cache read / M | Batch | Access |
|---|---|---|---|---|---|
Google | $0.3 | $2.5 | $0.03 | ✓ | First-party API |
Price history
Has this model gotten cheaper or more expensive?
More from Google
Other models from the same provider.
Stop estimating. Track it.
SuperPenguin meters every Gemini 3.5 Flash-Lite call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.
Rates are generated from SuperPenguin's live pricing catalog.