Z.ai

GLM-5.3-Flash

Z.ai's first natively multimodal GLM-5 model for visual coding and long-horizon agent work, with a 1M-token context window and always-on reasoning at Flash-tier rates.

ActiveReleased August 26, 2026Available from 24 providers
Tool callingCan invoke external tools and APIs (function calling) mid-response.Structured outputCan return JSON or schema-constrained responses instead of free text.ReasoningUses extended thinking / chain-of-thought for harder multi-step problems.Effort controlLets you dial thinking depth (low / medium / high) to trade cost for quality.Fine-tuningSupports training a custom variant on your own data (vendor fine-tune API or self-hosted open weights).

Source: Z.ai GLM-5.3-Flash model documentation, Z.ai GLM-5.3-Flash announcement, Z.ai GLM API pricing, Z.ai deep thinking documentation, Z.ai context caching documentation

Price per 1M tokens via DeepInfra (OpenRouter)
$0.075 in / $0.25 out
Context window1,000,000
Max output131,072
Cache read$0.015 / M
ModalitiesText, Image, Video → Text
Priced across24 providers

Pricing

Every billing leg from the live catalog, in USD per 1M tokens.

Billing legRate
Input$0.075 / M
Output$0.25 / M
Cache read$0.015 / M

Cost calculator

Billing-grade math over the rate card above, including the thinking tokens most estimates miss.

Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.

Estimated monthly cost
$30.00
at current list prices
Fresh input$15.00
Cached input$0.00
Output incl. thinking$15.00
Track this automatically →

Compare providers

The same model, priced across every platform that serves it. Lowest combined in/out rate is flagged.

PlatformInput / MOutput / MCache read / MBatchAccess
RelaceLowest
$0.07125$0.2375$0.01425-via OpenRouter
DeepInfra
$0.075$0.25$0.015-via OpenRouter
GMICloud
$0.075$0.25$0.015-via OpenRouter
Novita
$0.075$0.25$0.015-via OpenRouter
Z.ai
$0.075$0.25$0.015-via OpenRouter
Wafer
$0.1$0.35$0.02-via OpenRouter
Morph
$0.13$0.45$0.02-via OpenRouter
Makora
$0.14$0.47$0.024-via OpenRouter
Modal
$0.149985$0.49995$0.029997-via OpenRouter
Baseten
$0.15$0.5$0.03-via OpenRouter
Cloudflare
$0.15$0.5$0.03-via OpenRouter
DigitalOcean
$0.15$0.5$0.03-via OpenRouter
Fireworks AI
$0.15$0.5$0.03-via OpenRouter
Friendli
$0.15$0.5$0.03-via OpenRouter
Io-net
$0.15$0.5$0.03-via OpenRouter
NextBit
$0.15$0.5$0.03-via OpenRouter
Parasail
$0.15$0.5$0.03-via OpenRouter
Phala
$0.15$0.5$0.03-via OpenRouter
Reka
$0.15$0.5$0.03-via OpenRouter
Sail-research
$0.15$0.5$0.03-via OpenRouter
SiliconFlow
$0.15$0.5$0.03-via OpenRouter
StreamLake
$0.15$0.5$0.03-via OpenRouter
Together AI
$0.15$0.5$0.03-via OpenRouter
Venice
$0.15$0.5$0.03-via OpenRouter

Price history

Has this model gotten cheaper or more expensive?

EffectiveInput / OutputStatus
August 27, 2026 – August 30, 2026$0.15 / $0.5past
August 30, 2026$0.075 / $0.25current

More from Z.ai

Other models from the same provider.

Stop estimating. Track it.

SuperPenguin meters every GLM-5.3-Flash call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.

Get started free

Rates are generated from SuperPenguin's live pricing catalog.