Z.ai

GLM-5.3

Z.ai's latest flagship model for complex software engineering and long-horizon agent tasks, with a 1M-token context window and always-on reasoning.

ActiveReleased August 14, 2026Available from 24 providers
Tool callingCan invoke external tools and APIs (function calling) mid-response.Structured outputCan return JSON or schema-constrained responses instead of free text.ReasoningUses extended thinking / chain-of-thought for harder multi-step problems.Effort controlLets you dial thinking depth (low / medium / high) to trade cost for quality.Fine-tuningSupports training a custom variant on your own data (vendor fine-tune API or self-hosted open weights).

Source: Z.ai GLM-5.3 model documentation, Z.ai GLM API pricing, Z.ai release notes, Z.ai GLM-5.3 announcement, Z.ai context caching documentation, Z.ai deep thinking documentation

Price per 1M tokens via Baseten (OpenRouter)
$1.4 in / $4.4 out
Context window1,000,000
Max output131,072
Cache read$0.14 / M
ModalitiesText → Text
Priced across24 providers

Pricing

Every billing leg from the live catalog, in USD per 1M tokens.

Billing legRate
Input$1.4 / M
Output$4.4 / M
Cache read$0.14 / M

Cost calculator

Billing-grade math over the rate card above, including the thinking tokens most estimates miss.

Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.

Estimated monthly cost
$544
at current list prices
Fresh input$280
Cached input$0.00
Output incl. thinking$264
Track this automatically →

Compare providers

The same model, priced across every platform that serves it. Lowest combined in/out rate is flagged.

PlatformInput / MOutput / MCache read / MBatchAccess
RekaLowest
$1.15$3.5$0.1-via OpenRouter
Decart
$1.16$3.65$0.1909-via OpenRouter
Akashml
$1.17$3.96$0.234-via OpenRouter
DeepInfra
$1.2$4$0.24-via OpenRouter
Friendli
$1.26$3.96$0.234-via OpenRouter
Io-net
$1.19$4.18$0.247-via OpenRouter
Wafer
$1.19$4.4$0.26-via OpenRouter
Inceptron
$1.25$4.4$0.26-via OpenRouter
Morph
$1.25$4.4$0.26-via OpenRouter
Makora
$1.35$4.4$0.23-via OpenRouter
Baseten
$1.4$4.4$0.14-via OpenRouter
Cloudflare
$1.4$4.4$0.26-via OpenRouter
DigitalOcean
$1.4$4.4$0.26-via OpenRouter
Fireworks AI
$1.4$4.4$0.26-via OpenRouter
GMICloud
$1.4$4.4$0.26-via OpenRouter
Modal
$1.4$4.4$0.26-via OpenRouter
Novita
$1.4$4.4$0.26-via OpenRouter
Parasail
$1.4$4.4$0.26-via OpenRouter
Phala
$1.4$4.4$0.26-via OpenRouter
Sail-research
$1.4$4.4$0.26-via OpenRouter
SiliconFlow
$1.4$4.4$0.26-via OpenRouter
Together AI
$1.4$4.4$0.26-via OpenRouter
Z.ai
$1.4$4.4$0.26-via OpenRouter
Venice
$1.75$5.5$0.325-via OpenRouter

Price history

Has this model gotten cheaper or more expensive?

No changePriced at $1.4 / $4.4 per 1M since August 29, 2026.

More from Z.ai

Other models from the same provider.

Stop estimating. Track it.

SuperPenguin meters every GLM-5.3 call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.

Get started free

Rates are generated from SuperPenguin's live pricing catalog.