GLM-5.3
Z.ai's latest flagship model for complex software engineering and long-horizon agent tasks, with a 1M-token context window and always-on reasoning.
Source: Z.ai GLM-5.3 model documentation, Z.ai GLM API pricing, Z.ai release notes, Z.ai GLM-5.3 announcement, Z.ai context caching documentation, Z.ai deep thinking documentation
Pricing
Every billing leg from the live catalog, in USD per 1M tokens.
| Billing leg | Rate |
|---|---|
| Input | $1.4 / M |
| Output | $4.4 / M |
| Cache read | $0.26 / M |
Cost calculator
Billing-grade math over the rate card above, including the thinking tokens most estimates miss.
Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.
Provider
Where this model runs, and what it charges.
| Platform | Input / M | Output / M | Cache read / M | Batch | Access |
|---|---|---|---|---|---|
Z.ai | $1.4 | $4.4 | $0.26 | — | via OpenRouter |
Price history
Has this model gotten cheaper or more expensive?
More from Z.ai
Other models from the same provider.
Stop estimating. Track it.
SuperPenguin meters every GLM-5.3 call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.
Rates are generated from SuperPenguin's live pricing catalog.