GLM-5.3-Flash
Z.ai's first natively multimodal GLM-5 model for visual coding and long-horizon agent work, with a 1M-token context window and always-on reasoning at Flash-tier rates.
Source: Z.ai GLM-5.3-Flash model documentation, Z.ai GLM-5.3-Flash announcement, Z.ai GLM API pricing, Z.ai deep thinking documentation, Z.ai context caching documentation
Pricing
Every billing leg from the live catalog, in USD per 1M tokens.
| Billing leg | Rate |
|---|---|
| Input | $0.075 / M |
| Output | $0.25 / M |
| Cache read | $0.015 / M |
Cost calculator
Billing-grade math over the rate card above, including the thinking tokens most estimates miss.
Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.
Compare providers
The same model, priced across every platform that serves it. Lowest combined in/out rate is flagged.
| Platform | Input / M | Output / M | Cache read / M | Batch | Access |
|---|---|---|---|---|---|
GMICloudLowest | $0.075 | $0.25 | $0.015 | - | via OpenRouter |
Novita | $0.075 | $0.25 | $0.015 | - | via OpenRouter |
Z.ai | $0.075 | $0.25 | $0.015 | - | via OpenRouter |
Venice | $0.09375 | $0.3125 | $0.01875 | - | via OpenRouter |
Modal | $0.149985 | $0.49995 | $0.029997 | - | via OpenRouter |
Baseten | $0.15 | $0.5 | $0.03 | - | via OpenRouter |
Cloudflare | $0.15 | $0.5 | $0.03 | - | via OpenRouter |
DeepInfra | $0.15 | $0.5 | $0.03 | - | via OpenRouter |
Io-net | $0.15 | $0.5 | $0.03 | - | via OpenRouter |
Parasail | $0.15 | $0.5 | - | - | via OpenRouter |
Price history
Has this model gotten cheaper or more expensive?
More from Z.ai
Other models from the same provider.
Stop estimating. Track it.
SuperPenguin meters every GLM-5.3-Flash call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.
Rates are generated from SuperPenguin's live pricing catalog.

