Z.ai

GLM-5

The GLM-5 release: a strong open-weight reasoning and coding model, available across inference gateways.

ActiveReleased February 12, 2026Available from 17 providers
Tool callingCan invoke external tools and APIs (function calling) mid-response.Structured outputCan return JSON or schema-constrained responses instead of free text.ReasoningUses extended thinking / chain-of-thought for harder multi-step problems.Effort controlLets you dial thinking depth (low / medium / high) to trade cost for quality.Fine-tuningSupports training a custom variant on your own data (vendor fine-tune API or self-hosted open weights).

Source: Z.ai GLM API pricing, Z.ai GLM models, Metadata verification, Field resolution verification

Price per 1M tokens via GMICloud (OpenRouter)
$0.6 in / $1.92 out
Context window204,800
Max output131,072
Cache read$0.12 / M
Input
  • Text
Output
  • Text
Priced across17 providers

Pricing

Every billing leg from the live catalog, in USD per 1M tokens.

Billing legRate
Input$0.6 / M
Output$1.92 / M
Cache read$0.12 / M

Cost calculator

Billing-grade math over the rate card above, including the thinking tokens most estimates miss.

Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.

Estimated monthly cost
$235
at current list prices
Fresh input$120
Cached input$0.00
Output incl. thinking$115
Track this automatically →

Compare providers

The same model, priced across every platform that serves it. Lowest combined in/out rate is flagged.

PlatformRegionInput / MOutput / MCache read / MBatchAccess
GMICloudLowest
-$0.6$1.92$0.12-via OpenRouter
StreamLake
-$0.6$1.92$0.12-via OpenRouter
DeepInfra
-$0.6$2.08$0.12-via OpenRouter
Baidu
-$0.7$2.24$0.14-via OpenRouter
Chutes
-$0.95$2.55$0.475-via OpenRouter
SiliconFlow
-$0.95$2.55$0.2-via OpenRouter
AtlasCloud
-$0.95$3.15$0.19-via OpenRouter
Together AI
-$1$3.2--Direct
AWS Bedrock
$1$3.2-✓Direct
AWS Bedrock
-$1$3.2--via OpenRouter
DigitalOcean
-$1$3.2$0.2-via OpenRouter
Novita
-$1$3.2$0.2-via OpenRouter
Parasail
-$1$3.2$0.2-via OpenRouter
Venice
-$1$3.2$0.2-via OpenRouter
Z.ai
-$1$3.2$0.2-via OpenRouter
Azure AI Foundry (Fireworks)
-$1.1$3.52$0.22-Direct
Phala
-$1.2$3.5$0.475-via OpenRouter

Some hosts price by region. Shown rates match the selected region.

Price history

Has this model gotten cheaper or more expensive?

No changePriced at $0.6 / $1.92 per 1M since July 10, 2026.

More from Z.ai

Other models from the same provider.

Stop estimating. Track it.

SuperPenguin meters every GLM-5 call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.

Get started free

Rates are generated from SuperPenguin's live pricing catalog.