Inkling
Thinking Machines Lab's flagship open-weight multimodal MoE for reasoning, tool use, and fine-tuning. Available on Tinker (64K and 256K context), Together AI serverless (1M context), and other inference gateways.
Source: Tinker models & pricing, Inkling model card, Inkling announcement, Inkling on Together AI
- Text
- Image
- Audio
- Text
Pricing
Every billing leg from the live catalog, in USD per 1M tokens.
| Billing leg | Standard | Long context |
|---|---|---|
| Input | $1.87 / M | $3.74 / M |
| Output | $4.68 / M | $9.36 / M |
| Cache read | $0.374 / M | $0.748 / M |
Long-context rates apply above 65,536 input tokens.
Cost calculator
Billing-grade math over the rate card above, including the thinking tokens most estimates miss.
Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.
Compare providers
The same model, priced across every platform that serves it. Lowest combined in/out rate is flagged.
| Platform | Input / M | Output / M | Cache read / M | Batch | Access |
|---|---|---|---|---|---|
Thinking Machines | $1.87 | $4.68 | $0.374 | - | First-party API |
DeepInfraLowest | $0.95 | $4.05 | $0.16 | - | via OpenRouter |
Together AI | $1 | $4.05 | $0.17 | - | Direct |
Baseten | $1 | $4.05 | $0.17 | - | via OpenRouter |
Together AI | $1 | $4.05 | $0.17 | - | via OpenRouter |
Price history
Has this model gotten cheaper or more expensive?
| Effective | Input / Output | Status |
|---|---|---|
| July 15, 2026 – August 12, 2026 | $1.87 / $4.68 | past |
| August 12, 2026 – August 21, 2026 | $0.95 / $4.05 | past |
| August 21, 2026 | $1.87 / $4.68 | current |
More from Thinking Machines Lab
Other models from the same provider.
Stop estimating. Track it.
SuperPenguin meters every Inkling call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.
Rates are generated from SuperPenguin's live pricing catalog.
Thinking Machines Lab