Nemotron 3 Ultra
NVIDIA's Nemotron 3 Ultra (550B total / 55B active MoE) for frontier agentic reasoning and long-running orchestration, on Together AI, Fireworks, and other GPU hosts.
Source: NVIDIA Nemotron, Nemotron 3 family, AWS Bedrock pricing (NVIDIA)
Specs
Details that do not fit the rate card above.
Pricing
Every billing leg from the live catalog, in USD per 1M tokens.
| Billing leg | Rate |
|---|---|
| Input | $0.6 / M |
| Output | $3.6 / M |
| Cache read | $0.2 / M |
Cost calculator
Billing-grade math over the rate card above, including the thinking tokens most estimates miss.
Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.
Compare providers
The same model, priced across every platform that serves it. Lowest combined in/out rate is flagged.
| Platform | Input / M | Output / M | Cache read / M | Batch | Access |
|---|---|---|---|---|---|
Together AI | $0.6 | $3.6 | $0.2 | — | Direct |
DeepInfraLowest | $0.5 | $2.2 | $0.1 | — | via OpenRouter |
Fireworks AI | $0.6 | $2.4 | $0.12 | ✓ | Direct |
Baseten | $0.6 | $2.4 | $0.12 | — | via OpenRouter |
Venice | $0.625 | $3.13 | $0.1875 | — | via OpenRouter |
Together AI | $0.6 | $3.6 | $0.2 | — | via OpenRouter |
Price history
Has this model gotten cheaper or more expensive?
More from NVIDIA
Other models from the same provider.
Stop estimating. Track it.
SuperPenguin meters every Nemotron 3 Ultra call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.
Rates are generated from SuperPenguin's live pricing catalog.