Nemotron 3 Super
NVIDIA's Nemotron 3 Super (120B total / 12B active) for multi-agent reasoning and tool use, with a 1M-token context window on AWS Bedrock and other GPU hosts.
Source: NVIDIA Nemotron, Nemotron 3 family, AWS Bedrock pricing (NVIDIA)
Specs
Details that do not fit the rate card above.
Pricing
Every billing leg from the live catalog, in USD per 1M tokens.
| Billing leg | Rate | Batch |
|---|---|---|
| Input | $0.15 / M | $0.075 / M |
| Output | $0.65 / M | $0.325 / M |
Cost calculator
Billing-grade math over the rate card above, including the thinking tokens most estimates miss.
Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.
Compare providers
The same model, priced across every platform that serves it. Lowest combined in/out rate is flagged.
| Platform | Input / M | Output / M | Cache read / M | Batch | Access |
|---|---|---|---|---|---|
AWS Bedrock | $0.15 | $0.65 | — | ✓ | Direct |
DeepInfraLowest | $0.085 | $0.4 | — | — | via OpenRouter |
DigitalOcean | $0.165 | $0.3575 | $0.06 | — | via OpenRouter |
AWS Bedrock | $0.18 | $0.78 | — | — | Direct |
AWS Bedrock · Regional | from$0.18 | $0.78 | — | — | Direct |
Nebius | $0.3 | $0.9 | — | — | via OpenRouter |
Regional deployments are priced per region; the lowest available region rate is shown.
Price history
Has this model gotten cheaper or more expensive?
More from NVIDIA
Other models from the same provider.
Stop estimating. Track it.
SuperPenguin meters every Nemotron 3 Super call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.
Rates are generated from SuperPenguin's live pricing catalog.