NVIDIA

Nemotron 3 Nano

NVIDIA's efficient Nemotron 3 Nano MoE (~30B total / ~3B active) for specialized agentic sub-tasks at low inference cost, with up to a 1M-token context window on AWS Bedrock and other GPU hosts.

ActiveAvailable from 4 providers
Tool callingCan invoke external tools and APIs (function calling) mid-response.Structured outputCan return JSON or schema-constrained responses instead of free text.ReasoningUses extended thinking / chain-of-thought for harder multi-step problems.Effort controlLets you dial thinking depth (low / medium / high) to trade cost for quality.Fine-tuningSupports training a custom variant on your own data (vendor fine-tune API or self-hosted open weights).

Source: NVIDIA Nemotron, Nemotron 3 family, AWS Bedrock pricing (NVIDIA)

Price per 1M tokens
$0.06 in / $0.24 out
Context window1,048,576
Max output131,072
Batch discount-50%
ModalitiesText → Text
Priced across4 providers

Specs

Details that do not fit the rate card above.

ArchitectureHybrid Mamba-Transformer MoE; ~3B active / ~30B total

Pricing

Every billing leg from the live catalog, in USD per 1M tokens.

Billing legRateBatch
Input$0.06 / M$0.03 / M
Output$0.24 / M$0.12 / M

Cost calculator

Billing-grade math over the rate card above, including the thinking tokens most estimates miss.

Hidden reasoning tokens bill as output. 1.0x = thinking off. Medium-effort reasoning typically lands near 3x; heavy reasoning runs 6x to 8x. Set it to match your workload.

Estimated monthly cost
$26.40
at current list prices
Fresh input$12.00
Output incl. thinking$14.40
Track this automatically →

Compare providers

The same model, priced across every platform that serves it. Lowest combined in/out rate is flagged.

PlatformInput / MOutput / MCache read / MBatchAccess
AWS Bedrock
$0.06$0.24Direct
CrusoeLowest
$0.05$0.2$0.03via OpenRouter
DeepInfra
$0.05$0.2$0.025via OpenRouter
Novita
$0.05$0.2via OpenRouter

Price history

Has this model gotten cheaper or more expensive?

No changePriced at $0.06 / $0.24 per 1M since July 10, 2026.

More from NVIDIA

Other models from the same provider.

Stop estimating. Track it.

SuperPenguin meters every Nemotron 3 Nano call your team makes, prices it with this exact rate card, and shows spend by feature, customer, and team.

Get started free

Rates are generated from SuperPenguin's live pricing catalog.