nvidia

NVIDIA: Nemotron 3 Ultra

NVIDIA's Nemotron 3 Ultra is designed to handle text-based tasks with an impressive context length of up to one million tokens and supports both reasoning and tool integration for enhanced functionality. While it excels in processing complex text inputs, structured output capabilities are currently unimplemented. Given its pricing at $0.6 per million input tokens and $3.6 per million output tokens, Nemotron 3 Ultra might be a cost-effective choice for organizations requiring robust text analysis or reasoning with access to tools. However, without independent benchmarking, users should proceed with caution, especially when seeking definitive performance metrics.

Quality Score
97/100
price + capability + benchmarks
Input Price
$0.60
per 1M tokens
Output Price
$3.60
per 1M tokens
Context Window
1,000,000
tokens
Model ID
nvidia/nemotron-3-ultra-550b-a55b
Vendor
nvidia
Released
June 2026
Tokenizer
Other
Input Modalities
text
Output Modalities
text
Max Output
default
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
✓ supported
Vision
text only
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.60/M input and $3.60/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out $0.02
Classify 1,000 customer emails 500k in / 50k out $0.48
A month of a busy support chatbot 5M in / 2M out $10.20

Similar models

Quick answers

How much does NVIDIA: Nemotron 3 Ultra cost?
$0.60 per million input tokens and $3.60 per million output tokens.
What is NVIDIA: Nemotron 3 Ultra's context window?
1,000,000 tokens, roughly 1,500 pages of text in a single request.
Does NVIDIA: Nemotron 3 Ultra support tool calling?
It supports tool calling, structured output, a reasoning mode.
Can NVIDIA: Nemotron 3 Ultra process images?
No, it is text-only on the input side.