NVIDIA: Nemotron 3 Ultra
NVIDIA's Nemotron 3 Ultra is designed to handle text-based tasks with an impressive context length of up to one million tokens and supports both reasoning and tool integration for enhanced functionality. While it excels in processing complex text inputs, structured output capabilities are currently unimplemented. Given its pricing at $0.6 per million input tokens and $3.6 per million output tokens, Nemotron 3 Ultra might be a cost-effective choice for organizations requiring robust text analysis or reasoning with access to tools. However, without independent benchmarking, users should proceed with caution, especially when seeking definitive performance metrics.
- Model ID
- nvidia/nemotron-3-ultra-550b-a55b
- Vendor
- nvidia
- Released
- June 2026
- Tokenizer
- Other
- Input Modalities
- text
- Output Modalities
- text
- Max Output
- default
- Tool Calling
- ✓ supported
- Structured Output
- ✓ supported
- Reasoning Mode
- ✓ supported
- Vision
- text only
- Audio
- no
- Moderated
- no
What it costs in practice
Computed from the current $0.60/M input and $3.60/M output rates. Run your own numbers →
| Job | Tokens | Cost |
|---|---|---|
| Summarize a 50-page report | 30k in / 1.5k out | $0.02 |
| Classify 1,000 customer emails | 500k in / 50k out | $0.48 |
| A month of a busy support chatbot | 5M in / 2M out | $10.20 |
Similar models
NVIDIA: Nemotron 3 Super
NVIDIA: Nemotron 3 Nano 30B A3B
NVIDIA: Nemotron 3 Nano Omni (free)
NVIDIA: Nemotron 3 Super (free)
NVIDIA: Nemotron Nano 12B 2 VL (free)
NVIDIA: Nemotron 3 Ultra (free)
Quick answers
- How much does NVIDIA: Nemotron 3 Ultra cost?
- $0.60 per million input tokens and $3.60 per million output tokens.
- What is NVIDIA: Nemotron 3 Ultra's context window?
- 1,000,000 tokens, roughly 1,500 pages of text in a single request.
- Does NVIDIA: Nemotron 3 Ultra support tool calling?
- It supports tool calling, structured output, a reasoning mode.
- Can NVIDIA: Nemotron 3 Ultra process images?
- No, it is text-only on the input side.