minimax
// nvidia
Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b
Nemotron 3 Ultra costs $0.600 per 1M input tokens and $2.40 per 1M output tokens with a 262K-token context window.
A reasoning, tool-calling model from nvidia, released Jun 4, 2026. On LiveBench it scores 67.4 overall — #49 of 55 benchmarked models — and ranks highest in instruction following (#12).
- Input / 1M
- $0.600
- Output / 1M
- $2.40
- Context
- 262K
- LiveBench overall
- 67.4 · #49
Specs and pricing
| Input price USD per 1M prompt tokens | $0.600 |
|---|---|
| Output price USD per 1M completion tokens, reasoning included | $2.40 |
| Blended price 3:1 input:output — #132 most expensive of 343 listed models | $1.05 |
| Cached input read Blank means not priced separately — not free | $0.120 |
| Cache write | — |
| Context window | 262,144 tokens |
| Max output | 182,520 tokens |
| Input modalities | text |
| Tool calling | Yes |
| Extended reasoning | Yes |
| Knowledge cutoff | — |
| Listed | Jun 4, 2026 |
Also listed as nvidia/nemotron-3-ultra-550b-a55b:free (Free blended)— the same model at a different price or rate limit.
Benchmarks by category
LiveBench scores out of 100, with Nemotron 3 Ultra's rank among the 55 benchmarked models listed here. Run: nemotron-3-ultra-550b-a55b.
$0.0483 measured per point
What Nemotron 3 Ultra costs to run
Monthly list cost across five workload shapes, using the published cached-input rate where there is one. A floor, not a quote — batch discounts and cache writes aren't included.
| Workload | Per month |
|---|---|
| Support chatbot 1.2K in / 400 out × 200K requests | $301.44/mo |
| RAG assistant 8K in / 600 out × 100K requests | $432.00/mo |
| Coding agent 40K in / 4K out × 20K requests | $403.20/mo |
| Document extraction 20K in / 1.5K out × 50K requests | $756.00/mo |
| Bulk classification 500 in / 20 out × 5M requests | $1,500/mo |
Compare Nemotron 3 Ultra with
The benchmarked models closest to it on overall score.
x-ai
Nemotron 3 Ultra vs SpaceXAI: Grok Build 0.1
openai
Nemotron 3 Ultra vs GPT-5.4 Mini
moonshotai
Nemotron 3 Ultra vs Kimi K2.7 Code
qwen
Nemotron 3 Ultra vs Qwen3.6 Plus
deepseek
Nemotron 3 Ultra vs DeepSeek V4 Flash 0423
Nemotron 3 Ultra FAQ
How much does Nemotron 3 Ultra cost?
Nemotron 3 Ultra lists at $0.600 per 1M input tokens and $2.40 per 1M output tokens, with cached input reads at $0.120 per 1M. A RAG assistant handling 100K requests a month (8K tokens in, 600 out) comes to about $432.00 at list price.
What is the context window of Nemotron 3 Ultra?
Nemotron 3 Ultra accepts up to 262,144 tokens of context and can return up to 182,520 output tokens per response.
Does Nemotron 3 Ultra support tool calling?
Yes — Nemotron 3 Ultra supports tool (function) calling and exposes extended reasoning. It accepts text input.
How does Nemotron 3 Ultra score on benchmarks?
On LiveBench release 2026-06-25 (run: nemotron-3-ultra-550b-a55b), Nemotron 3 Ultra scores 67.4 overall, ranking #49 of 55 benchmarked models.
What is the API model ID for Nemotron 3 Ultra?
On OpenRouter the model ID is "nvidia/nemotron-3-ultra-550b-a55b". It was listed on Jun 4, 2026.
More from nvidia
How these numbers are produced
- Price and specs — provider list data from OpenRouter, refreshed every 15 minutes. “Blended” is a 3:1 input:output mix.
- Scores and ranks — LiveBench release 2026-06-25. Ranks count only models that have a published run; a blank means “not evaluated”, never “bad”.
- Cost per point — the measured dollars LiveBench spent on the run, divided by the score it earned.
Published benchmarks rank models on someone else's tasks. Before committing, see LLM & agent evaluation for building an eval on your own.