// nvidia

Nemotron 3 Ultra

nvidia/nemotron-3-ultra-550b-a55b

reasoningtool callingprompt caching
OpenRouter + LiveBenchAll modelsFull leaderboard

Nemotron 3 Ultra costs $0.600 per 1M input tokens and $2.40 per 1M output tokens with a 262K-token context window.

A reasoning, tool-calling model from nvidia, released Jun 4, 2026. On LiveBench it scores 67.4 overall — #49 of 55 benchmarked models — and ranks highest in instruction following (#12).

Input / 1M
$0.600
Output / 1M
$2.40
Context
262K
LiveBench overall
67.4 · #49

Specs and pricing

Input price

USD per 1M prompt tokens

$0.600
Output price

USD per 1M completion tokens, reasoning included

$2.40
Blended price

3:1 input:output — #132 most expensive of 343 listed models

$1.05
Cached input read

Blank means not priced separately — not free

$0.120
Cache write
Context window262,144 tokens
Max output182,520 tokens
Input modalitiestext
Tool callingYes
Extended reasoningYes
Knowledge cutoff
ListedJun 4, 2026
All NVIDIA API pricing

Also listed as nvidia/nemotron-3-ultra-550b-a55b:free (Free blended)— the same model at a different price or rate limit.

Benchmarks by category

LiveBench scores out of 100, with Nemotron 3 Ultra's rank among the 55 benchmarked models listed here. Run: nemotron-3-ultra-550b-a55b.

Overall#49 of 55
67.4

$0.2118 measured per point

38.7

$1.77 measured per point

Coding#48 of 55
70.7

$0.0158 measured per point

Reasoning#49 of 55
74.7

$0.0829 measured per point

Mathematics#32 of 55
88.7

$0.1018 measured per point

Data analysis#53 of 55
54.5

$0.1500 measured per point

Language#52 of 55
70.8

$0.0483 measured per point

73.4

$0.0453 measured per point

What Nemotron 3 Ultra costs to run

Monthly list cost across five workload shapes, using the published cached-input rate where there is one. A floor, not a quote — batch discounts and cache writes aren't included.

WorkloadPer month
Support chatbot

1.2K in / 400 out × 200K requests

$301.44/mo
RAG assistant

8K in / 600 out × 100K requests

$432.00/mo
Coding agent

40K in / 4K out × 20K requests

$403.20/mo
Document extraction

20K in / 1.5K out × 50K requests

$756.00/mo
Bulk classification

500 in / 20 out × 5M requests

$1,500/mo
Price your own workload

Compare Nemotron 3 Ultra with

The benchmarked models closest to it on overall score.

Nemotron 3 Ultra FAQ

How much does Nemotron 3 Ultra cost?

Nemotron 3 Ultra lists at $0.600 per 1M input tokens and $2.40 per 1M output tokens, with cached input reads at $0.120 per 1M. A RAG assistant handling 100K requests a month (8K tokens in, 600 out) comes to about $432.00 at list price.

What is the context window of Nemotron 3 Ultra?

Nemotron 3 Ultra accepts up to 262,144 tokens of context and can return up to 182,520 output tokens per response.

Does Nemotron 3 Ultra support tool calling?

Yes — Nemotron 3 Ultra supports tool (function) calling and exposes extended reasoning. It accepts text input.

How does Nemotron 3 Ultra score on benchmarks?

On LiveBench release 2026-06-25 (run: nemotron-3-ultra-550b-a55b), Nemotron 3 Ultra scores 67.4 overall, ranking #49 of 55 benchmarked models.

What is the API model ID for Nemotron 3 Ultra?

On OpenRouter the model ID is "nvidia/nemotron-3-ultra-550b-a55b". It was listed on Jun 4, 2026.

More from nvidia

How these numbers are produced

  • Price and specs — provider list data from OpenRouter, refreshed every 15 minutes. “Blended” is a 3:1 input:output mix.
  • Scores and ranksLiveBench release 2026-06-25. Ranks count only models that have a published run; a blank means “not evaluated”, never “bad”.
  • Cost per point — the measured dollars LiveBench spent on the run, divided by the score it earned.

Published benchmarks rank models on someone else's tasks. Before committing, see LLM & agent evaluation for building an eval on your own.