// live_model_index

LLM model comparison

Every other leaderboard tells you which model scores highest. This one also tells you what that score cost — joining LiveBench's published scores and measured spend to live pricing, so you can see dollars per point of capability. Refreshed every 15 minutes; new models appear on their own.

LiveSource: OpenRouterUpdated Sep 8, 2026, 8:44 PM UTCCompare two modelsCost calculatorBack to topics

Models tracked

425

57 providers

New in 30 days

48

104 in 90 days

Cheapest blended

$0.022

Mistral: Mistral Nemo

Largest context

2M

SpaceXAI: Grok 4.20 Multi-Agent

Release cadence

New models per month, trailing 12 months

02142OctNovDecJanFebMarAprMayJunJulAugSep
Models added to the catalogue per month. Single series — colour carries no extra meaning.

Cost vs context

Where each model sits on price and window size

8K32K128K1M$0.100$1.00$10.00$100.00Blended cost per 1M tokens (log)
Each dot is a model. Cheaper is left, larger context is up — so the upper-left corner is the value frontier.
Category

Value frontier — Agentic coding

Benchmark score against the measured cost of earning it

20406070$1$10$100DeepSeek: DeepSeek V4…DeepSeek: DeepSeek V4…Claude Opus 5Measured cost to run the benchmark, USD (log)
Highlighted models sit on the value frontier — nothing is both cheaper and better. Anything above the dashed line is unreachable; anything far below it is overpaying.

Showing 425 of 425 models

Live comparison of language models by release date, context window, and price
Capabilities
Anthropic: Claude Fable 5.1New
anthropic/claude-fable-5.1
66.1$2.260Sep 1, 20267d ago1M$10.00$50.00$20.00
ReasoningToolsVision
Claude Opus 5
anthropic/claude-opus-5
65.2$1.964Jul 24, 202646d ago1M$5.00$25.00$10.00
ReasoningToolsVision
DeepSeek: DeepSeek V4 Flash Vision ExpNew
deepseek/deepseek-v4-flash-vision-exp
65.1$0.040Aug 21, 202618d ago1.0M$0.220$0.660$0.330
ReasoningToolsVision
Meta: Muse Spark 1.3New
meta/muse-spark-1.3
64.1$0.439Sep 2, 20266d ago1.0M$1.25$4.25$2.00
ReasoningToolsVision
MoonshotAI: Kimi K3
moonshotai/kimi-k3
62.2$0.680Jul 16, 202654d ago1.0M$3.00$15.00$6.00
ReasoningToolsVision
Anthropic: Claude Fable 5
anthropic/claude-fable-5
62.2$3.422Jun 9, 202691d ago1M$10.00$50.00$20.00
ReasoningToolsVision
Qwen: Qwen3.8 27BNew
qwen/qwen3.8-27b
61.4$0.371Aug 14, 202625d ago1M$0.420$3.00$1.06
ReasoningToolsVision
Z.ai: GLM 5.3New
z-ai/glm-5.3
60.9$0.630Aug 18, 202620d ago1.3M$1.40$4.40$2.15
ReasoningTools
Anthropic: Claude Sonnet 5
anthropic/claude-sonnet-5
59.4$0.864Jun 30, 202670d ago1M$2.00$10.00$4.00
ReasoningToolsVision
Meta: Muse Spark 1.1
meta/muse-spark-1.1
58.5$0.710Jul 16, 202654d ago1.0M$1.25$4.25$2.00
ReasoningToolsVision
Google: Gemini 3.7 FlashNew
google/gemini-3.7-flash
58.3$0.343Aug 13, 202626d ago1.0M$0.750$3.75$1.50
ReasoningToolsVision
Meta: Muse Spark 1.2
meta/muse-spark-1.2
57.6$1.636Aug 5, 202634d ago1.0M$1.25$4.25$2.00
ReasoningToolsVision
OpenAI: GPT-6 AstraNew
openai/gpt-6-astra
57.3$1.018Sep 4, 20264d ago1.1M$10.00$50.00$20.00
ReasoningToolsVision
SpaceXAI: Grok 4.6New
x-ai/grok-4.6
57.0$0.752Aug 12, 202627d ago500K$2.00$6.00$3.00
ReasoningToolsVision
Z.ai: GLM 5.3 FlashNew
z-ai/glm-5.3-flash
56.8$0.047Aug 26, 202613d ago1.3M$0.075$0.250$0.119
ReasoningToolsVision

Loading more… (15 of 425)

Rendering 15 of 425 rows — scroll the table to load more.

What this measures — and what it doesn't

Live and measured

  • Release date — every model carries a publication timestamp, so new launches appear here automatically.
  • Price — list price per million input and output tokens. “Blended” is a 3:1 input:output mix, the usual shape of real traffic.
  • Context window and max output tokens.
  • Capabilities — reasoning, tool calling, and image input, as declared by the provider.
  • All seven LiveBench categories — reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following, from LiveBench release 2026-06-25, using their own category map. Overall is the mean of the seven, which reproduces their headline number exactly. Covers 50 models.
  • $ / point — LiveBench also publishes the measured dollars spent running each task. Dividing that by the score gives cost per point of capability, which neither a leaderboard nor a pricing page shows on its own. Across scored models the spread is about 96×.

LiveBench scores a model once per effort setting; the table shows the best published run, and hovering a score names the exact run. They rank models under that setup, not under yours.

Not shown, and why

  • A score for every model. LiveBench evaluates a curated set, not the whole catalogue — 50 of 425 models here carry a score. A blank means “not evaluated,” never “bad.”
  • Guessed matches. LiveBench rows are per effort setting (“claude-opus-5-xhigh-effort”). We strip the effort suffix and require an exact match to a catalogue id — no fuzzy matching, because in testing it confidently mapped Opus 5 onto Opus 4.5's score. Unmatched rows are dropped rather than guessed.
  • A single “best” number. Effort setting moves these scores several points. The table shows each model's strongest published run; hover a score to see which run it was.
  • Latency and throughput. The catalogue exposes these fields but returns them empty without an authenticated key.

Price and capability are facts a provider publishes. Quality is a measurement someone has to run — treat the two differently.

For how to evaluate models on your own workload rather than a public leaderboard, see LLM & agent evaluation.

LLM model comparison FAQ

Common questions about comparing language models on score, price, and cost per point of capability. Figures refresh with the table above.

What does cost per point mean in an LLM comparison?

Cost per point is the measured dollar cost of running a benchmark divided by the score it earned — dollars per point of capability. A leaderboard tells you which model scores highest; a pricing page tells you what a model charges per token. Neither tells you what a point of capability actually costs, because a reasoning model can burn many times more tokens than its per-token price suggests. LiveBench publishes the measured spend for each run, so dividing that by the score gives a like-for-like value number that token pricing alone hides.

Which LLM gives the best value for money?

By cost per point of overall capability, DeepSeek: DeepSeek V4 Flash 0423 is currently the best value on this page — it scores 65.5 at about $0.0083 per point. The most expensive model per point, Anthropic: Claude Fable 5, costs roughly 96× more for each point it earns. Best value is not the same as best: the cheapest model per point is rarely the highest scorer, so pick the cheapest model that clears the score you actually need rather than the cheapest overall.

Is the most expensive LLM the best one?

No. Anthropic: Claude Fable 5.1 has the highest overall LiveBench score here at 83.4, but it ranks 49 of 50 on cost per point. Price and score are only loosely related — several models score within a few points of the leader for a fraction of the cost. The value frontier chart on this page plots exactly that: any model above the line is unreachable, and anything far below it is overpaying for the score it delivers.

What is the cheapest LLM per million tokens?

The cheapest model in the catalogue by blended price is Mistral: Mistral Nemo at $0.0218 per million tokens, using a 3:1 input:output mix. Blended pricing matters because headline input prices understate the real bill — output tokens typically cost three to six times more than input, and reasoning models emit far more of them. A low sticker price also does not imply low cost per point; a cheap model that needs more attempts can end up costing more per unit of useful work.

How do LiveBench scores compare to price?

They correlate weakly, which is the whole reason this page exists. Sorting by score gives you a ranking; sorting by cost per point usually gives a very different one. This page joins both — LiveBench scores across reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following, against list pricing and the measured spend behind each run — so you can sort by either and see how far apart they are.

How often is this LLM comparison updated?

The model catalogue refreshes every 15 minutes from OpenRouter, so new releases and price changes appear without a redeploy — 48 models were added in the last 30 days. Benchmark scores update when LiveBench publishes a new release; the current one is 2026-06-25, covering 50 of 425 models.

Why do some models have no benchmark score?

LiveBench evaluates a curated set rather than every model that exists — 50 of 425 models here carry a score. A blank means "not evaluated", never "bad". Scores are also matched to catalogue ids by exact match only, with the effort suffix stripped; fuzzy matching was rejected because in testing it confidently mapped a new model onto an older version's score. Unmatched rows are dropped rather than guessed.