// live_model_index
LLM model comparison
Every other leaderboard tells you which model scores highest. This one also tells you what that score cost — joining LiveBench's published scores and measured spend to live pricing, so you can see dollars per point of capability. Refreshed every 15 minutes; new models appear on their own.
Models tracked
425
57 providers
New in 30 days
48
104 in 90 days
Cheapest blended
$0.022
Mistral: Mistral Nemo
Largest context
2M
SpaceXAI: Grok 4.20 Multi-Agent
Release cadence
New models per month, trailing 12 months
Cost vs context
Where each model sits on price and window size
Value frontier — Agentic coding
Benchmark score against the measured cost of earning it
Showing 425 of 425 models
| Capabilities | ||||||||
|---|---|---|---|---|---|---|---|---|
Anthropic: Claude Fable 5.1New anthropic/claude-fable-5.1 | 66.1 | $2.260 | Sep 1, 20267d ago | 1M | $10.00 | $50.00 | $20.00 | ReasoningToolsVision |
Claude Opus 5 anthropic/claude-opus-5 | 65.2 | $1.964 | Jul 24, 202646d ago | 1M | $5.00 | $25.00 | $10.00 | ReasoningToolsVision |
DeepSeek: DeepSeek V4 Flash Vision ExpNew deepseek/deepseek-v4-flash-vision-exp | 65.1 | $0.040 | Aug 21, 202618d ago | 1.0M | $0.220 | $0.660 | $0.330 | ReasoningToolsVision |
Meta: Muse Spark 1.3New meta/muse-spark-1.3 | 64.1 | $0.439 | Sep 2, 20266d ago | 1.0M | $1.25 | $4.25 | $2.00 | ReasoningToolsVision |
MoonshotAI: Kimi K3 moonshotai/kimi-k3 | 62.2 | $0.680 | Jul 16, 202654d ago | 1.0M | $3.00 | $15.00 | $6.00 | ReasoningToolsVision |
Anthropic: Claude Fable 5 anthropic/claude-fable-5 | 62.2 | $3.422 | Jun 9, 202691d ago | 1M | $10.00 | $50.00 | $20.00 | ReasoningToolsVision |
Qwen: Qwen3.8 27BNew qwen/qwen3.8-27b | 61.4 | $0.371 | Aug 14, 202625d ago | 1M | $0.420 | $3.00 | $1.06 | ReasoningToolsVision |
Z.ai: GLM 5.3New z-ai/glm-5.3 | 60.9 | $0.630 | Aug 18, 202620d ago | 1.3M | $1.40 | $4.40 | $2.15 | ReasoningTools |
Anthropic: Claude Sonnet 5 anthropic/claude-sonnet-5 | 59.4 | $0.864 | Jun 30, 202670d ago | 1M | $2.00 | $10.00 | $4.00 | ReasoningToolsVision |
Meta: Muse Spark 1.1 meta/muse-spark-1.1 | 58.5 | $0.710 | Jul 16, 202654d ago | 1.0M | $1.25 | $4.25 | $2.00 | ReasoningToolsVision |
Google: Gemini 3.7 FlashNew google/gemini-3.7-flash | 58.3 | $0.343 | Aug 13, 202626d ago | 1.0M | $0.750 | $3.75 | $1.50 | ReasoningToolsVision |
Meta: Muse Spark 1.2 meta/muse-spark-1.2 | 57.6 | $1.636 | Aug 5, 202634d ago | 1.0M | $1.25 | $4.25 | $2.00 | ReasoningToolsVision |
OpenAI: GPT-6 AstraNew openai/gpt-6-astra | 57.3 | $1.018 | Sep 4, 20264d ago | 1.1M | $10.00 | $50.00 | $20.00 | ReasoningToolsVision |
SpaceXAI: Grok 4.6New x-ai/grok-4.6 | 57.0 | $0.752 | Aug 12, 202627d ago | 500K | $2.00 | $6.00 | $3.00 | ReasoningToolsVision |
Z.ai: GLM 5.3 FlashNew z-ai/glm-5.3-flash | 56.8 | $0.047 | Aug 26, 202613d ago | 1.3M | $0.075 | $0.250 | $0.119 | ReasoningToolsVision |
Loading more… (15 of 425)
Rendering 15 of 425 rows — scroll the table to load more.
What this measures — and what it doesn't
Live and measured
- Release date — every model carries a publication timestamp, so new launches appear here automatically.
- Price — list price per million input and output tokens. “Blended” is a 3:1 input:output mix, the usual shape of real traffic.
- Context window and max output tokens.
- Capabilities — reasoning, tool calling, and image input, as declared by the provider.
- All seven LiveBench categories — reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following, from LiveBench release 2026-06-25, using their own category map. Overall is the mean of the seven, which reproduces their headline number exactly. Covers 50 models.
- $ / point — LiveBench also publishes the measured dollars spent running each task. Dividing that by the score gives cost per point of capability, which neither a leaderboard nor a pricing page shows on its own. Across scored models the spread is about 96×.
LiveBench scores a model once per effort setting; the table shows the best published run, and hovering a score names the exact run. They rank models under that setup, not under yours.
Not shown, and why
- A score for every model. LiveBench evaluates a curated set, not the whole catalogue — 50 of 425 models here carry a score. A blank means “not evaluated,” never “bad.”
- Guessed matches. LiveBench rows are per effort setting (“claude-opus-5-xhigh-effort”). We strip the effort suffix and require an exact match to a catalogue id — no fuzzy matching, because in testing it confidently mapped Opus 5 onto Opus 4.5's score. Unmatched rows are dropped rather than guessed.
- A single “best” number. Effort setting moves these scores several points. The table shows each model's strongest published run; hover a score to see which run it was.
- Latency and throughput. The catalogue exposes these fields but returns them empty without an authenticated key.
Price and capability are facts a provider publishes. Quality is a measurement someone has to run — treat the two differently.
For how to evaluate models on your own workload rather than a public leaderboard, see LLM & agent evaluation.
LLM model comparison FAQ
Common questions about comparing language models on score, price, and cost per point of capability. Figures refresh with the table above.
What does cost per point mean in an LLM comparison?
Cost per point is the measured dollar cost of running a benchmark divided by the score it earned — dollars per point of capability. A leaderboard tells you which model scores highest; a pricing page tells you what a model charges per token. Neither tells you what a point of capability actually costs, because a reasoning model can burn many times more tokens than its per-token price suggests. LiveBench publishes the measured spend for each run, so dividing that by the score gives a like-for-like value number that token pricing alone hides.
Which LLM gives the best value for money?
By cost per point of overall capability, DeepSeek: DeepSeek V4 Flash 0423 is currently the best value on this page — it scores 65.5 at about $0.0083 per point. The most expensive model per point, Anthropic: Claude Fable 5, costs roughly 96× more for each point it earns. Best value is not the same as best: the cheapest model per point is rarely the highest scorer, so pick the cheapest model that clears the score you actually need rather than the cheapest overall.
Is the most expensive LLM the best one?
No. Anthropic: Claude Fable 5.1 has the highest overall LiveBench score here at 83.4, but it ranks 49 of 50 on cost per point. Price and score are only loosely related — several models score within a few points of the leader for a fraction of the cost. The value frontier chart on this page plots exactly that: any model above the line is unreachable, and anything far below it is overpaying for the score it delivers.
What is the cheapest LLM per million tokens?
The cheapest model in the catalogue by blended price is Mistral: Mistral Nemo at $0.0218 per million tokens, using a 3:1 input:output mix. Blended pricing matters because headline input prices understate the real bill — output tokens typically cost three to six times more than input, and reasoning models emit far more of them. A low sticker price also does not imply low cost per point; a cheap model that needs more attempts can end up costing more per unit of useful work.
How do LiveBench scores compare to price?
They correlate weakly, which is the whole reason this page exists. Sorting by score gives you a ranking; sorting by cost per point usually gives a very different one. This page joins both — LiveBench scores across reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following, against list pricing and the measured spend behind each run — so you can sort by either and see how far apart they are.
How often is this LLM comparison updated?
The model catalogue refreshes every 15 minutes from OpenRouter, so new releases and price changes appear without a redeploy — 48 models were added in the last 30 days. Benchmark scores update when LiveBench publishes a new release; the current one is 2026-06-25, covering 50 of 425 models.
Why do some models have no benchmark score?
LiveBench evaluates a curated set rather than every model that exists — 50 of 425 models here carry a score. A blank means "not evaluated", never "bad". Scores are also matched to catalogue ids by exact match only, with the effort suffix stripped; fuzzy matching was rejected because in testing it confidently mapped a new model onto an older version's score. Unmatched rows are dropped rather than guessed.