// live_model_index
LLM model comparison
Every other leaderboard tells you which model scores highest. This one also tells you what that score cost — joining LiveBench's published scores and measured spend to live pricing, so you can see dollars per point of capability. Refreshed every 15 minutes; new models appear on their own.
Models tracked
339
57 providers
New in 30 days
28
72 in 90 days
Cheapest blended
$0.015
inclusionAI: Ling-2.6-flash
Largest context
2M
xAI: Grok 4.20 Multi-Agent
Release cadence
New models per month, trailing 12 months
Cost vs context
Where each model sits on price and window size
Value frontier — Agentic coding
Benchmark score against the measured cost of earning it
Showing 339 of 339 models
| Capabilities | ||||||||
|---|---|---|---|---|---|---|---|---|
OpenAI: GPT-5.6 TerraNew openai/gpt-5.6-terra | 68.0 | $1.240 | Jul 9, 202616d ago | 1.1M | $2.50 | $15.00 | $5.63 | ReasoningToolsVision |
OpenAI: GPT-5.6 SolNew openai/gpt-5.6-sol | 65.6 | $1.678 | Jul 9, 202616d ago | 1.1M | $5.00 | $30.00 | $11.25 | ReasoningToolsVision |
Meta: Muse Spark 1.1New meta/muse-spark-1.1 | 65.1 | $0.830 | Jul 16, 20269d ago | 1.0M | $1.25 | $4.25 | $2.00 | ReasoningToolsVision |
Claude Opus 5New anthropic/claude-opus-5 | 61.3 | $1.343 | Jul 24, 20261d ago | 1M | $5.00 | $25.00 | $10.00 | ReasoningToolsVision |
xAI: Grok 4.5New x-ai/grok-4.5 | 59.8 | $0.259 | Jul 8, 202617d ago | 500K | $2.00 | $6.00 | $3.00 | ReasoningToolsVision |
MoonshotAI: Kimi K3New moonshotai/kimi-k3 | 57.6 | $0.896 | Jul 16, 20269d ago | 1.0M | $3.00 | $15.00 | $6.00 | ReasoningToolsVision |
Anthropic: Claude Opus 4.8 anthropic/claude-opus-4.8 | 56.1 | $1.608 | May 27, 202659d ago | 1M | $5.00 | $25.00 | $10.00 | ReasoningToolsVision |
OpenAI: GPT-5.6 LunaNew openai/gpt-5.6-luna | 53.8 | $0.421 | Jul 9, 202616d ago | 1.1M | $1.00 | $6.00 | $2.25 | ReasoningToolsVision |
OpenAI: GPT-5.4 openai/gpt-5.4 | 53.8 | $0.771 | Mar 5, 2026142d ago | 1.1M | $2.50 | $15.00 | $5.63 | ReasoningToolsVision |
OpenAI: GPT-5.5 openai/gpt-5.5 | 52.1 | $1.617 | Apr 24, 202692d ago | 1.1M | $5.00 | $30.00 | $11.25 | ReasoningToolsVision |
Z.ai: GLM 5.2 z-ai/glm-5.2 | 51.9 | $0.499 | Jun 16, 202639d ago | 1.0M | $0.762 | $2.39 | $1.17 | ReasoningTools |
Anthropic: Claude Sonnet 5New anthropic/claude-sonnet-5 | 51.1 | $0.872 | Jun 30, 202625d ago | 1M | $2.00 | $10.00 | $4.00 | ReasoningToolsVision |
Anthropic: Claude Opus 4.7 anthropic/claude-opus-4.7 | 50.7 | $1.250 | Apr 16, 2026100d ago | 1M | $5.00 | $25.00 | $10.00 | ReasoningToolsVision |
OpenAI: GPT-5.2 openai/gpt-5.2 | 50.3 | $0.679 | Dec 10, 2025227d ago | 400K | $1.75 | $14.00 | $4.81 | ReasoningToolsVision |
OpenAI: GPT-5.2-Codex openai/gpt-5.2-codex | 49.4 | $0.289 | Jan 14, 2026192d ago | 400K | $1.75 | $14.00 | $4.81 | ReasoningToolsVision |
Loading more… (15 of 339)
Rendering 15 of 339 rows — scroll the table to load more.
What this measures — and what it doesn't
Live and measured
- Release date — every model carries a publication timestamp, so new launches appear here automatically.
- Price — list price per million input and output tokens. “Blended” is a 3:1 input:output mix, the usual shape of real traffic.
- Context window and max output tokens.
- Capabilities — reasoning, tool calling, and image input, as declared by the provider.
- All seven LiveBench categories — reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following, from LiveBench release 2026-06-25, using their own category map. Overall is the mean of the seven, which reproduces their headline number exactly. Covers 36 models.
- $ / point — LiveBench also publishes the measured dollars spent running each task. Dividing that by the score gives cost per point of capability, which neither a leaderboard nor a pricing page shows on its own. Across scored models the spread is about 105×.
LiveBench scores a model once per effort setting; the table shows the best published run, and hovering a score names the exact run. They rank models under that setup, not under yours.
Not shown, and why
- A score for every model. LiveBench evaluates a curated set, not the whole catalogue — 36 of 339 models here carry a score. A blank means “not evaluated,” never “bad.”
- Guessed matches. LiveBench rows are per effort setting (“claude-opus-5-xhigh-effort”). We strip the effort suffix and require an exact match to a catalogue id — no fuzzy matching, because in testing it confidently mapped Opus 5 onto Opus 4.5's score. Unmatched rows are dropped rather than guessed.
- A single “best” number. Effort setting moves these scores several points. The table shows each model's strongest published run; hover a score to see which run it was.
- Latency and throughput. The catalogue exposes these fields but returns them empty without an authenticated key.
Price and capability are facts a provider publishes. Quality is a measurement someone has to run — treat the two differently.
For how to evaluate models on your own workload rather than a public leaderboard, see LLM & agent evaluation.
LLM model comparison FAQ
Common questions about comparing language models on score, price, and cost per point of capability. Figures refresh with the table above.
What does cost per point mean in an LLM comparison?
Cost per point is the measured dollar cost of running a benchmark divided by the score it earned — dollars per point of capability. A leaderboard tells you which model scores highest; a pricing page tells you what a model charges per token. Neither tells you what a point of capability actually costs, because a reasoning model can burn many times more tokens than its per-token price suggests. LiveBench publishes the measured spend for each run, so dividing that by the score gives a like-for-like value number that token pricing alone hides.
Which LLM gives the best value for money?
By cost per point of overall capability, DeepSeek: DeepSeek V4 Flash is currently the best value on this page — it scores 65.5 at about $0.0083 per point. The most expensive model per point, Anthropic: Claude Fable 5, costs roughly 105× more for each point it earns. Best value is not the same as best: the cheapest model per point is rarely the highest scorer, so pick the cheapest model that clears the score you actually need rather than the cheapest overall.
Is the most expensive LLM the best one?
No. OpenAI: GPT-5.6 Sol has the highest overall LiveBench score here at 82.4, but it ranks 34 of 36 on cost per point. Price and score are only loosely related — several models score within a few points of the leader for a fraction of the cost. The value frontier chart on this page plots exactly that: any model above the line is unreachable, and anything far below it is overpaying for the score it delivers.
What is the cheapest LLM per million tokens?
The cheapest model in the catalogue by blended price is inclusionAI: Ling-2.6-flash at $0.0150 per million tokens, using a 3:1 input:output mix. Blended pricing matters because headline input prices understate the real bill — output tokens typically cost three to six times more than input, and reasoning models emit far more of them. A low sticker price also does not imply low cost per point; a cheap model that needs more attempts can end up costing more per unit of useful work.
How do LiveBench scores compare to price?
They correlate weakly, which is the whole reason this page exists. Sorting by score gives you a ranking; sorting by cost per point usually gives a very different one. This page joins both — LiveBench scores across reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following, against list pricing and the measured spend behind each run — so you can sort by either and see how far apart they are.
How often is this LLM comparison updated?
The model catalogue refreshes every 15 minutes from OpenRouter, so new releases and price changes appear without a redeploy — 28 models were added in the last 30 days. Benchmark scores update when LiveBench publishes a new release; the current one is 2026-06-25, covering 36 of 339 models.
Why do some models have no benchmark score?
LiveBench evaluates a curated set rather than every model that exists — 36 of 339 models here carry a score. A blank means "not evaluated", never "bad". Scores are also matched to catalogue ids by exact match only, with the effort suffix stripped; fuzzy matching was rejected because in testing it confidently mapped a new model onto an older version's score. Unmatched rows are dropped rather than guessed.