// ranked · September 2026

Best LLM Overall

Ranked by LiveBench overall score — the mean of all seven category scores, on contamination-limited questions refreshed every release.

OpenRouter + LiveBenchAll rankingsAll models

Claude Fable 5.1 (83.4) and Claude Opus 5.5 (83.2) are effectively tied on LiveBench overall — less than a point apart, which effort settings alone can move. Claude Fable 5 follows at 83.0.

Claude Fable 5.1 vs Claude Opus 5.5, head to head

Top 10

#ModelOverall scoreBlended / 1MContextCapabilities
1Claude Fable 5.1

anthropic/claude-fable-5.1

83.4$20.001M
reasoningtoolsvision
2Claude Opus 5.5

anthropic/claude-opus-5.5

83.2$8.001M
reasoningtoolsvision
3Claude Fable 5

anthropic/claude-fable-5

83.0$20.001M
reasoningtoolsvision
4GPT-6 Astra

openai/gpt-6-astra

82.2$20.001.1M
reasoningtoolsvision
5Muse Spark 1.3

meta/muse-spark-1.3

81.6$2.001.0M
reasoningtoolsvision
6DeepSeek V4.1 Flash

deepseek/deepseek-v4.1-flash

81.1$0.2001.0M
reasoningtoolsvision
7GPT-5.6 Sol

openai/gpt-5.6-sol

81.1$4.001.1M
reasoningtoolsvision
8GPT-5.5

openai/gpt-5.5

80.2$11.251.1M
reasoningtoolsvision
9Claude Opus 5

anthropic/claude-opus-5

80.1$10.001M
reasoningtoolsvision
10GPT-6 Sol

openai/gpt-6-sol

79.2$4.001.1M
reasoningtoolsvision

45 more models qualify — see the full leaderboard.

FAQ

What is the best LLM Overall in September 2026?

Claude Fable 5.1 (83.4) and Claude Opus 5.5 (83.2) are effectively tied on LiveBench overall — less than a point apart, which effort settings alone can move. Claude Fable 5 follows at 83.0.

What is the best value in the top 10 for overall?

DeepSeek V4.1 Flash. Lowest measured cost per point among the leaders — $0.0157 per point for a score of 81.1.

What is the best under $1.00 / 1m for overall?

DeepSeek V4.1 Flash. Highest score at a blended list price of $1.00 per 1M tokens or less — 81.1 at $0.200.

How is this ranking produced?

Ranked by LiveBench overall score — the mean of all seven category scores, on contamination-limited questions refreshed every release.

Other rankings

How this ranking is produced

  • One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
  • Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
  • Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.

A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.