// ranked · September 2026

Best LLM for Reasoning

Ranked by LiveBench’s reasoning category: logic puzzles, deduction and spatial reasoning tasks written after most training cutoffs.

OpenRouter + LiveBenchAll rankingsAll models

GPT-6 Astra (92.7) and Claude Opus 5.5 (92.2) are effectively tied on LiveBench reasoning — less than a point apart, which effort settings alone can move. Claude Fable 5.1 follows at 91.7.

GPT-6 Astra vs Claude Opus 5.5, head to head

Top 10

#ModelReasoning scoreBlended / 1MContextCapabilities
1GPT-6 Astra

openai/gpt-6-astra

92.7$20.001.1M
reasoningtoolsvision
2Claude Opus 5.5

anthropic/claude-opus-5.5

92.2$8.001M
reasoningtoolsvision
3Claude Fable 5.1

anthropic/claude-fable-5.1

91.7$20.001M
reasoningtoolsvision
4GPT-5.6 Sol

openai/gpt-5.6-sol

91.7$4.001.1M
reasoningtoolsvision
5Claude Opus 5

anthropic/claude-opus-5

91.2$10.001M
reasoningtoolsvision
6Kimi K3

moonshotai/kimi-k3

90.7$6.001.0M
reasoningtoolsvision
7GPT-5.6 Terra

openai/gpt-5.6-terra

90.6$4.501.1M
reasoningtoolsvision
8SpaceXAI: Grok 4.6

x-ai/grok-4.6

90.5$3.00500K
reasoningtoolsvision
9Muse Spark 1.2

meta/muse-spark-1.2

90.0$2.001.0M
reasoningtoolsvision
10Muse Spark 1.3

meta/muse-spark-1.3

89.7$2.001.0M
reasoningtoolsvision

45 more models qualify — see the full leaderboard.

FAQ

What is the best LLM for Reasoning in September 2026?

GPT-6 Astra (92.7) and Claude Opus 5.5 (92.2) are effectively tied on LiveBench reasoning — less than a point apart, which effort settings alone can move. Claude Fable 5.1 follows at 91.7.

What is the best value in the top 10 for reasoning?

SpaceXAI: Grok 4.6. Lowest measured cost per point among the leaders — $0.0485 per point for a score of 90.5.

What is the best under $1.00 / 1m for reasoning?

DeepSeek V4.1 Flash. Highest score at a blended list price of $1.00 per 1M tokens or less — 86.7 at $0.200.

How is this ranking produced?

Ranked by LiveBench’s reasoning category: logic puzzles, deduction and spatial reasoning tasks written after most training cutoffs.

Other rankings

How this ranking is produced

  • One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
  • Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
  • Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.

A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.