// ranked · September 2026
Best LLM for Reasoning
Ranked by LiveBench’s reasoning category: logic puzzles, deduction and spatial reasoning tasks written after most training cutoffs.
GPT-6 Astra (92.7) and Claude Opus 5.5 (92.2) are effectively tied on LiveBench reasoning — less than a point apart, which effort settings alone can move. Claude Fable 5.1 follows at 91.7.
GPT-6 Astra vs Claude Opus 5.5, head to headBest value in the top 10
SpaceXAI: Grok 4.6
Lowest measured cost per point among the leaders — $0.0485 per point for a score of 90.5.
Best under $1.00 / 1M
DeepSeek V4.1 Flash
Highest score at a blended list price of $1.00 per 1M tokens or less — 86.7 at $0.200.
Top 10
| # | Model | Reasoning score | Blended / 1M | Context | Capabilities |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra openai/gpt-6-astra | 92.7 | $20.00 | 1.1M | reasoningtoolsvision |
| 2 | Claude Opus 5.5 anthropic/claude-opus-5.5 | 92.2 | $8.00 | 1M | reasoningtoolsvision |
| 3 | Claude Fable 5.1 anthropic/claude-fable-5.1 | 91.7 | $20.00 | 1M | reasoningtoolsvision |
| 4 | GPT-5.6 Sol openai/gpt-5.6-sol | 91.7 | $4.00 | 1.1M | reasoningtoolsvision |
| 5 | Claude Opus 5 anthropic/claude-opus-5 | 91.2 | $10.00 | 1M | reasoningtoolsvision |
| 6 | Kimi K3 moonshotai/kimi-k3 | 90.7 | $6.00 | 1.0M | reasoningtoolsvision |
| 7 | GPT-5.6 Terra openai/gpt-5.6-terra | 90.6 | $4.50 | 1.1M | reasoningtoolsvision |
| 8 | SpaceXAI: Grok 4.6 x-ai/grok-4.6 | 90.5 | $3.00 | 500K | reasoningtoolsvision |
| 9 | Muse Spark 1.2 meta/muse-spark-1.2 | 90.0 | $2.00 | 1.0M | reasoningtoolsvision |
| 10 | Muse Spark 1.3 meta/muse-spark-1.3 | 89.7 | $2.00 | 1.0M | reasoningtoolsvision |
45 more models qualify — see the full leaderboard.
FAQ
What is the best LLM for Reasoning in September 2026?
GPT-6 Astra (92.7) and Claude Opus 5.5 (92.2) are effectively tied on LiveBench reasoning — less than a point apart, which effort settings alone can move. Claude Fable 5.1 follows at 91.7.
What is the best value in the top 10 for reasoning?
SpaceXAI: Grok 4.6. Lowest measured cost per point among the leaders — $0.0485 per point for a score of 90.5.
What is the best under $1.00 / 1m for reasoning?
DeepSeek V4.1 Flash. Highest score at a blended list price of $1.00 per 1M tokens or less — 86.7 at $0.200.
How is this ranking produced?
Ranked by LiveBench’s reasoning category: logic puzzles, deduction and spatial reasoning tasks written after most training cutoffs.
Other rankings
How this ranking is produced
- One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
- Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
- Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.
A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.