// ranked · September 2026
Best LLM for Math
Ranked by LiveBench’s mathematics category: recent competition problems and proof-style questions.
Claude Opus 5.5 (97.1) and Claude Fable 5.1 (97.0) are effectively tied on LiveBench mathematics — less than a point apart, which effort settings alone can move. GPT-6 Astra follows at 96.8.
Claude Opus 5.5 vs Claude Fable 5.1, head to headBest value in the top 10
GPT-6 Sol
Lowest measured cost per point among the leaders — $0.0840 per point for a score of 96.4.
Best under $1.00 / 1M
DeepSeek V4 Pro 0813
Highest score at a blended list price of $1.00 per 1M tokens or less — 95.1 at $0.990.
Top 10
| # | Model | Mathematics score | Blended / 1M | Context | Capabilities |
|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 anthropic/claude-opus-5.5 | 97.1 | $8.00 | 1M | reasoningtoolsvision |
| 2 | Claude Fable 5.1 anthropic/claude-fable-5.1 | 97.0 | $20.00 | 1M | reasoningtoolsvision |
| 3 | GPT-6 Astra openai/gpt-6-astra | 96.8 | $20.00 | 1.1M | reasoningtoolsvision |
| 4 | GPT-6 Sol openai/gpt-6-sol | 96.4 | $4.00 | 1.1M | reasoningtoolsvision |
| 5 | GPT-5.6 Sol openai/gpt-5.6-sol | 96.2 | $4.00 | 1.1M | reasoningtoolsvision |
| 6 | Claude Fable 5 anthropic/claude-fable-5 | 96.0 | $20.00 | 1M | reasoningtoolsvision |
| 7 | Muse Spark 1.3 meta/muse-spark-1.3 | 95.9 | $2.00 | 1.0M | reasoningtoolsvision |
| 8 | GPT-5.5 openai/gpt-5.5 | 95.9 | $11.25 | 1.1M | reasoningtoolsvision |
| 9 | Claude Opus 5 anthropic/claude-opus-5 | 95.7 | $10.00 | 1M | reasoningtoolsvision |
| 10 | SpaceXAI: Grok 4.7 x-ai/grok-4.7 | 95.7 | $2.40 | 500K | reasoningtoolsvision |
45 more models qualify — see the full leaderboard.
FAQ
What is the best LLM for Math in September 2026?
Claude Opus 5.5 (97.1) and Claude Fable 5.1 (97.0) are effectively tied on LiveBench mathematics — less than a point apart, which effort settings alone can move. GPT-6 Astra follows at 96.8.
What is the best value in the top 10 for mathematics?
GPT-6 Sol. Lowest measured cost per point among the leaders — $0.0840 per point for a score of 96.4.
What is the best under $1.00 / 1m for mathematics?
DeepSeek V4 Pro 0813. Highest score at a blended list price of $1.00 per 1M tokens or less — 95.1 at $0.990.
How is this ranking produced?
Ranked by LiveBench’s mathematics category: recent competition problems and proof-style questions.
Other rankings
How this ranking is produced
- One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
- Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
- Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.
A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.