// ranked · September 2026

Best LLM for Math

Ranked by LiveBench’s mathematics category: recent competition problems and proof-style questions.

OpenRouter + LiveBenchAll rankingsAll models

Claude Opus 5.5 (97.1) and Claude Fable 5.1 (97.0) are effectively tied on LiveBench mathematics — less than a point apart, which effort settings alone can move. GPT-6 Astra follows at 96.8.

Claude Opus 5.5 vs Claude Fable 5.1, head to head

Top 10

#ModelMathematics scoreBlended / 1MContextCapabilities
1Claude Opus 5.5

anthropic/claude-opus-5.5

97.1$8.001M
reasoningtoolsvision
2Claude Fable 5.1

anthropic/claude-fable-5.1

97.0$20.001M
reasoningtoolsvision
3GPT-6 Astra

openai/gpt-6-astra

96.8$20.001.1M
reasoningtoolsvision
4GPT-6 Sol

openai/gpt-6-sol

96.4$4.001.1M
reasoningtoolsvision
5GPT-5.6 Sol

openai/gpt-5.6-sol

96.2$4.001.1M
reasoningtoolsvision
6Claude Fable 5

anthropic/claude-fable-5

96.0$20.001M
reasoningtoolsvision
7Muse Spark 1.3

meta/muse-spark-1.3

95.9$2.001.0M
reasoningtoolsvision
8GPT-5.5

openai/gpt-5.5

95.9$11.251.1M
reasoningtoolsvision
9Claude Opus 5

anthropic/claude-opus-5

95.7$10.001M
reasoningtoolsvision
10SpaceXAI: Grok 4.7

x-ai/grok-4.7

95.7$2.40500K
reasoningtoolsvision

45 more models qualify — see the full leaderboard.

FAQ

What is the best LLM for Math in September 2026?

Claude Opus 5.5 (97.1) and Claude Fable 5.1 (97.0) are effectively tied on LiveBench mathematics — less than a point apart, which effort settings alone can move. GPT-6 Astra follows at 96.8.

What is the best value in the top 10 for mathematics?

GPT-6 Sol. Lowest measured cost per point among the leaders — $0.0840 per point for a score of 96.4.

What is the best under $1.00 / 1m for mathematics?

DeepSeek V4 Pro 0813. Highest score at a blended list price of $1.00 per 1M tokens or less — 95.1 at $0.990.

How is this ranking produced?

Ranked by LiveBench’s mathematics category: recent competition problems and proof-style questions.

Other rankings

How this ranking is produced

  • One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
  • Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
  • Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.

A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.