// ranked · September 2026

Best LLM for Coding

Ranked by LiveBench’s coding category: code generation and completion on recent competitive-programming and LeetCode-style problems.

OpenRouter + LiveBenchAll rankingsAll models

Claude Opus 5.5 leads on LiveBench coding at 89.3, ahead of Claude Fable 5.1 (86.4) and Claude Fable 5 (86.0).

Claude Opus 5.5 vs Claude Fable 5.1, head to head

Top 10

#ModelCoding scoreBlended / 1MContextCapabilities
1Claude Opus 5.5

anthropic/claude-opus-5.5

89.3$8.001M
reasoningtoolsvision
2Claude Fable 5.1

anthropic/claude-fable-5.1

86.4$20.001M
reasoningtoolsvision
3Claude Fable 5

anthropic/claude-fable-5

86.0$20.001M
reasoningtoolsvision
4GPT-5.6 Sol

openai/gpt-5.6-sol

83.9$4.001.1M
reasoningtoolsvision
5GPT-5.2-Codex

openai/gpt-5.2-codex

83.6$4.81400K
reasoningtoolsvision
6GPT-5.6 Luna

openai/gpt-5.6-luna

82.9$0.4501.1M
reasoningtoolsvision
7GPT-5.5

openai/gpt-5.5

82.1$11.251.1M
reasoningtoolsvision
8Claude Opus 4.7

anthropic/claude-opus-4.7

82.1$10.001M
reasoningtoolsvision
9Claude Opus 4.8

anthropic/claude-opus-4.8

81.8$10.001M
reasoningtoolsvision
10GPT-6 Sol

openai/gpt-6-sol

81.8$4.001.1M
reasoningtoolsvision

45 more models qualify — see the full leaderboard.

FAQ

What is the best LLM for Coding in September 2026?

Claude Opus 5.5 leads on LiveBench coding at 89.3, ahead of Claude Fable 5.1 (86.4) and Claude Fable 5 (86.0).

What is the best value in the top 10 for coding?

Claude Opus 4.7. Lowest measured cost per point among the leaders — $0.0204 per point for a score of 82.1.

What is the best under $1.00 / 1m for coding?

GPT-5.6 Luna. Highest score at a blended list price of $1.00 per 1M tokens or less — 82.9 at $0.450.

How is this ranking produced?

Ranked by LiveBench’s coding category: code generation and completion on recent competitive-programming and LeetCode-style problems.

Other rankings

How this ranking is produced

  • One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
  • Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
  • Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.

A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.