// ranked · September 2026

Best LLM for Agentic Coding

Ranked by LiveBench’s agentic coding category: multi-step tasks where the model works in a real repository with tools, closer to how coding agents are actually used.

OpenRouter + LiveBenchAll rankingsAll models

DeepSeek V4.1 Flash leads on LiveBench agentic coding at 77.3, ahead of Claude Opus 5.5 (71.7) and Claude Fable 5.1 (66.1).

DeepSeek V4.1 Flash vs Claude Opus 5.5, head to head

Top 10

#ModelAgentic coding scoreBlended / 1MContextCapabilities
1DeepSeek V4.1 Flash

deepseek/deepseek-v4.1-flash

77.3$0.2001.0M
reasoningtoolsvision
2Claude Opus 5.5

anthropic/claude-opus-5.5

71.7$8.001M
reasoningtoolsvision
3Claude Fable 5.1

anthropic/claude-fable-5.1

66.1$20.001M
reasoningtoolsvision
4Claude Opus 5

anthropic/claude-opus-5

65.2$10.001M
reasoningtoolsvision
5DeepSeek V4 Flash Vision Exp

deepseek/deepseek-v4-flash-vision-exp

65.1$0.3301.0M
reasoningtoolsvision
6Muse Spark 1.3

meta/muse-spark-1.3

64.1$2.001.0M
reasoningtoolsvision
7Kimi K3

moonshotai/kimi-k3

62.2$6.001.0M
reasoningtoolsvision
8Claude Fable 5

anthropic/claude-fable-5

62.2$20.001M
reasoningtoolsvision
9Qwen3.8 27B

qwen/qwen3.8-27b

61.4$1.061M
reasoningtoolsvision
10GLM 5.3

z-ai/glm-5.3

60.9$0.8621.3M
reasoningtools

45 more models qualify — see the full leaderboard.

FAQ

What is the best LLM for Agentic Coding in September 2026?

DeepSeek V4.1 Flash leads on LiveBench agentic coding at 77.3, ahead of Claude Opus 5.5 (71.7) and Claude Fable 5.1 (66.1).

How is this ranking produced?

Ranked by LiveBench’s agentic coding category: multi-step tasks where the model works in a real repository with tools, closer to how coding agents are actually used.

Other rankings

How this ranking is produced

  • One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
  • Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
  • Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.

A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.