// ranked · September 2026
Best LLM for Agentic Coding
Ranked by LiveBench’s agentic coding category: multi-step tasks where the model works in a real repository with tools, closer to how coding agents are actually used.
DeepSeek V4.1 Flash leads on LiveBench agentic coding at 77.3, ahead of Claude Opus 5.5 (71.7) and Claude Fable 5.1 (66.1).
DeepSeek V4.1 Flash vs Claude Opus 5.5, head to headTop 10
| # | Model | Agentic coding score | Blended / 1M | Context | Capabilities |
|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 Flash deepseek/deepseek-v4.1-flash | 77.3 | $0.200 | 1.0M | reasoningtoolsvision |
| 2 | Claude Opus 5.5 anthropic/claude-opus-5.5 | 71.7 | $8.00 | 1M | reasoningtoolsvision |
| 3 | Claude Fable 5.1 anthropic/claude-fable-5.1 | 66.1 | $20.00 | 1M | reasoningtoolsvision |
| 4 | Claude Opus 5 anthropic/claude-opus-5 | 65.2 | $10.00 | 1M | reasoningtoolsvision |
| 5 | DeepSeek V4 Flash Vision Exp deepseek/deepseek-v4-flash-vision-exp | 65.1 | $0.330 | 1.0M | reasoningtoolsvision |
| 6 | Muse Spark 1.3 meta/muse-spark-1.3 | 64.1 | $2.00 | 1.0M | reasoningtoolsvision |
| 7 | Kimi K3 moonshotai/kimi-k3 | 62.2 | $6.00 | 1.0M | reasoningtoolsvision |
| 8 | Claude Fable 5 anthropic/claude-fable-5 | 62.2 | $20.00 | 1M | reasoningtoolsvision |
| 9 | Qwen3.8 27B qwen/qwen3.8-27b | 61.4 | $1.06 | 1M | reasoningtoolsvision |
| 10 | GLM 5.3 z-ai/glm-5.3 | 60.9 | $0.862 | 1.3M | reasoningtools |
45 more models qualify — see the full leaderboard.
FAQ
What is the best LLM for Agentic Coding in September 2026?
DeepSeek V4.1 Flash leads on LiveBench agentic coding at 77.3, ahead of Claude Opus 5.5 (71.7) and Claude Fable 5.1 (66.1).
How is this ranking produced?
Ranked by LiveBench’s agentic coding category: multi-step tasks where the model works in a real repository with tools, closer to how coding agents are actually used.
Other rankings
How this ranking is produced
- One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
- Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
- Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.
A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.