// ranked · September 2026
Best LLM for Instruction Following
Ranked by LiveBench’s instruction-following category: producing output under explicit format and content constraints.
Gemini 3.8 Flash leads on LiveBench instruction following at 81.4, ahead of Gemini 3.7 Flash (79.9) and Gemini 3.1 Pro Preview (79.1).
Gemini 3.8 Flash vs Gemini 3.7 Flash, head to headBest value in the top 10
Muse Spark 1.2
Lowest measured cost per point among the leaders — $0.0485 per point for a score of 74.3.
Best under $1.00 / 1M
DeepSeek V4 Flash Vision Exp
Highest score at a blended list price of $1.00 per 1M tokens or less — 71.0 at $0.330.
Top 10
| # | Model | Instruction following score | Blended / 1M | Context | Capabilities |
|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash google/gemini-3.8-flash | 81.4 | $1.50 | 1.0M | reasoningtoolsvision |
| 2 | Gemini 3.7 Flash google/gemini-3.7-flash | 79.9 | $1.50 | 1.0M | reasoningtoolsvision |
| 3 | Gemini 3.1 Pro Preview google/gemini-3.1-pro-preview | 79.1 | $4.50 | 1.0M | reasoningtoolsvision |
| 4 | Muse Spark 1.3 meta/muse-spark-1.3 | 78.0 | $2.00 | 1.0M | reasoningtoolsvision |
| 5 | Claude Fable 5 anthropic/claude-fable-5 | 75.8 | $20.00 | 1M | reasoningtoolsvision |
| 6 | Gemini 3.5 Flash google/gemini-3.5-flash | 75.6 | $3.38 | 1.0M | reasoningtoolsvision |
| 7 | GPT-6 Astra openai/gpt-6-astra | 75.6 | $20.00 | 1.1M | reasoningtoolsvision |
| 8 | Gemini 3.6 Flash google/gemini-3.6-flash | 75.4 | $1.50 | 1.0M | reasoningtoolsvision |
| 9 | SpaceXAI: Grok 4.7 x-ai/grok-4.7 | 75.3 | $2.40 | 500K | reasoningtoolsvision |
| 10 | Muse Spark 1.2 meta/muse-spark-1.2 | 74.3 | $2.00 | 1.0M | reasoningtoolsvision |
45 more models qualify — see the full leaderboard.
FAQ
What is the best LLM for Instruction Following in September 2026?
Gemini 3.8 Flash leads on LiveBench instruction following at 81.4, ahead of Gemini 3.7 Flash (79.9) and Gemini 3.1 Pro Preview (79.1).
What is the best value in the top 10 for instruction following?
Muse Spark 1.2. Lowest measured cost per point among the leaders — $0.0485 per point for a score of 74.3.
What is the best under $1.00 / 1m for instruction following?
DeepSeek V4 Flash Vision Exp. Highest score at a blended list price of $1.00 per 1M tokens or less — 71.0 at $0.330.
How is this ranking produced?
Ranked by LiveBench’s instruction-following category: producing output under explicit format and content constraints.
Other rankings
How this ranking is produced
- One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
- Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
- Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.
A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.