// ranked · September 2026

Best LLM for Instruction Following

Ranked by LiveBench’s instruction-following category: producing output under explicit format and content constraints.

OpenRouter + LiveBenchAll rankingsAll models

Gemini 3.8 Flash leads on LiveBench instruction following at 81.4, ahead of Gemini 3.7 Flash (79.9) and Gemini 3.1 Pro Preview (79.1).

Gemini 3.8 Flash vs Gemini 3.7 Flash, head to head

Top 10

#ModelInstruction following scoreBlended / 1MContextCapabilities
1Gemini 3.8 Flash

google/gemini-3.8-flash

81.4$1.501.0M
reasoningtoolsvision
2Gemini 3.7 Flash

google/gemini-3.7-flash

79.9$1.501.0M
reasoningtoolsvision
3Gemini 3.1 Pro Preview

google/gemini-3.1-pro-preview

79.1$4.501.0M
reasoningtoolsvision
4Muse Spark 1.3

meta/muse-spark-1.3

78.0$2.001.0M
reasoningtoolsvision
5Claude Fable 5

anthropic/claude-fable-5

75.8$20.001M
reasoningtoolsvision
6Gemini 3.5 Flash

google/gemini-3.5-flash

75.6$3.381.0M
reasoningtoolsvision
7GPT-6 Astra

openai/gpt-6-astra

75.6$20.001.1M
reasoningtoolsvision
8Gemini 3.6 Flash

google/gemini-3.6-flash

75.4$1.501.0M
reasoningtoolsvision
9SpaceXAI: Grok 4.7

x-ai/grok-4.7

75.3$2.40500K
reasoningtoolsvision
10Muse Spark 1.2

meta/muse-spark-1.2

74.3$2.001.0M
reasoningtoolsvision

45 more models qualify — see the full leaderboard.

FAQ

What is the best LLM for Instruction Following in September 2026?

Gemini 3.8 Flash leads on LiveBench instruction following at 81.4, ahead of Gemini 3.7 Flash (79.9) and Gemini 3.1 Pro Preview (79.1).

What is the best value in the top 10 for instruction following?

Muse Spark 1.2. Lowest measured cost per point among the leaders — $0.0485 per point for a score of 74.3.

What is the best under $1.00 / 1m for instruction following?

DeepSeek V4 Flash Vision Exp. Highest score at a blended list price of $1.00 per 1M tokens or less — 71.0 at $0.330.

How is this ranking produced?

Ranked by LiveBench’s instruction-following category: producing output under explicit format and content constraints.

Other rankings

How this ranking is produced

  • One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
  • Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
  • Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.

A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.