// ranked · September 2026

LLMs with the Longest Context Window

Ranked by maximum context window in tokens, cheaper first on ties. A window is a ceiling, not a guarantee — recall degrades well before it on most models.

OpenRouter + LiveBenchAll rankingsAll models

SpaceXAI: Grok 4.20 Multi-Agent has the largest context window listed, at 2,000,000 tokens. SpaceXAI: Grok 4.20 matches it.

SpaceXAI: Grok 4.20 Multi-Agent vs SpaceXAI: Grok 4.20, head to head

Top 10

#ModelContextBlended / 1MLiveBenchCapabilities
1SpaceXAI: Grok 4.20 Multi-Agent

x-ai/grok-4.20-multi-agent

2M$1.56
reasoningvision
2SpaceXAI: Grok 4.20

x-ai/grok-4.20

2M$1.56
reasoningtoolsvision
3Llama 4 Scout

meta-llama/llama-4-scout

1.3M$0.150
toolsvision
4DeepSeek V4 Flash 0731

deepseek/deepseek-v4-flash-0731

1.3M$0.19074.2
reasoningtools
5GLM 5.3 Flash

z-ai/glm-5.3-flash

1.3M$0.23771.6
reasoningtoolsvision
6GLM 5.3

z-ai/glm-5.3

1.3M$0.86276.1
reasoningtools
7MiMo-V2.5

xiaomi/mimo-v2.5

1.1M$0.175
reasoningtoolsvision
8GPT-6 Luna Pro

openai/gpt-6-luna-pro

1.1M$0.200
reasoningtoolsvision
9GPT-6 Luna

openai/gpt-6-luna

1.1M$0.20072.0
reasoningtoolsvision
10GPT-5.6 Luna Pro

openai/gpt-5.6-luna-pro

1.1M$0.450
reasoningtoolsvision

333 more models qualify — see the full leaderboard.

FAQ

Which LLM has the longest context window in September 2026?

SpaceXAI: Grok 4.20 Multi-Agent has the largest context window listed, at 2,000,000 tokens. SpaceXAI: Grok 4.20 matches it.

How is this ranking produced?

Ranked by maximum context window in tokens, cheaper first on ties. A window is a ceiling, not a guarantee — recall degrades well before it on most models.

Other rankings

How this ranking is produced

  • One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
  • Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
  • Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.

A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.