// ranked · September 2026
LLMs with the Longest Context Window
Ranked by maximum context window in tokens, cheaper first on ties. A window is a ceiling, not a guarantee — recall degrades well before it on most models.
SpaceXAI: Grok 4.20 Multi-Agent has the largest context window listed, at 2,000,000 tokens. SpaceXAI: Grok 4.20 matches it.
SpaceXAI: Grok 4.20 Multi-Agent vs SpaceXAI: Grok 4.20, head to headTop 10
| # | Model | Context | Blended / 1M | LiveBench | Capabilities |
|---|---|---|---|---|---|
| 1 | SpaceXAI: Grok 4.20 Multi-Agent x-ai/grok-4.20-multi-agent | 2M | $1.56 | — | reasoningvision |
| 2 | SpaceXAI: Grok 4.20 x-ai/grok-4.20 | 2M | $1.56 | — | reasoningtoolsvision |
| 3 | Llama 4 Scout meta-llama/llama-4-scout | 1.3M | $0.150 | — | toolsvision |
| 4 | DeepSeek V4 Flash 0731 deepseek/deepseek-v4-flash-0731 | 1.3M | $0.190 | 74.2 | reasoningtools |
| 5 | GLM 5.3 Flash z-ai/glm-5.3-flash | 1.3M | $0.237 | 71.6 | reasoningtoolsvision |
| 6 | GLM 5.3 z-ai/glm-5.3 | 1.3M | $0.862 | 76.1 | reasoningtools |
| 7 | MiMo-V2.5 xiaomi/mimo-v2.5 | 1.1M | $0.175 | — | reasoningtoolsvision |
| 8 | GPT-6 Luna Pro openai/gpt-6-luna-pro | 1.1M | $0.200 | — | reasoningtoolsvision |
| 9 | GPT-6 Luna openai/gpt-6-luna | 1.1M | $0.200 | 72.0 | reasoningtoolsvision |
| 10 | GPT-5.6 Luna Pro openai/gpt-5.6-luna-pro | 1.1M | $0.450 | — | reasoningtoolsvision |
333 more models qualify — see the full leaderboard.
FAQ
Which LLM has the longest context window in September 2026?
SpaceXAI: Grok 4.20 Multi-Agent has the largest context window listed, at 2,000,000 tokens. SpaceXAI: Grok 4.20 matches it.
How is this ranking produced?
Ranked by maximum context window in tokens, cheaper first on ties. A window is a ceiling, not a guarantee — recall degrades well before it on most models.
Other rankings
How this ranking is produced
- One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
- Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
- Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.
A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.