// ranked · September 2026
Best Value LLM
Ranked by measured cost per point of LiveBench overall score: the dollars LiveBench actually spent running the model, divided by the score it earned. It captures verbosity and reasoning-token spend that list prices hide.
DeepSeek V4 Flash 0423 delivers LiveBench capability most cheaply, at $0.0083 of measured spend per overall point (score 65.5). GPT-6 Luna is next at $0.0134.
DeepSeek V4 Flash 0423 vs GPT-6 Luna, head to headTop 10
| # | Model | $ per point | Blended / 1M | Context | LiveBench | Capabilities |
|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash 0423 deepseek/deepseek-v4-flash | $0.0083 | $0.061 | 1.0M | 65.5 | reasoningtools |
| 2 | GPT-6 Luna openai/gpt-6-luna | $0.0134 | $0.200 | 1.1M | 72.0 | reasoningtoolsvision |
| 3 | SpaceXAI: Grok Build 0.1 x-ai/grok-build-0.1 | $0.0144 | $1.25 | 256K | 67.8 | reasoningtoolsvision |
| 4 | DeepSeek V4.1 Flash deepseek/deepseek-v4.1-flash | $0.0157 | $0.200 | 1.0M | 81.1 | reasoningtoolsvision |
| 5 | GLM 5.3 Flash z-ai/glm-5.3-flash | $0.0161 | $0.237 | 1.3M | 71.6 | reasoningtoolsvision |
| 6 | DeepSeek V4 Pro 0813 deepseek/deepseek-v4-pro-0813 | $0.0241 | $0.990 | 1.0M | 77.4 | reasoningtools |
| 7 | DeepSeek V4 Pro 0423 deepseek/deepseek-v4-pro | $0.0261 | $1.08 | 1.0M | 71.6 | reasoningtools |
| 8 | DeepSeek V4 Flash Vision Exp deepseek/deepseek-v4-flash-vision-exp | $0.0277 | $0.330 | 1.0M | 76.8 | reasoningtoolsvision |
| 9 | SpaceXAI: Grok 4.3 x-ai/grok-4.3 | $0.0325 | $1.56 | 1M | 62.2 | reasoningtoolsvision |
| 10 | MiniMax M3 minimax/minimax-m3 | $0.0339 | $0.525 | 1.0M | 67.3 | reasoningtoolsvision |
45 more models qualify — see the full leaderboard.
FAQ
Which LLM gives the best value for money in September 2026?
DeepSeek V4 Flash 0423 delivers LiveBench capability most cheaply, at $0.0083 of measured spend per overall point (score 65.5). GPT-6 Luna is next at $0.0134.
How is this ranking produced?
Ranked by measured cost per point of LiveBench overall score: the dollars LiveBench actually spent running the model, divided by the score it earned. It captures verbosity and reasoning-token spend that list prices hide.
Other rankings
How this ranking is produced
- One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
- Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
- Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.
A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.