// ranked · September 2026

Best Value LLM

Ranked by measured cost per point of LiveBench overall score: the dollars LiveBench actually spent running the model, divided by the score it earned. It captures verbosity and reasoning-token spend that list prices hide.

OpenRouter + LiveBenchAll rankingsAll models

DeepSeek V4 Flash 0423 delivers LiveBench capability most cheaply, at $0.0083 of measured spend per overall point (score 65.5). GPT-6 Luna is next at $0.0134.

DeepSeek V4 Flash 0423 vs GPT-6 Luna, head to head

Top 10

#Model$ per pointBlended / 1MContextLiveBenchCapabilities
1DeepSeek V4 Flash 0423

deepseek/deepseek-v4-flash

$0.0083$0.0611.0M65.5
reasoningtools
2GPT-6 Luna

openai/gpt-6-luna

$0.0134$0.2001.1M72.0
reasoningtoolsvision
3SpaceXAI: Grok Build 0.1

x-ai/grok-build-0.1

$0.0144$1.25256K67.8
reasoningtoolsvision
4DeepSeek V4.1 Flash

deepseek/deepseek-v4.1-flash

$0.0157$0.2001.0M81.1
reasoningtoolsvision
5GLM 5.3 Flash

z-ai/glm-5.3-flash

$0.0161$0.2371.3M71.6
reasoningtoolsvision
6DeepSeek V4 Pro 0813

deepseek/deepseek-v4-pro-0813

$0.0241$0.9901.0M77.4
reasoningtools
7DeepSeek V4 Pro 0423

deepseek/deepseek-v4-pro

$0.0261$1.081.0M71.6
reasoningtools
8DeepSeek V4 Flash Vision Exp

deepseek/deepseek-v4-flash-vision-exp

$0.0277$0.3301.0M76.8
reasoningtoolsvision
9SpaceXAI: Grok 4.3

x-ai/grok-4.3

$0.0325$1.561M62.2
reasoningtoolsvision
10MiniMax M3

minimax/minimax-m3

$0.0339$0.5251.0M67.3
reasoningtoolsvision

45 more models qualify — see the full leaderboard.

FAQ

Which LLM gives the best value for money in September 2026?

DeepSeek V4 Flash 0423 delivers LiveBench capability most cheaply, at $0.0083 of measured spend per overall point (score 65.5). GPT-6 Luna is next at $0.0134.

How is this ranking produced?

Ranked by measured cost per point of LiveBench overall score: the dollars LiveBench actually spent running the model, divided by the score it earned. It captures verbosity and reasoning-token spend that list prices hide.

Other rankings

How this ranking is produced

  • One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
  • Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
  • Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.

A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.