// cache_model

Prompt caching calculator

Caching discounts the part of your prompt that never changes — and can charge extra to store it. Split your prompt into its stable prefix and the part that changes, set how often the cache is warm, and see what caching does to the bill on every model that publishes cache pricing.

208 models with cache pricing · OpenRouterHow prompt caching worksFull cost calculatorPrice table

At an 85% hit rate, caching cuts the typical bill by 44%

Across 208 models that publish a cached-input price, savings range from 0% to 62% of the monthly bill. Your prompt is 87% cacheable prefix — output tokens and the changing part of the input are billed in full either way, which caps what caching can do.

Monthly cost with and without prompt caching, per model
ModelInput / cached read / writeNo cacheWith cacheSavedBreak-even
MiMo-V2.6-Pro-UltraSpeedxiaomi$4.35 / $0.036 / $4.35*$4,698$1,764$2,934(62%)any
MiMo-V2.6-Proxiaomi$0.435 / $0.0036 / $0.435*$469.80$176.45$293.35(62%)any
MiMo-V2.5-Proxiaomi$0.435 / $0.0036 / $0.435*$469.80$176.45$293.35(62%)any
GPT-5 Image Miniopenai$2.50 / $0.250 / $2.50*$2,460$930.00$1,530(62%)any
MiMo-V2.6-Flashxiaomi$0.140 / $0.0028 / $0.140*$151.20$57.90$93.30(62%)any
MiMo-V2.5xiaomi$0.140 / $0.0028 / $0.140*$151.20$57.90$93.30(62%)any
Muse Spark 1.3 Contributormeta$0.100 / $0.0020 / $0.100*$108.00$41.36$66.64(62%)any
Muse Spark 1.2 Contributormeta$0.100 / $0.0020 / $0.100*$108.00$41.36$66.64(62%)any
Ministral 3 14B 2512mistralai$0.200 / $0.020 / $0.200*$200.00$77.60$122.40(61%)any
Ministral 3 8B 2512mistralai$0.150 / $0.015 / $0.150*$150.00$58.20$91.80(61%)any
Ministral 3 3B 2512mistralai$0.100 / $0.010 / $0.100*$100.00$38.80$61.20(61%)any
GPT-5 Imageopenai$10.00 / $1.25 / $10.00*$10,000$4,050$5,950(60%)any
DeepSeek V4 Pro 0423deepseek$0.865 / $0.072 / $0.865*$934.53$395.15$539.37(58%)any
DeepSeek V4 Flash Vision Expdeepseek$0.220 / $0.0070 / $0.220*$255.20$110.36$144.84(57%)any
Laguna S 2.1poolside$0.090 / $0.0090 / $0.090*$97.20$42.12$55.08(57%)any
DeepSeek V4 Pro 0813deepseek$0.660 / $0.022 / $0.660*$765.60$331.76$433.84(57%)any
Hy4 previewtencent$0.834 / $0.042 / $0.834*$967.36$428.80$538.56(56%)any
Gemini 3.8 Flashgoogle$0.750 / $0.075 / $0.042$990.00$446.00$544.00(55%)any
Gemini 3.7 Flashgoogle$0.750 / $0.075 / $0.042$990.00$446.00$544.00(55%)any
Gemini 3.6 Flashgoogle$0.750 / $0.075 / $0.042$990.00$446.00$544.00(55%)any
LongCat 2.0meituan$0.300 / $0.0060 / $0.300*$372.00$172.08$199.92(54%)any
SpaceXAI: Grok 4.3x-ai$1.25 / $0.200 / $1.25*$1,350$636.00$714.00(53%)any
SpaceXAI: Grok 4.20 Multi-Agentx-ai$1.25 / $0.200 / $1.25*$1,350$636.00$714.00(53%)any
SpaceXAI: Grok 4.20x-ai$1.25 / $0.200 / $1.25*$1,350$636.00$714.00(53%)any
Paretounbiased$2.50 / $0.250 / $2.50*$2,900$1,370$1,530(53%)any
* No cache-write price published; writes billed at the normal input rate. Prices per 1M tokens.

Prompt caching FAQ

What is prompt caching?

When the start of a prompt is identical across requests — a system prompt, tool definitions, a document you keep asking about — the provider can keep its processed form and reuse it. Reusing it (a cache hit) is billed at a steep discount to normal input; storing it (a cache write) is billed at or above normal input. The changing part of the prompt and every output token are billed in full either way.

How much does prompt caching save?

For an assistant with a 4,000-token system prompt, 600 tokens of user input, 400 tokens of output and an 85% hit rate, the median saving across 208 models with published cache prices is 44% of the monthly bill. The bigger the stable prefix relative to everything else, the bigger the saving — a long document questioned repeatedly saves far more than a short prompt.

Can prompt caching cost more than not caching?

Yes, when the provider charges a premium to write the cache and too few requests reuse it. Claude Fable 5.1 charges $12.50 per 1M tokens to write against $10.00 for normal input and $0.250 to read, so caching only pays once more than 20% of requests hit the cache. Below that, every miss costs more than it would have uncached. The break-even column in the table shows this for every model.

How do I get a high cache hit rate?

Put everything stable at the very start of the prompt — system instructions, tool schemas, reference documents — and everything that changes at the end. A single changed token early in the prompt invalidates the cache from that point on. Caches expire after a few minutes without use, so steady traffic hits far more often than occasional bursts, and most providers only cache prefixes above a minimum length, often around a thousand tokens.

What does this calculator leave out?

Storage-time fees (some providers bill cached context per hour it is held), batch discounts and negotiated rates. Where a model publishes a cached-read price but no write price, writes are billed at the normal input rate — how automatic caching is usually priced. All prices are OpenRouter list prices, refreshed every 15 minutes.

For the whole bill rather than just the caching delta, use the LLM cost calculator; to find cheaper models outright, see the best-value ranking. Worked figures above are recomputed from list prices on every refresh.