// ranked · September 2026
Cheapest LLMs with Tool Calling
Every paid model that supports tool (function) calling, ranked by blended list price — 3:1 input:output per 1M tokens. Free tiers are on their own page.
Mistral Nemo is the cheapest paid model with tool calling, at $0.022 per 1M tokens blended. Ling 3.0 Flash follows at $0.032.
Mistral Nemo vs Ling 3.0 Flash, head to headTop 10
| # | Model | Blended / 1M | Context | LiveBench | Capabilities |
|---|---|---|---|---|---|
| 1 | Mistral Nemo mistralai/mistral-nemo | $0.022 | 131K | — | tools |
| 2 | Ling 3.0 Flash inclusionai/ling-3.0-flash | $0.032 | 262K | — | reasoningtools |
| 3 | gpt-oss-20b openai/gpt-oss-20b | $0.036 | 131K | — | reasoningtools |
| 4 | Qwen3.7 Flash qwen/qwen3.7-flash | $0.055 | 1M | — | reasoningtoolsvision |
| 5 | Llama 3.1 8B Instruct meta-llama/llama-3.1-8b-instruct | $0.057 | 131K | — | tools |
| 6 | DeepSeek V4 Flash 0423 deepseek/deepseek-v4-flash | $0.061 | 1.0M | 65.5 | reasoningtools |
| 7 | Nova Micro 1.0 amazon/nova-micro-v1 | $0.061 | 128K | — | tools |
| 8 | Mercury 2.5 inception/mercury-2.5 | $0.068 | 260K | — | reasoningtools |
| 9 | Laguna XS 2.1 poolside/laguna-xs-2.1 | $0.075 | 262K | — | reasoningtools |
| 10 | Gemma 3 12B google/gemma-3-12b-it | $0.075 | 131K | — | toolsvision |
264 more models qualify — see the full leaderboard.
FAQ
What is the cheapest LLM with tool calling in September 2026?
Mistral Nemo is the cheapest paid model with tool calling, at $0.022 per 1M tokens blended. Ling 3.0 Flash follows at $0.032.
How is this ranking produced?
Every paid model that supports tool (function) calling, ranked by blended list price — 3:1 input:output per 1M tokens. Free tiers are on their own page.
Other rankings
How this ranking is produced
- One metric, stated above — nothing here is weighted or scored by us. Scores come from LiveBench release 2026-06-25; models without a published run don't appear in score-based rankings.
- Price, context and capabilities — live from OpenRouter, refreshed every 15 minutes.
- Ties — scores less than a point apart are called a tie; effort settings alone move a LiveBench score by more than that.
A public benchmark is someone else's workload. Before committing, see LLM & agent evaluation.