Ranking · data through 2026-08-21
The cheapest LLM APIs, ranked.
Every provider’s lightest tier, ranked by list-price input rate, checked nightly against official pricing pages. As of 2026-08-21, the cheapest tracked rate is Llama 3.1 8B on DeepInfra at $0.02 per 1M input tokens.
| # | Provider | Lightest tier | Input / 1M | Output / 1M | Context | The Hobbit* |
|---|---|---|---|---|---|---|
| 1 | DeepInfra | Llama 3.1 8B | $0.02 | $0.04 | 128K | $0.0051 |
| 2 | Cohere | Command R7B (12-2024) | $0.0375 | $0.15 | 128K | $0.0191 |
| 3 | Together AI | GPT-OSS 20B | $0.05 | $0.2 | 128K | $0.0254 |
| 4 | Fireworks AI | GPT-OSS 20B | $0.07 | $0.3 | 128K | $0.0381 |
| 5 | Groq | GPT-OSS 20B | $0.075 | $0.3 | 131K | $0.0381 |
| 6 | OpenRouter | Llama 3.3 70B | $0.1 | $0.32 | 128K | $0.0407 |
| 7 | Mistral AI | Mistral Small 4 | $0.15 | $0.6 | 32K | $0.0763 |
| 8 | Zhipu GLM | GLM-4.5-Air | $0.2 | $1.10 | 200K | $0.1399 |
| 9 | OpenAI | GPT-5.6 Luna | $0.2 | $1.20 | 1M | $0.1526 |
| 10 | SambaNova | GPT-OSS 120B | $0.22 | $0.59 | 128K | $0.075 |
| 11 | Alibaba Qwen | Qwen3.6-Flash | $0.25 | $1.50 | 1M | $0.1907 |
| 12 | MiniMax | MiniMax M2 | $0.3 | $1.20 | 200K | $0.1526 |
| 13 | Google Gemini | Flash-Lite 3.5 | $0.3 | $2.50 | 1M | $0.3179 |
| 14 | Cerebras | GPT-OSS 120B | $0.35 | $0.75 | 128K | $0.0954 |
| 15 | DeepSeek | V4 Flash · peak rate | $0.44 | $1.32 | 1M | $0.1678 |
| 16 | Moonshot Kimi | Kimi K2.5 | $0.6 | $3.00 | 256K | $0.3814 |
| 17 | Perplexity† | Sonar | $1.00 | $1.00 | 127K | $0.1271 |
| 18 | xAI Grok | Grok Build 0.1 | $1.00 | $2.00 | 256K | $0.2543 |
| 19 | Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | 200K | $0.6357 |
| 20 | AWS Bedrock | Claude Haiku 4.5 (Bedrock) | $1.00 | $5.00 | 200K | $0.6357 |
| 21 | Meta | Muse Spark 1.2 | $1.25 | $4.25 | 1M | $0.5403 |
| 22 | Sakana AI | Fugu Ultra | $5.00 | $30.00 | 1M | $3.81 |
Scroll the table sideways to see every column.
* Cost to generate The Hobbit (95,356 words ≈ 127,141 tokens) as output. † Perplexity bills web search separately per request; token prices exclude that fee.
How to read this table
- “Cheapest” means the lowest list-price input rate on the provider’s lightest tracked tier. Capability differs widely between tiers; cheap and fit-for-purpose are not the same thing.
- The spread is large: the priciest tracked flagship (Claude Fable 5, $10.00/1M input) costs about 500× the cheapest rate in this table.
- Most providers also offer batch and cached-input discounts that can cut these rates substantially. The calculator compares standard, batch and cache-hit pricing for a real workload.
About these numbers. Prices are checked nightly against each provider’s official pricing page and recorded in an append-only history (data through 2026-08-21). List prices only: promotions never enter the ranking. Corrections are published, not buried: see the methodology.