The cheapest LLM APIs, ranked.
Every provider’s lightest tier, ranked by list-price input rate from the nightly checked record. Each row states whether its latest observation was verified or carried forward. As of 2026-10-05, the cheapest tracked rate is Llama 3.1 8B on DeepInfra at $0.02 per 1M input tokens.
| # | Provider | Lightest tier | Input / 1M | Output / 1M | Context | The Hobbit* |
|---|---|---|---|---|---|---|
| 1 | DeepInfra TokenScale observation: verified 2026-10-05 | Llama 3.1 8B | $0.02 | $0.04 | 128K | $0.0051 |
| 2 | Cohere TokenScale observation: verified 2026-10-05 | Command R7B (12-2024) | $0.0375 | $0.15 | 128K | $0.0191 |
| 3 | Together AI TokenScale observation: carried 2026-10-05 | GPT-OSS 20B | $0.05 | $0.2 | 128K | $0.0254 |
| 4 | Fireworks AI TokenScale observation: carried 2026-10-05 | GPT-OSS 20B | $0.07 | $0.3 | 128K | $0.0381 |
| 5 | Groq TokenScale observation: verified 2026-10-05 | GPT-OSS 20B | $0.075 | $0.3 | 131K | $0.0381 |
| 6 | OpenRouter TokenScale observation: verified 2026-10-05 | Llama 3.3 70B | $0.1 | $0.32 | 128K | $0.0407 |
| 7 | OpenAI TokenScale observation: verified 2026-10-05 | GPT-6 Luna | $0.1 | $0.5 | 1.05M | $0.0636 |
| 8 | Mistral AI TokenScale observation: verified 2026-10-05 | Mistral Small 4 | $0.15 | $0.6 | 32K | $0.0763 |
| 9 | Zhipu GLM TokenScale observation: verified 2026-10-05 | GLM-4.5-Air | $0.2 | $1.10 | 200K | $0.1399 |
| 10 | SambaNova TokenScale observation: verified 2026-10-05 | GPT-OSS 120B | $0.22 | $0.59 | 128K | $0.075 |
| 11 | Alibaba Qwen TokenScale observation: verified 2026-10-05 | Qwen3.6-Flash | $0.25 | $1.50 | 1M | $0.1907 |
| 12 | DeepSeek TokenScale observation: verified 2026-10-05 | V4.1 Flash · peak rate | $0.3 | $1.20 | 1M | $0.1526 |
| 13 | MiniMax TokenScale observation: verified 2026-10-05 | MiniMax M2 | $0.3 | $1.20 | 200K | $0.1526 |
| 14 | Google Gemini TokenScale observation: verified 2026-10-05 | Flash-Lite 3.5 | $0.3 | $2.50 | 1M | $0.3179 |
| 15 | Cerebras TokenScale observation: verified 2026-10-05 | GPT-OSS 120B | $0.35 | $0.75 | 128K | $0.0954 |
| 16 | Moonshot Kimi TokenScale observation: carried 2026-10-05 | Kimi K2.5 | $0.6 | $3.00 | 256K | $0.3814 |
| 17 | Perplexity† TokenScale observation: carried 2026-10-05 | Sonar | $1.00 | $1.00 | 127K | $0.1271 |
| 18 | xAI Grok TokenScale observation: verified 2026-10-05 | Grok Build 0.1 | $1.00 | $2.00 | 256K | $0.2543 |
| 19 | Anthropic TokenScale observation: verified 2026-10-05 | Claude Haiku 4.5 | $1.00 | $5.00 | 200K | $0.6357 |
| 20 | AWS Bedrock TokenScale observation: carried 2026-10-05 | Claude Haiku 4.5 (Bedrock) | $1.00 | $5.00 | 200K | $0.6357 |
| 21 | Meta TokenScale observation: verified 2026-10-05 | Muse Spark 1.2 | $1.25 | $4.25 | 1M | $0.5403 |
| 22 | Sakana AI TokenScale observation: verified 2026-10-05 | Fugu Ultra | $5.00 | $30.00 | 1M | $3.81 |
Scroll the table sideways to see every column.
* Cost to generate The Hobbit (95,356 words ≈ 127,141 tokens) as output. † Perplexity bills web search separately per request; token prices exclude that fee.
How to read this table
- “Cheapest” means the lowest list-price input rate on the provider’s lightest tracked tier. Capability differs widely between tiers; cheap and fit-for-purpose are not the same thing.
- The spread is large: the priciest tracked flagship (Fugu Ultra, $5.00/1M input) costs about 250× the cheapest rate in this table.
- Some providers offer batch or cached-input discounts, but the rules vary by model and provider. These rankings and the calculator use standard on-demand pricing, with any carried observation marked, rather than applying one universal discount.
About these numbers. Provider pages are authoritative for current published prices. TokenScale checks those sources nightly and independently records what was verified, changed, corrected or carried forward (data through 2026-10-05). List prices only: promotions never enter the ranking. Corrections are published, not buried: see the methodology.
Use this ranking with
More pricing pages
- All tracked model pricing pages
- Model cost comparisons
- Public historical pricing dataset
- Claude API pricing
- OpenAI API pricing
- Gemini API pricing
- The state of AI pricing: the market, measured
- Price a real workload on the TokenScale calculator
- Night-by-night price history, one chart per provider
- The Chart: every model ranked, replayed at each price change