AI APIs · Price comparison
AI API pricing comparison
Compare official vendor list prices separately from third-party channel prices. Every number below is priced per 1 million tokens, includes its applicable prompt range, and links back to a source.
What is the cheapest AI API here?
Cheapest input tokens
GPT-5.6 Luna
$0.20 per 1M input tokens
Official list price for < 272K.
Cheapest output tokens
GPT-5.6 Luna
$1.20 per 1M output tokens
Official list price for < 272K.
This is a rate-card answer, not a quality ranking. A model that needs more output tokens, retries, or human correction can cost more per completed task even when its listed token price is lower.
LLM API price comparison per 1M tokens
This first table is official list pricing only. Every context-dependent rate is expanded into its own row. OpenRouter and other retail/channel prices are kept in the separate table below and never determine the “cheapest official API” headline.
| Model | Provider | Input / 1M | Output / 1M | Cached input | Applicable prompt | Context |
|---|---|---|---|---|---|---|
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | — | < 272K | 1.05M |
| GPT-5.6 Luna | OpenAI | $0.40 | $1.80 | — | ≥ 272K | 1.05M |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | — | All prompts | Not stated | |
| Qwen3.5 Plus | Alibaba | $0.40 | $2.40 | — | ≤ 256K | 1M |
| Qwen3.5 Plus | Alibaba | $0.50 | $3.00 | — | > 256K | 1M |
| Gemini 3.7 Flash | $0.75 | $3.75 | — | All prompts | Not stated | |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $0.10 | All prompts | Not stated |
| Gemini 3.6 Flash | $1.50 | $7.50 | — | All prompts | Not stated | |
| Grok 4.6 | xAI | $2.00 | $6.00 | — | < Not stated | Not stated |
| Grok 4.6 | xAI | $4.00 | $12.00 | — | All prompts | Not stated |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | — | < 272K | 1.05M |
| GPT-5.6 Terra | OpenAI | $4.00 | $18.00 | — | ≥ 272K | 1.05M |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $0.20 | All prompts | 1M |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | — | < 272K | 1.05M |
| GPT-5.6 Sol | OpenAI | $10.00 | $45.00 | — | ≥ 272K | 1.05M |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | — | < Not stated | Not stated |
| Claude Opus 5 | Anthropic | $10.00 | $50.00 | — | All prompts | Not stated |
Models are sorted by base input price. Prices exclude batch discounts, promotions, enterprise agreements and tool-call charges. “Not stated” means the cited source did not publish a context window; it is not an estimate.
OpenRouter channel pricing
These are third-party channel rates, not vendor canonical list prices. They can include retail discounts or promotions and may differ from buying from the vendor.
| Model | Input / 1M | Output / 1M | Cached input | Applicable prompt | Context |
|---|---|---|---|---|---|
| Qwen3.7 Flash | $0.03 | $0.13 | $0.006 | < 32K | 1M |
| Qwen3.7 Flash | $0.10 | $0.40 | $0.02 | ≥ 32K – < 256K | 1M |
| Qwen3.7 Flash | $0.20 | $0.80 | $0.04 | ≥ 256K | 1M |
| Qwen3.5 Flash | $0.065 | $0.26 | — | All prompts | 1M |
| Qwen3.6 Flash | $0.1875 | $1.125 | — | < 256K | 1M |
| Qwen3.6 Flash | $0.75 | $3.00 | — | ≥ 256K | 1M |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 | $0.025 | All prompts | 1.048576M |
| Qwen3.5 Plus | $0.30 | $1.80 | — | < 256K | 1M |
| Qwen3.5 Plus | $0.375 | $2.25 | — | ≥ 256K | 1M |
| Qwen3.7 Plus | $0.32 | $1.28 | $0.064 | < 256K | 1M |
| Qwen3.7 Plus | $0.96 | $3.84 | $0.192 | ≥ 256K | 1M |
| Qwen3.6 Plus | $0.325 | $1.95 | — | < 256K | 1M |
| Qwen3.6 Plus | $1.30 | $3.90 | — | ≥ 256K | 1M |
| Kimi K2.7 Code | $0.71 | $3.50 | $0.15 | All prompts | 262.144K |
| Qwen3.6 Max Preview | $1.027 | $6.162 | — | < 128K | 262.144K |
| Qwen3.6 Max Preview | $1.58 | $9.48 | — | ≥ 128K | 262.144K |
| Qwen3.7 Max | $1.475 | $4.425 | — | All prompts | 1M |
| GPT-5.6 Sol | $2.50 | $15.00 | $0.25 | All prompts | 1.05M |
| Kimi K3 | $3.00 | $15.00 | $0.30 | All prompts | 1M |
AI token cost calculator
Enter the same workload for every model. The calculator combines input and output cost and marks the lowest list-price total for that token mix.
Cost calculator
Enter monthly volume and the typical prompt size. The applicable context tier is selected automatically.
| Tier | Input cost | Output cost | Monthly total |
|---|---|---|---|
| Luna | $0.20 | $0.24 | $0.44cheapest |
| Gemini 3.5 Flash-Lite | $0.30 | $0.50 | $0.80 |
| Qwen3.5 Plus (Alibaba international, ≤256K) | $0.40 | $0.48 | $0.88 |
| Gemini 3.7 Flash (introductory, through 2026-12-31) | $0.75 | $0.75 | $1.50 |
| Haiku 4.5 | $1.00 | $1.00 | $2.00 |
| Gemini 3.6 Flash | $1.50 | $1.50 | $3.00 |
| Grok 4.6 | $2.00 | $1.20 | $3.20 |
| Terra | $2.00 | $2.40 | $4.40 |
| Sonnet 5 (introductory, through 2026-08-31) | $2.00 | $2.00 | $4.00 |
| Sol | $5.00 | $6.00 | $11.00 |
| Opus 5 | $5.00 | $5.00 | $10.00 |
Estimates only — list prices, before any caching discount, batch pricing, or enterprise agreement.
How to compare AI API pricing without fooling yourself
Separate input and output
Output is often several times more expensive. A chat app and a document classifier can rank models differently even at the same total token volume.
Measure tokens per accepted result
Reasoning verbosity, retries and rejected answers change the real bill. Benchmark the task you actually run, not only the vendor's per-token rate.
Recheck the rate card
Vendors change prices and promotions. Each row carries sourced data, and the page shows when this comparison was last verified.
Need detail on one family? Start with the GPT-5.6 price breakdown or the Kimi K3 cost analysis.
Sources
- OpenAI — API pricing — retrieved 2026-08-14
- OpenAI — API changelog (July 30: Terra costs 20% less) — retrieved 2026-08-14
- Google — Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber (launch post) — retrieved 2026-07-22
- Alibaba Cloud Model Studio — Qwen3.5 Plus international official pricing — retrieved 2026-08-10
- OpenRouter — Qwen3.5 Plus 2026-04-20 channel pricing and 256K override — retrieved 2026-08-10
- Google — Introducing Gemini 3.7 Flash (launch and introductory pricing) — retrieved 2026-08-14
- Anthropic — API pricing (introductory rate and its 2026-08-31 expiry) — retrieved 2026-08-01
- xAI — Introducing Grok 4.6 (availability and API pricing) — retrieved 2026-08-13
- OpenRouter — GPT-5.6 Sol (channel price) — retrieved 2026-08-18
- Anthropic — Introducing Claude Opus 5 (launch post, rates and Fast Mode) — retrieved 2026-07-25
- Artificial Analysis — Opus 5: Fable 5 level intelligence at a lower cost per task — retrieved 2026-07-25
- Artificial Analysis — Claude Opus 5 (max) model page — retrieved 2026-07-25
- Implicator.ai — Opus 5 cut tokens 17% but cost more to benchmark than Opus 4.8 — retrieved 2026-07-25
- OpenRouter — Qwen model catalogue — retrieved 2026-07-31
- OpenRouter — Gemini 3.1 Flash Lite (channel price) — retrieved 2026-08-03
- OpenRouter — Kimi K2.7 Code (channel price) — retrieved 2026-08-17
- OpenRouter — Kimi K3 — retrieved 2026-07-21
- Artificial Analysis — Kimi K3 achieves #3 in the Intelligence Index — retrieved 2026-07-21
- Artificial Analysis — Kimi K3 model page (live index and price) — retrieved 2026-07-25
- Hugging Face — moonshotai organisation (checked: no Kimi K3 repository) — retrieved 2026-07-25
- Hugging Face — Kimi-K2.7-Code model card (Modified MIT License precedent) — retrieved 2026-07-25
- Tom's Hardware — Kimi K3 beats Claude Fable 5 in Frontend Code Arena — retrieved 2026-07-25
- The Decoder — Alibaba's Qwen takes on Kimi K3 with open-weight Qwen 3.8 — retrieved 2026-07-21
- Hugging Face community post — Kimi K3 architecture, MXFP4 quantization and weight-release notes — retrieved 2026-07-21
Figures on this page last checked against these sources on 2026-08-17. Vendors change pricing and specs without notice — if a number here disagrees with the vendor's own page, trust the vendor.