Moonshot AI · API pricing
Kimi K3 pricing
Moonshot AI shipped Kimi K3 on July 16, 2026 at $3.00 per 1M input tokens and $15.00 per 1M output. That is five times cheaper per token than GPT-5.6 Sol. It is not five times cheaper per finished job, and the gap between those two sentences is what this page is about.
The short answer
- $3.00 in / $15.00 out per 1M tokens, on Moonshot's first-party API and on OpenRouter.
- Cached input drops to $0.30 per 1M — a 90% discount, and the single biggest lever on your bill if you reuse a long system prompt or codebase context.
- Cost per completed task: $0.94 in Artificial Analysis' agentic evaluation — versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8. A 10% saving against Sol, not an 80% one.
- The catch is verbosity. Thinking mode is always on, and K3 emitted roughly 130M output tokens across the evaluation suite against a 63M median — output is the expensive side of the meter.
The rate card
| Meter | Per 1M tokens | Notes |
|---|---|---|
| Input (cache miss) | $3.00 | Standard rate for fresh context |
| Input (cache hit) | $0.30 | 90% off — applies to repeated prefixes |
| Output | $15.00 | Includes thinking-mode tokens |
Confirmedrates Artificial Analysis reports a blended rate of $2.31 per 1M tokens at their 3:1 input-to-output mix. Your blend will be worse than that if you run agentic loops, because reasoning traces land on the $15.00 side of the meter.
Why per-token price misleads here
Almost every page comparing K3 to the frontier stops at the rate card, where K3 looks like a rout: $3.00 against $5.00 for GPT-5.6 Sol on input, $15.00 against $30.00 on output. Same headline numbers, one fifth the price.
The rate card is not the bill. K3's thinking mode cannot be switched off, so every request carries a reasoning trace you pay for at the output rate. In Artificial Analysis' run across their evaluation suite, K3 produced about 130 million output tokens where the median model produced 63 million — roughly double. Multiply a 5× cheaper token by 2× the tokens and the advantage collapses to something like 2.5×; measured end to end on real tasks, Artificial Analysis puts it at $0.94 versus $1.04.
| Model | Cost per task | GDPval-AA v2 Elo |
|---|---|---|
| Kimi K3 | $0.94 | 1668 |
| GPT-5.6 Sol | $1.04 | — |
| Claude Opus 4.8 | $1.80 | 1600 |
| GPT-5.5 | — | 1494 |
| Claude Fable 5 | — | 1760 |
Cost-per-task and Elo figures as published by Artificial Analysis, retrieved July 21, 2026. Dashes mean the figure was not given in the source — we have not filled them in by inference.
The cache discount is the real saving
If you are going to save money on K3, this is where. Cached input bills at $0.30 per 1M against $3.00 — a 90% cut on any prefix the model has already seen. With a 1M token context window, the natural pattern is to pin a large stable prefix (a system prompt, a repository, a document set) and vary only the tail.
Concretely: 200K tokens of pinned context, hit 1,000 times in a month, costs $600 at the miss rate and $60 at the hit rate. That $540 dwarfs anything you will save by shaving output. Structure prompts so the stable part comes first, and do not interleave per-request data into the prefix.
What you give up: speed
K3 is cheap and slow. Artificial Analysis measures 39.5 output tokens per second, ranking it 138th of 186 models tracked, with a 4.23 second time to first token. Combined with the verbosity noted above — twice the tokens at a third the speed — a K3 turn can take several times longer in wall-clock terms than a frontier model doing the same job.
That makes the choice fairly clean. For batch work, overnight agentic runs, and long-context analysis where nobody is watching a cursor blink, the economics are good. For anything a user is waiting on interactively, the latency is the cost that matters, and it is not denominated in dollars.
Price it against the alternatives
Enter your monthly volume. Note that this calculator uses list rates at cache-miss prices — it is the ceiling, not the bill you will get if your caching works.
Cost calculator
Enter your monthly token volume to see what each tier would cost.
| Tier | Input cost | Output cost | Monthly total |
|---|---|---|---|
| Kimi K3 | $3.00 | $3.00 | $6.00 |
| Sol | $5.00 | $6.00 | $11.00 |
| Terra | $2.50 | $3.00 | $5.50 |
| Qwen3.7 Max | $1.48 | $0.885 | $2.36cheapest |
Estimates only — list prices, before any caching discount, batch pricing, or enterprise agreement.
One number that already moved
At launch, Artificial Analysis placed K3 #3 on its Intelligence Index with a score of 57. As of July 21 the same score reads #4 of 186, behind Fable 5 (59.86), GPT-5.6 Sol max (58.89) and GPT-5.6 Sol xhigh (57.65). K3 did not get worse — the leaderboard got more crowded, and the reasoning-effort variants of Sol are counted separately.
It is worth flagging because “#3 in the world” is now propagating across dozens of pages that will never update it. The underlying claim — a Chinese model within three points of the frontier at a fifth of the token price — holds either way.
What changes next
Moonshot has promised open weights by July 27, 2026. If that lands, the pricing question splits in two: Moonshot's API rate, and whatever third-party hosts charge to serve the weights — which historically undercuts the first-party price. We are tracking the date on the Kimi K3 open weights page.
Also on this site: Qwen 3.8 pricing — Alibaba's answer to K3, previewed three days later with no per-token price at all — and GPT-5.6 pricing.
Sources
- OpenRouter — Kimi K3 — retrieved 2026-07-21
- Artificial Analysis — Kimi K3 achieves #3 in the Intelligence Index — retrieved 2026-07-21
- Artificial Analysis — Kimi K3 model page (live index and price) — retrieved 2026-07-21
- The Decoder — Alibaba's Qwen takes on Kimi K3 with open-weight Qwen 3.8 — retrieved 2026-07-21
- Hugging Face community post — Kimi K3 architecture, MXFP4 quantization and weight-release notes — retrieved 2026-07-21
- OpenRouter — GPT-5.6 Sol — retrieved 2026-07-17
- OpenRouter — GPT-5.6 Terra — retrieved 2026-07-17
- OpenRouter — Qwen model catalogue — retrieved 2026-07-20
Figures on this page last checked against these sources on 2026-07-21. Vendors change pricing and specs without notice — if a number here disagrees with the vendor's own page, trust the vendor.