Moonshot AI · API pricing
Kimi K3 pricing
Moonshot AI shipped Kimi K3 on July 16, 2026 at $3.00 per 1M input tokens and $15.00 per 1M output. That is five times cheaper per token than GPT-5.6 Sol. It is not five times cheaper per finished job, and the gap between those two sentences is what this page is about.
The short answer
- $3.00 in / $15.00 out per 1M tokens, on Moonshot's first-party API and on OpenRouter.
- Cached input drops to $0.30 per 1M — a 90% discount, and the single biggest lever on your bill if you reuse a long system prompt or codebase context.
- Cost per completed task: $0.94 in Artificial Analysis' agentic evaluation — versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8. A 10% saving against Sol, not an 80% one.
- The catch is verbosity. Thinking mode is always on, and K3 emitted roughly 130M output tokens across the evaluation suite against a 63M median — output is the expensive side of the meter.
The rate card
| Meter | Per 1M tokens | Notes |
|---|---|---|
| Input (cache miss) | $3.00 | Standard rate for fresh context |
| Input (cache hit) | $0.30 | 90% off — applies to repeated prefixes |
| Output | $15.00 | Includes thinking-mode tokens |
Confirmedrates Artificial Analysis reports a blended rate of $2.31 per 1M tokens at their 3:1 input-to-output mix. Your blend will be worse than that if you run agentic loops, because reasoning traces land on the $15.00 side of the meter.
Why per-token price misleads here
Almost every page comparing K3 to the frontier stops at the rate card, where K3 looks like a rout: $3.00 against $5.00 for GPT-5.6 Sol on input, $15.00 against $30.00 on output. Same headline numbers, one fifth the price.
The rate card is not the bill. K3's thinking mode cannot be switched off, so every request carries a reasoning trace you pay for at the output rate. In Artificial Analysis' run across their evaluation suite, K3 produced about 130 million output tokens where the median model produced 63 million — roughly double. Multiply a 5× cheaper token by 2× the tokens and the advantage collapses to something like 2.5×; measured end to end on real tasks, Artificial Analysis puts it at $0.94 versus $1.04.
| Model | Cost per task | GDPval-AA v2 Elo |
|---|---|---|
| Kimi K3 | $0.94 | 1668 |
| GPT-5.6 Sol | $1.04 | — |
| Claude Opus 4.8 | $1.80 | 1600 |
| Claude Opus 5 (max) | $2.03 | 1861 |
| Claude Fable 5 | $2.75 | 1760 |
| GPT-5.5 | — | 1494 |
Cost-per-task and Elo figures as published by Artificial Analysis. Kimi K3, Sol, Opus 4.8 and GPT-5.5 retrieved July 21, 2026; the Opus 5 row and Fable 5's cost per task added July 25, 2026 after Opus 5 shipped and took the top of this benchmark — see Opus 5 pricing. Dashes mean the figure was not given in the source — we have not filled them in by inference. Elo drifts between retrievals as comparisons accumulate, so rows dated differently are not exactly simultaneous.
The cache discount is the real saving
If you are going to save money on K3, this is where. Cached input bills at $0.30 per 1M against $3.00 — a 90% cut on any prefix the model has already seen. With a 1M token context window, the natural pattern is to pin a large stable prefix (a system prompt, a repository, a document set) and vary only the tail.
Concretely: 200K tokens of pinned context, hit 1,000 times in a month, costs $600 at the miss rate and $60 at the hit rate. That $540 dwarfs anything you will save by shaving output. Structure prompts so the stable part comes first, and do not interleave per-request data into the prefix.
What you give up: speed
K3 is cheap and slow, and the gap has widened since launch. As of July 25, 2026 Artificial Analysis measures 32.1 output tokens per second, ranking it 149th of 190 models tracked — down from the 39.5 tok/s it was posting a week earlier.
The number that should actually decide your architecture is the latency: 161 seconds to first answer token. That figure is not a network artifact. Artificial Analysis measures time to the first token of the answer, which for a reasoning model includes the entire thinking block — and K3's thinking mode cannot be turned off. Combined with the verbosity noted above, you are waiting out roughly two and a half minutes of reasoning before the first useful character arrives. Any shorter time-to-first-token figure you see quoted for K3 is measuring the stream opening, not an answer beginning.
That makes the choice fairly clean. For batch work, overnight agentic runs, and long-context analysis where nobody is watching a cursor blink, the economics are good. For anything a user is waiting on interactively, the latency is the cost that matters, and it is not denominated in dollars.
Price it against the alternatives
Enter your monthly volume. Note that this calculator uses list rates at cache-miss prices — it is the ceiling, not the bill you will get if your caching works.
Cost calculator
Enter monthly volume and the typical prompt size. The applicable context tier is selected automatically.
| Tier | Input cost | Output cost | Monthly total |
|---|---|---|---|
| Kimi K3 | $3.00 | $3.00 | $6.00 |
| Sol | $5.00 | $6.00 | $11.00 |
| Terra | $2.00 | $2.40 | $4.40 |
| Qwen3.7 Max | $1.48 | $0.885 | $2.36cheapest |
Estimates only — list prices, before any caching discount, batch pricing, or enterprise agreement.
One number that already moved
K3's Intelligence Index score has been 57since the day it launched. Its rank has moved three times in nine days:
- At launch (July 16): #3, which is the figure most launch coverage froze on.
- July 21: #4 of 186 — behind Fable 5 (59.86), GPT-5.6 Sol max (58.89) and GPT-5.6 Sol xhigh (57.65), because the reasoning-effort variants of Sol are counted separately.
- July 25: #7 of 190. Four more models entered the index in four days.
The model did not change. The score did not change. A benchmark rank is a position in a moving field, not a property of the thing being measured — and “#3 in the world” is still propagating across dozens of pages that will never go back and check. The underlying claim — a Chinese model within three points of the frontier at a fifth of the token price — holds either way, which is the point of quoting the score rather than the rank.
What changes next
Moonshot has promised open weights by July 27, 2026. If that lands, the pricing question splits in two: Moonshot's API rate, and whatever third-party hosts charge to serve the weights — which historically undercuts the first-party price. We are tracking the date on the Kimi K3 open weights page.
Also on this site: Qwen 3.8 pricing — Alibaba's answer to K3, previewed three days later with no per-token price at all — and GPT-5.6 pricing.
Sources
- OpenRouter — Kimi K3 — retrieved 2026-07-21
- Artificial Analysis — Kimi K3 achieves #3 in the Intelligence Index — retrieved 2026-07-21
- Artificial Analysis — Kimi K3 model page (live index and price) — retrieved 2026-07-25
- Hugging Face — moonshotai organisation (checked: no Kimi K3 repository) — retrieved 2026-07-25
- Hugging Face — Kimi-K2.7-Code model card (Modified MIT License precedent) — retrieved 2026-07-25
- Tom's Hardware — Kimi K3 beats Claude Fable 5 in Frontend Code Arena — retrieved 2026-07-25
- The Decoder — Alibaba's Qwen takes on Kimi K3 with open-weight Qwen 3.8 — retrieved 2026-07-21
- Hugging Face community post — Kimi K3 architecture, MXFP4 quantization and weight-release notes — retrieved 2026-07-21
- OpenAI — API pricing — retrieved 2026-07-31
- OpenAI — API changelog (July 30: Terra costs 20% less) — retrieved 2026-07-31
- OpenRouter — Qwen model catalogue — retrieved 2026-07-31
Figures on this page last checked against these sources on 2026-07-25. Vendors change pricing and specs without notice — if a number here disagrees with the vendor's own page, trust the vendor.