Moonshot AI · API pricing
Kimi K3 pricing
OpenRouter lists Kimi K3 at $0.64 per 1M input tokens and $13.50 per 1M output, verified October 10, 2026. These are channel rates; this check does not establish Moonshot's direct API rates. Cost per completed task also depends on token usage, caching and reasoning output.
The short answer
- $0.64 in / $13.50 out per 1M tokens on OpenRouter. Provider routing may affect the bill.
- Cached input drops to $0.28 per 1M — a 56% discount on cache hits. Cache eligibility depends on the selected provider.
- Historical cost per completed task: $0.94 in Artificial Analysis' evaluation retrieved July 21, 2026, versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8. These figures have not been rerun at today's prices.
- The catch is verbosity. Thinking mode is always on, and K3 emitted roughly 130M output tokens across the evaluation suite against a 63M median — output is the expensive side of the meter.
The rate card
| Meter | Per 1M tokens | Notes |
|---|---|---|
| Input (cache miss) | $0.64 | Standard rate for fresh context |
| Input (cache hit) | $0.28 | 56% off the listed input rate; provider-dependent |
| Output | $13.50 | Includes thinking-mode tokens |
Confirmedchannel rates At a 3:1 input-to-output mix, the listed rates blend to $3.85 per 1M tokens, before cache discounts. Your blend will be higher if you run agentic loops, because reasoning traces land on the $13.50 side of the meter.
Why per-token price misleads here
Almost every page comparing K3 to the frontier stops at the rate card, where K3 looks like a rout: $0.64 against $5.00 for GPT-5.6 Sol on input, $13.50 against $30.00 on output. This comparison mixes K3 channel rates with Sol's listed rate; it is not a controlled task-cost comparison.
The rate card is not the bill. K3's thinking mode cannot be switched off, so every request carries a reasoning trace you pay for at the output rate. In Artificial Analysis' run across their evaluation suite, K3 produced about 130 million output tokens where the median model produced 63 million — roughly double. That historical run reported $0.94 versus $1.04 per task. Its token usage helps explain why a rate card alone cannot predict a bill; its dollar totals do not estimate today's task costs.
| Model | Cost per task | GDPval-AA v2 Elo |
|---|---|---|
| Kimi K3 | $0.94 | 1668 |
| GPT-5.6 Sol | $1.04 | — |
| Claude Opus 4.8 | $1.80 | 1600 |
| Claude Opus 5 (max) | $2.03 | 1861 |
| Claude Fable 5 | $2.75 | 1760 |
| GPT-5.5 | — | 1494 |
Cost-per-task and Elo figures as published by Artificial Analysis. Kimi K3, Sol, Opus 4.8 and GPT-5.5 retrieved July 21, 2026; the Opus 5 row and Fable 5's cost per task added July 25, 2026 after Opus 5 shipped and took the top of this benchmark — see Opus 5 pricing. Dashes mean the figure was not given in the source — we have not filled them in by inference. Elo drifts between retrievals as comparisons accumulate, so rows dated differently are not exactly simultaneous.
The cache discount is the real saving
If you are going to save money on K3, this is where. Cached input bills at $0.28 per 1M against $0.64 — a 56% cut for eligible cache hits. With a 1M token context window, the natural pattern is to pin a large stable prefix (a system prompt, a repository, a document set) and vary only the tail.
Concretely: 200K tokens of pinned context, hit 1,000 times in a month, costs $128.00 at the miss rate and $56.00 at the hit rate, saving $72.00 on input alone. This assumes every repeated prefix is cache-eligible. Structure prompts so the stable part comes first, and do not interleave per-request data into the prefix.
What you give up: speed
K3 is cheap and slow, and the gap has widened since launch. As of July 25, 2026 Artificial Analysis measures 32.1 output tokens per second, ranking it 149th of 190 models tracked — down from the 39.5 tok/s it was posting a week earlier.
The number that should actually decide your architecture is the latency: 161 seconds to first answer token. That figure is not a network artifact. Artificial Analysis measures time to the first token of the answer, which for a reasoning model includes the entire thinking block — and K3's thinking mode cannot be turned off. Combined with the verbosity noted above, you are waiting out roughly two and a half minutes of reasoning before the first useful character arrives. Any shorter time-to-first-token figure you see quoted for K3 is measuring the stream opening, not an answer beginning.
That makes the choice fairly clean. For batch work, overnight agentic runs, and long-context analysis where nobody is watching a cursor blink, the economics are good. For anything a user is waiting on interactively, the latency is the cost that matters, and it is not denominated in dollars.
Price it against the alternatives
Enter your monthly volume. Note that this calculator uses list rates at cache-miss prices — it is the ceiling, not the bill you will get if your caching works.
Cost calculator
Enter monthly volume and the typical prompt size. The applicable context tier is selected automatically.
| Tier | Input cost | Output cost | Monthly total |
|---|---|---|---|
| Kimi K3 | $0.64 | $2.70 | $3.34 |
| Sol | $5.00 | $6.00 | $11.00 |
| Terra | $2.00 | $2.40 | $4.40 |
| Qwen3.7 Max | $1.48 | $0.885 | $2.36cheapest |
Estimates only — list prices, before any caching discount, batch pricing, or enterprise agreement.
One number that already moved
K3's Intelligence Index score has been 57since the day it launched. Its rank has moved three times in nine days:
- At launch (July 16): #3, which is the figure most launch coverage froze on.
- July 21: #4 of 186 — behind Fable 5 (59.86), GPT-5.6 Sol max (58.89) and GPT-5.6 Sol xhigh (57.65), because the reasoning-effort variants of Sol are counted separately.
- July 25: #7 of 190. Four more models entered the index in four days.
The model did not change. The score did not change. A benchmark rank is a position in a moving field, not a property of the thing being measured — and “#3 in the world” is still propagating across dozens of pages that will never go back and check. The underlying claim — a Chinese model within three points of the frontier at a fifth of the token price — holds either way, which is the point of quoting the score rather than the rank.
What changes next
Moonshot has promised open weights by July 27, 2026. If that lands, the pricing question splits in two: Moonshot's API rate, and whatever third-party hosts charge to serve the weights — which historically undercuts the first-party price. We are tracking the date on the Kimi K3 open weights page.
Also on this site: Qwen 3.8 pricing — Alibaba's answer to K3, previewed three days later with no per-token price at all — and GPT-5.6 pricing.
Sources
- OpenRouter — Kimi K3 (channel price; models API verified) — retrieved 2026-10-10
- Artificial Analysis — Kimi K3 achieves #3 in the Intelligence Index — retrieved 2026-07-21
- Artificial Analysis — Kimi K3 model page (live index and price) — retrieved 2026-07-25
- Hugging Face — moonshotai organisation (checked: no Kimi K3 repository) — retrieved 2026-07-25
- Hugging Face — Kimi-K2.7-Code model card (Modified MIT License precedent) — retrieved 2026-07-25
- Tom's Hardware — Kimi K3 beats Claude Fable 5 in Frontend Code Arena — retrieved 2026-07-25
- The Decoder — Alibaba's Qwen takes on Kimi K3 with open-weight Qwen 3.8 — retrieved 2026-07-21
- Hugging Face community post — Kimi K3 architecture, MXFP4 quantization and weight-release notes — retrieved 2026-07-21
- OpenAI — API pricing — retrieved 2026-08-14
- OpenRouter — GPT-5.6 Sol (channel price) — retrieved 2026-09-03
- OpenAI — API changelog (July 30: Terra costs 20% less) — retrieved 2026-08-14
- OpenRouter — Qwen model catalogue — retrieved 2026-07-31
Figures on this page last checked against these sources on 2026-10-10. Vendors change pricing and specs without notice — if a number here disagrees with the vendor's own page, trust the vendor.