Moonshot AI · API pricing

Kimi K3 pricing

Moonshot AI shipped Kimi K3 on July 16, 2026 at $3.00 per 1M input tokens and $15.00 per 1M output. That is five times cheaper per token than GPT-5.6 Sol. It is not five times cheaper per finished job, and the gap between those two sentences is what this page is about.

The short answer

  • $3.00 in / $15.00 out per 1M tokens, on Moonshot's first-party API and on OpenRouter.
  • Cached input drops to $0.30 per 1M — a 90% discount, and the single biggest lever on your bill if you reuse a long system prompt or codebase context.
  • Cost per completed task: $0.94 in Artificial Analysis' agentic evaluation — versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8. A 10% saving against Sol, not an 80% one.
  • The catch is verbosity. Thinking mode is always on, and K3 emitted roughly 130M output tokens across the evaluation suite against a 63M median — output is the expensive side of the meter.

The rate card

MeterPer 1M tokensNotes
Input (cache miss)$3.00Standard rate for fresh context
Input (cache hit)$0.3090% off — applies to repeated prefixes
Output$15.00Includes thinking-mode tokens

Confirmedrates Artificial Analysis reports a blended rate of $2.31 per 1M tokens at their 3:1 input-to-output mix. Your blend will be worse than that if you run agentic loops, because reasoning traces land on the $15.00 side of the meter.

Why per-token price misleads here

Almost every page comparing K3 to the frontier stops at the rate card, where K3 looks like a rout: $3.00 against $5.00 for GPT-5.6 Sol on input, $15.00 against $30.00 on output. Same headline numbers, one fifth the price.

The rate card is not the bill. K3's thinking mode cannot be switched off, so every request carries a reasoning trace you pay for at the output rate. In Artificial Analysis' run across their evaluation suite, K3 produced about 130 million output tokens where the median model produced 63 million — roughly double. Multiply a 5× cheaper token by 2× the tokens and the advantage collapses to something like 2.5×; measured end to end on real tasks, Artificial Analysis puts it at $0.94 versus $1.04.

ModelCost per taskGDPval-AA v2 Elo
Kimi K3$0.941668
GPT-5.6 Sol$1.04
Claude Opus 4.8$1.801600
GPT-5.51494
Claude Fable 51760

Cost-per-task and Elo figures as published by Artificial Analysis, retrieved July 21, 2026. Dashes mean the figure was not given in the source — we have not filled them in by inference.

The cache discount is the real saving

If you are going to save money on K3, this is where. Cached input bills at $0.30 per 1M against $3.00 — a 90% cut on any prefix the model has already seen. With a 1M token context window, the natural pattern is to pin a large stable prefix (a system prompt, a repository, a document set) and vary only the tail.

Concretely: 200K tokens of pinned context, hit 1,000 times in a month, costs $600 at the miss rate and $60 at the hit rate. That $540 dwarfs anything you will save by shaving output. Structure prompts so the stable part comes first, and do not interleave per-request data into the prefix.

What you give up: speed

K3 is cheap and slow. Artificial Analysis measures 39.5 output tokens per second, ranking it 138th of 186 models tracked, with a 4.23 second time to first token. Combined with the verbosity noted above — twice the tokens at a third the speed — a K3 turn can take several times longer in wall-clock terms than a frontier model doing the same job.

That makes the choice fairly clean. For batch work, overnight agentic runs, and long-context analysis where nobody is watching a cursor blink, the economics are good. For anything a user is waiting on interactively, the latency is the cost that matters, and it is not denominated in dollars.

Price it against the alternatives

Enter your monthly volume. Note that this calculator uses list rates at cache-miss prices — it is the ceiling, not the bill you will get if your caching works.

Cost calculator

Enter your monthly token volume to see what each tier would cost.

TierInput costOutput costMonthly total
Kimi K3$3.00$3.00$6.00
Sol$5.00$6.00$11.00
Terra$2.50$3.00$5.50
Qwen3.7 Max$1.48$0.885$2.36cheapest

Estimates only — list prices, before any caching discount, batch pricing, or enterprise agreement.

One number that already moved

At launch, Artificial Analysis placed K3 #3 on its Intelligence Index with a score of 57. As of July 21 the same score reads #4 of 186, behind Fable 5 (59.86), GPT-5.6 Sol max (58.89) and GPT-5.6 Sol xhigh (57.65). K3 did not get worse — the leaderboard got more crowded, and the reasoning-effort variants of Sol are counted separately.

It is worth flagging because “#3 in the world” is now propagating across dozens of pages that will never update it. The underlying claim — a Chinese model within three points of the frontier at a fifth of the token price — holds either way.

What changes next

Moonshot has promised open weights by July 27, 2026. If that lands, the pricing question splits in two: Moonshot's API rate, and whatever third-party hosts charge to serve the weights — which historically undercuts the first-party price. We are tracking the date on the Kimi K3 open weights page.

Also on this site: Qwen 3.8 pricing — Alibaba's answer to K3, previewed three days later with no per-token price at all — and GPT-5.6 pricing.

Sources

Figures on this page last checked against these sources on 2026-07-21. Vendors change pricing and specs without notice — if a number here disagrees with the vendor's own page, trust the vendor.