Kimi K3: Pricing, Weights and Deployment
Kimi K3 brings a 2.8-trillion-parameter model, native vision and a one-million-token context window. Its public rates, license and deployment choices deserve separate evaluation.
Moonshot introduced Kimi K3 in July 2026, with downloadable weights following on July 27. The model is a substantial open-weight release. Its technical report, rate card and license are more useful decision inputs than an unsourced snapshot of a fast-moving leaderboard.
The price deserves a closer look than the headline. K3’s published $3 input and $15 output rates per million tokens are 70% below Fable 5’s $10/$50 launch rates. That is a substantial unit-price difference; whether it produces a cheaper accepted result depends on the workload, caching and retries.
| Kimi K3 | Primary-source specification |
|---|---|
| Model scale | 2.8T total parameters; mixture of experts |
| Context and input | 1M-token context; native vision |
| Published first-party rates /1M tokens | $3 cache-miss input; $0.30 cache-hit input; $15 output |
| Weights | Available under the custom Kimi K3 License; review its conditions |
Who's actually behind Kimi
Kimi is the product; Moonshot AI is the company — a Beijing startup founded in March 2023. The name to know is Yang Zhilin, the CEO: a Tsinghua and Carnegie Mellon researcher who co-authored Transformer-XL and XLNet and spent time at Google Brain and Meta before starting Moonshot. This is not a copy shop, and it isn't a team that discovered transformers through a venture-capital slide deck.
The relevant procurement question is the deployment you choose. Check the provider’s data terms, processing region, retention and access controls; for self-hosting, check the license and operating requirements. An investor list cannot answer those questions for every route.
What K3 actually is
Moonshot describes K3 as a 2.8-trillion-parameter mixture-of-experts model with native vision and a one-million-token context window. It combines Kimi Delta Attention and Attention Residuals, activating 16 of 896 experts. The launch used max thinking effort by default; architecture-level efficiency claims are vendor measurements, not guaranteed application speedups.
The technical report provides benchmark results and harness conditions. Moonshot’s own overview says K3’s overall performance still trails Fable 5 and Sol in its evaluation suite while performing strongly against other tested models. A lead on one frontend preference benchmark would not establish a general lead.
A leaderboard measures a bounded task under particular conditions. Evaluate K3 on representative work, including long sessions, failed attempts and verification, before transferring a reported score to production expectations.
The coding story — Kimi Code and the swarm
Kimi Code offers a terminal workflow for K3, and Kimi also provides agent orchestration features. Product-level concurrency and plan capacity should be checked separately from the model specification. Parallelism helps divisible work but can add coordination cost and reproduce a shared mistaken assumption.
Vendor coding comparisons need their exact harness, effort and task definitions. The K3 technical blog explains that some evaluations use different harnesses or calibrated tasks, and that fallback behavior affected some comparator runs. Treat those qualifications as part of the result.
What it costs
Moonshot’s K3 technical blog lists $3 per million cache-miss input tokens, $0.30 per million cache-hit input tokens and $15 per million output tokens. Against Fable 5’s $10/$50 launch rates, that is 70% lower on uncached input and output. Cache hits, generated volume and provider-specific terms still determine the actual task cost.
Subscription access is a separate product decision. Check the current regional offering for included usage, credits, concurrency and model access; a monthly fee is not a guaranteed amount of model work.
| Access route | Check separately |
|---|---|
| Subscription product | Current seat terms, included allowance, overage and concurrency |
| First-party or third-party API | Model version, rate card, retention and region |
| Self-hosted weights | License, hardware, isolation and operating costs |
Kimi vs Claude vs Codex — the honest read
Claude Code, Codex and Kimi Code combine models with different client workflows. Compare repository compatibility, tools, policy handling, accepted-result quality and the selected deployment’s data terms. No unnamed reviewer consensus establishes a universal quality or maturity ordering.
Routing across models is an option to test against a simpler single-model baseline. It adds maintenance and evaluation costs; use it where measured quality, availability or cost justifies that complexity.
Compare the complete task cost
A lower published rate is useful, but it is only one term in the cost of a verified task. Measure input, cache use, generated output, retries, tool charges and review effort against the same acceptance criteria. A model that needs more attempts can lose part of its unit-price advantage.
The decision is broader than a token-rate comparison. Reliability, licensing, data handling and developer experience all affect whether a model belongs in a production workflow. Evaluate those properties directly instead of inferring the vendor’s motives from its rate card.
Open, Chinese, and complicated
K3 weights are available under the custom Kimi K3 License, which includes commercial conditions rather than being plain MIT. Deployment is resource-intensive, but private hosting is a hardware and operational question, not logically impossible. Data handling depends on whether the chosen route is Moonshot, another provider or self-hosting; open weights alone do not establish privacy.
Behavior on politically sensitive prompts needs a K3-specific evaluation across relevant languages and deployment policies. A study of a predecessor does not establish K3’s responses or prove that a topic is irrelevant to every engineering workload.
K3 broadens the available open-weight model options. That is useful without turning a single benchmark into a geopolitical verdict. Keep the model evaluation, hosting decision and governance assessment separate.
So, is it worth it?
Consider K3 where its quality, cost, license and deployment requirements fit the workload. Compare it with Claude and Codex using the same acceptance criteria. Provider reputation and a leaderboard position are inputs to that decision, not substitutes for it.
| Decision | Evidence to collect |
|---|---|
| Choose a model | Accepted-result quality, reliability and complete task cost |
| Choose a deployment | Data terms, control requirements, license and infrastructure |
| Add model routing | Demonstrable benefit over a simpler baseline, including maintenance |
K3 adds a meaningful model option. Evaluate it before adding another production dependency, and retain a router only if its benefits exceed the operational complexity.