Kimi K3: Pricing, Weights and Deployment

Kimi K3 brings a 2.8-trillion-parameter model, native vision and a one-million-token context window. Its public rates, license and deployment choices deserve separate evaluation.

Kimi K3: Pricing, Weights and Deployment — AI

Moonshot introduced Kimi K3 in July 2026, with downloadable weights following on July 27. The model is a substantial open-weight release. Its technical report, rate card and license are more useful decision inputs than an unsourced snapshot of a fast-moving leaderboard.

The price deserves a closer look than the headline. K3’s published $3 input and $15 output rates per million tokens are 70% below Fable 5’s $10/$50 launch rates. That is a substantial unit-price difference; whether it produces a cheaper accepted result depends on the workload, caching and retries.

Kimi K3Primary-source specification
Model scale2.8T total parameters; mixture of experts
Context and input1M-token context; native vision
Published first-party rates /1M tokens$3 cache-miss input; $0.30 cache-hit input; $15 output
WeightsAvailable under the custom Kimi K3 License; review its conditions

Who's actually behind Kimi

Kimi is the product; Moonshot AI is the company — a Beijing startup founded in March 2023. The name to know is Yang Zhilin, the CEO: a Tsinghua and Carnegie Mellon researcher who co-authored Transformer-XL and XLNet and spent time at Google Brain and Meta before starting Moonshot. This is not a copy shop, and it isn't a team that discovered transformers through a venture-capital slide deck.

The relevant procurement question is the deployment you choose. Check the provider’s data terms, processing region, retention and access controls; for self-hosting, check the license and operating requirements. An investor list cannot answer those questions for every route.

What K3 actually is

Moonshot describes K3 as a 2.8-trillion-parameter mixture-of-experts model with native vision and a one-million-token context window. It combines Kimi Delta Attention and Attention Residuals, activating 16 of 896 experts. The launch used max thinking effort by default; architecture-level efficiency claims are vendor measurements, not guaranteed application speedups.

The technical report provides benchmark results and harness conditions. Moonshot’s own overview says K3’s overall performance still trails Fable 5 and Sol in its evaluation suite while performing strongly against other tested models. A lead on one frontend preference benchmark would not establish a general lead.

A leaderboard measures a bounded task under particular conditions. Evaluate K3 on representative work, including long sessions, failed attempts and verification, before transferring a reported score to production expectations.

The coding story — Kimi Code and the swarm

Kimi Code offers a terminal workflow for K3, and Kimi also provides agent orchestration features. Product-level concurrency and plan capacity should be checked separately from the model specification. Parallelism helps divisible work but can add coordination cost and reproduce a shared mistaken assumption.

Vendor coding comparisons need their exact harness, effort and task definitions. The K3 technical blog explains that some evaluations use different harnesses or calibrated tasks, and that fallback behavior affected some comparator runs. Treat those qualifications as part of the result.

What it costs

Moonshot’s K3 technical blog lists $3 per million cache-miss input tokens, $0.30 per million cache-hit input tokens and $15 per million output tokens. Against Fable 5’s $10/$50 launch rates, that is 70% lower on uncached input and output. Cache hits, generated volume and provider-specific terms still determine the actual task cost.

Subscription access is a separate product decision. Check the current regional offering for included usage, credits, concurrency and model access; a monthly fee is not a guaranteed amount of model work.

Access routeCheck separately
Subscription productCurrent seat terms, included allowance, overage and concurrency
First-party or third-party APIModel version, rate card, retention and region
Self-hosted weightsLicense, hardware, isolation and operating costs

Kimi vs Claude vs Codex — the honest read

Claude Code, Codex and Kimi Code combine models with different client workflows. Compare repository compatibility, tools, policy handling, accepted-result quality and the selected deployment’s data terms. No unnamed reviewer consensus establishes a universal quality or maturity ordering.

Routing across models is an option to test against a simpler single-model baseline. It adds maintenance and evaluation costs; use it where measured quality, availability or cost justifies that complexity.

Compare the complete task cost

A lower published rate is useful, but it is only one term in the cost of a verified task. Measure input, cache use, generated output, retries, tool charges and review effort against the same acceptance criteria. A model that needs more attempts can lose part of its unit-price advantage.

The decision is broader than a token-rate comparison. Reliability, licensing, data handling and developer experience all affect whether a model belongs in a production workflow. Evaluate those properties directly instead of inferring the vendor’s motives from its rate card.

Open, Chinese, and complicated

K3 weights are available under the custom Kimi K3 License, which includes commercial conditions rather than being plain MIT. Deployment is resource-intensive, but private hosting is a hardware and operational question, not logically impossible. Data handling depends on whether the chosen route is Moonshot, another provider or self-hosting; open weights alone do not establish privacy.

Behavior on politically sensitive prompts needs a K3-specific evaluation across relevant languages and deployment policies. A study of a predecessor does not establish K3’s responses or prove that a topic is irrelevant to every engineering workload.

K3 broadens the available open-weight model options. That is useful without turning a single benchmark into a geopolitical verdict. Keep the model evaluation, hosting decision and governance assessment separate.

So, is it worth it?

Consider K3 where its quality, cost, license and deployment requirements fit the workload. Compare it with Claude and Codex using the same acceptance criteria. Provider reputation and a leaderboard position are inputs to that decision, not substitutes for it.

DecisionEvidence to collect
Choose a modelAccepted-result quality, reliability and complete task cost
Choose a deploymentData terms, control requirements, license and infrastructure
Add model routingDemonstrable benefit over a simpler baseline, including maintenance

K3 adds a meaningful model option. Evaluate it before adding another production dependency, and retain a router only if its benefits exceed the operational complexity.