Moonshot AI released Kimi K3 on July 16, 2026, calling it the first open model to cross the 3-trillion-parameter class. The company's own benchmark table shows the model leading other open-weight systems while still sitting behind Anthropic's Claude Fable 5 and OpenAI's GPT 5.6 Sol on most measures.
Kimi K3 Reaches 2.8 Trillion Parameters as the Largest Open Model Yet
According to Moonshot's technical blog, Kimi K3 is a 2.8-trillion-parameter model built on two architectural changes the company calls Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), designed to improve how information moves across sequence length and model depth. The model uses a sparse Mixture-of-Experts setup — Stable LatentMoE — that activates only 16 of 896 experts on any given forward pass, and it ships with native vision support and a 1-million-token context window. Moonshot says the combined architectural changes produce roughly a 2.5x gain in scaling efficiency over its previous flagship, Kimi K2.
Full model weights are scheduled for release by July 27, 2026, with a technical report to follow separately. The model is already live through kimi.com, the Kimi Work desktop app, Kimi Code, and the Kimi API under the identifier kimi-k3.
Benchmark Scores Place K3 Behind Fable 5 and GPT 5.6 Sol, Ahead of Open Rivals
Moonshot's published benchmark table backs up its own framing: on comparable percentage-scale tests, Kimi K3 scores close to but consistently at or below Claude Fable 5 and GPT 5.6 Sol, while outperforming Claude Opus 4.8 on several of the same tests. On DeepSWE, K3 scored 67.5 against Fable 5's 70.0 and GPT 5.6 Sol's 73.0. On Terminal-Bench 2.1, the gap narrows to a fraction of a point, with K3 at 88.3 against GPT 5.6 Sol's 88.8. On GPQA-Diamond, a graduate-level science reasoning test, K3's 93.5 sits between Opus 4.8's 91.0 and GPT 5.6 Sol's 94.1.
The comparison carries a footnoted caveat worth weighing: Moonshot evaluated Kimi K3 largely through its own KimiCode harness, while GPT models were run through Codex and the Claude models through Claude Code or a third-party evaluation for Fable 5 that the footnotes say "may include fallback behavior" to Opus 4.8. That means the scores below reflect each vendor's preferred operating environment rather than one identical test harness across the board — a pattern also seen in GPT-5.6 Sol's own benchmark disclosures.
API Pricing Undercuts Rival Frontier Models by a Wide Margin
Beyond raw scores, Moonshot is pricing Kimi K3 aggressively for developers building on top of it. The Kimi API lists the model at $0.30 per million tokens for cache-hit input, $3.00 per million tokens for cache-miss input, and $15.00 per million tokens for output — with Moonshot noting that its Mooncake inference stack keeps cache-hit rates above 90% on typical coding workloads, meaning most real-world input traffic should land at the cheaper tier. That pricing structure positions K3 as a lower-cost option relative to the proprietary systems it trails on benchmarks, echoing the cost-versus-capability tradeoff already playing out between GPT-5.6 Sol and Claude Fable 5 and among open-weight competitors like Qwen 3.7 Max.
Moonshot Discloses Session-Switching Instability and a User-Experience Gap
Moonshot's release notes are unusually direct about where K3 falls short. The company warns that the model was trained with a "preserved thinking history" mode and that generation quality "may become highly unstable" if an agent harness fails to pass back the model's full prior reasoning, or if an ongoing session started with another model is switched mid-way to K3 — a compatibility issue the company says is verified only in its own Kimi Code harness. Moonshot also describes a tendency toward "excessive proactiveness," where the model may act on ambiguous instructions without checking in, and recommends teams add explicit behavioral constraints for applications that need tighter guardrails.
Most notably, Moonshot's own limitations section states plainly that Kimi K3 "exhibits a noticeable gap in user experience compared with Claude Fable 5 and GPT 5.6 Sol" — a rare instance of a lab naming its product's shortfall against named competitors rather than leaving it to independent testers. Separately, Moonshot describes internal tests where an early K3 build optimized GPU kernels and, in one case, designed a chip layout for a companion model, claiming the model "substantially outperformed" Opus 4.8, GPT 5.6 Sol, and GPT 5.5 on the kernel task. No numeric scores were published for that comparison, so it remains a vendor claim rather than a benchmarked result, unlike the scored table above.
For now, Kimi K3 lands as a capable, cheaper open alternative rather than a frontier leader: strong enough to top other open-weight systems on Moonshot's own numbers, but by the company's own account, still a step behind the two proprietary models it was built to chase.
Comments (0)
Please sign in to join the discussion.
No comments yet.
Be the first to share your perspective on this topic.