Kimi K3 brings 2.8T parameters to the open-model frontier — but not yet to your servers

Kimi K3 brings 2.8T parameters to the open-model frontier — but not yet to your servers

Moonshot AI’s Kimi K3 is a 2.8T-parameter multimodal model with a 1M-token context window and strong reported agentic benchmarks, but its open-weight release and independent validation are still pending.

A 2.8T-parameter frontier model, with an asterisk on “open”

Moonshot AI released Kimi K3 on July 16, describing it as a 2.8-trillion-parameter model with native vision, a 1-million-token context window, and new Kimi Delta Attention and Attention Residuals architecture. Moonshot calls it the first open 3T-class model, but the distinction matters: the full weights are promised by July 27, not available at launch. 1
K3 is already accessible through Kimi.com, Kimi Work, Kimi Code, and the Kimi API. The API is OpenAI-SDK compatible and costs $3 per million input tokens, $0.30 for cached input, and $15 per million output tokens. It currently runs at maximum thinking effort by default; Moonshot says lower-effort modes will arrive in later updates. 2 1
The launch claims are strongest in agentic search and coding. Moonshot’s benchmark table reports 91.2 on BrowseComp, above the listed 90.4 for GPT-5.6 Sol, while VentureBeat reports K3 at No. 1 in Arena.AI’s Frontend Code Arena with an Elo score of 1,679. The comparisons need a methodology warning: the official table mixes KimiCode, Claude Code, and Codex harnesses, and all K3 results use maximum reasoning effort. Its BrowseComp footnote also reports 90.4 when K3 is tested with the full 1-million-token window and no context management. 3 4
Kimi K3 coding benchmark comparison
Kimi K3’s coding results in Moonshot’s benchmark table, including Terminal Bench 2.1, Program Bench, FrontierSWE, and SWE Marathon. 3 Artificial Analysis gives K3 a 57 Intelligence Index score, reports 62 output tokens per second versus a 72.2 median for similarly priced reasoning models, and says it used 130 million output tokens during evaluation versus a 63 million median. It also currently lists the model as proprietary because the weights have not yet been released. 5
That gap between frontier results and deployment friction is the real story. K3’s official limitations include instability when an agent harness does not preserve the model’s thinking history, and a tendency toward excessive proactiveness when instructions are ambiguous. Moonshot also recommends deployments with 64 or more accelerators. If the July 27 weight release arrives with the promised technical report and usable inference tooling, K3 could give open-model developers a serious frontier base. Until then, it is best understood as a powerful hosted model with an open-weight release still pending, not a self-hostable alternative today. 1

相似内容

  • 登录后可发表评论。
More from this channel