
🔮 Will Kimi K3 change the economics of AI?|Kimi K3 会改变 AI 的经济学吗?|英文原文 + 中文翻译
完整呈现 Exponential View 2026 年 7 月 23 日公开可读的英文原文与中文翻译,追踪 Kimi K3、token 价格弹性、开源模型推理成本和中国 AI 实验室效率;付费墙后的内容不补写。
English original
🔮 Will Kimi K3 change the economics of AI?
The pressure is on
Jul 23, 2026
∙ Paid

Kimi K3 has caused quite an uproar since its release last week. It’s the first time a Chinese model has taken the lead on the frontend Code Arena benchmark. And that’s three months since Moonshot AI’s previous impressive flagship model, Kimi K2.6, was released. 1
Following in Moonshot’s steps, Alibaba announced over the weekend that Qwen3.8 – a 2.4 trillion-parameter model – is coming soon, and unlike its last release, this one will be an open-weight model. No benchmarks or further details have been released as of yet. 1
Open models are now estimated to be 4-7 months behind the frontier in cyber capabilities, down from 6-10 months in 2025. And despite compute constraints, efficiency improvements mean these labs are doing more with less. Comparing the compute availability and model performance between US labs and Chinese labs, we estimated Chinese labs to be getting 4-7x more out of their compute. 1
Hannah Petrovic and I spent some time with the Moonshot AI and Alibaba teams in China back in April and May, and we’ve had time to think about the economics of open-source models and how they affect the entire ecosystem. 1
Does Kimi K3 break the economic case for AI?
Some have claimed that Kimi K3’s performance breaks the economic case for AI as it lowers the cost to complete various tasks at frontier standards. For instance, Microsoft engineers are reportedly testing whether Kimi K3 can be used within Copilot. 1
We don’t think this is the case, and in today’s post we’ll work through what might happen next. 1
In The State of the AI Economy report, we found that token usage is elastic across providers. This means that every drop in token price leads to a larger increase in token volume, more than offsetting the difference. 1

For every 10% price cut, token consumption rises 12-18%. A paper by Demirer et al, found a similar effect: a 10% price cut resulted in an 11% or so increase in volumes, which economists call an elasticity of -1.11. 1
The net effect is a rise in total token spend. But note that the effect is a weak one, not the cantering Jevons’ paradox sometimes presented. Reality might tilt the scales further in favor of more, not less, demand. Workflows are becoming more token-intensive as we rely on reasoning models and verification and approval loops. And the early evidence suggests that firms that adopt AI early tend to increase their relative spend alongside growing headcount. These effects might be short-term elasticities rather than ones that can be sustained for decades, but for now they indicate that falling prices increase volumes and, with that, revenue. 1
Flowing down the stack
The model weights may be free, but the inference is not. Kimi K3 has 2.8 trillion parameters. The weights alone occupy 1.4 TB. It needs to be served on something like a 72-GPU NVIDIA GB200 NVL72 rack or equivalent. That’ll cost $3-4 million to buy and install. Operating it consumes about 120 kW continuously, over a million kWh per year, before you consider networking, storage, cooling, and humans. If you rented these in the open market, it would cost about $7 million a year. 1
But for infrastructure providers, the economics of hosting open-source models can be very attractive compared to serving closed-source models. A simple way to understand this is to think of the hyperscaler as needing to pay a license fee for a closed-source model but not for an open-source one1. 1
Public access boundaryThe publicly readable page ends here with the notice "This post is for paid subscribers." The remaining article is not accessible on the official page and is not translated or supplemented.
中文翻译
🔮 Kimi K3 会改变 AI 的经济学吗?
压力正在加大
2026 年 7 月 23 日
∙ 付费文章
今年早些时候我们访问 Moonshot AI 办公室时拍摄的照片 1
Kimi K3 上周发布后引起了不小的反响。这是中国模型首次登上 Frontend Code Arena 基准测试 的榜首。Moonshot AI 上一款表现同样亮眼的旗舰模型 Kimi K2.6,则是在三个月前发布的。1
阿里巴巴沿着 Moonshot 的步伐,在周末宣布,Qwen3.8,一个拥有 2.4 万亿参数的模型,即将推出。与上一次发布不同,这次会是一个开放权重模型。截至目前,官方还没有发布基准成绩或更多细节。1
目前估计,开放模型在网络安全能力上落后前沿模型 4 至 7 个月,2025 年这一差距是 6 至 10 个月。尽管算力受限,效率提升让这些实验室用更少资源做更多事情。比较美国与中国实验室的可用算力和模型表现后,我们估计中国实验室从算力中得到的产出多 4 至 7 倍。1
今年 4 月和 5 月,我与 Hannah Petrovic 在中国和 Moonshot AI、阿里巴巴的团队交流了一段时间,也有时间思考开源模型的经济学,以及它们如何影响整个生态系统。1
Kimi K3 会打破 AI 的经济基础吗?
有人认为,Kimi K3 按前沿标准完成各种任务的成本下降,因此它的表现打破了 AI 的经济基础。相关讨论包括成本下降以及以前沿水平完成任务;据报道,微软工程师正在测试 Kimi K3 是否能用于 Copilot。1
我们不这么认为。本文将梳理接下来可能发生什么。1
在《AI 经济现状》报告中,我们发现,不同服务商之间的 token 使用量具有弹性。这意味着 token 价格每下降一点,token 使用量就会更大幅增加,增量足以抵消价格差异。1
净效果是 token 总支出上升。但要注意,这个效应并不强,不是有时被描述的那种奔腾式 Jevons 悖论。现实可能还会进一步偏向更多、而不是更少的需求。随着我们依赖推理模型、验证与审批循环,工作流正在消耗更多 token。早期证据表明,较早采用 AI 的企业在员工人数增长的同时,往往也会增加相对支出。这些效应可能只是短期弹性,无法维持数十年,但目前它们表明,价格下降会增加使用量,并随之增加收入。1
成本沿着技术栈向下传导
模型权重可以免费,推理却不是。Kimi K3 有 2.8 万亿个参数,仅权重就占据 1.4 TB 的空间。它需要部署在类似一台配备 72 块 GPU 的 NVIDIA GB200 NVL72 机架上,或使用同等设备。购买和安装这样的设备要花 300 万至 400 万美元。设备持续运行时的功耗约为 120 kW,每年超过 100 万 kWh;这还没有算入网络、存储、冷却和人员成本。如果在公开市场租用这样的设备,每年成本约为 700 万美元。1
公开内容边界官方页面在上一个段落后显示「This post is for paid subscribers」。其余文章内容在当前页面上不可访问,本文不翻译,也不补写。
Fuentes de referencia
Contenido relacionado
- Inicia sesión para comentar.