Best of your X follows: Kimi K3, benchmark skepticism, and open weights

Best of your X follows: Kimi K3, benchmark skepticism, and open weights

Six original posts examine Kimi K3's release, the limits of Arena scores, open-weight business and policy, a reported court-workflow result, and Turso's broader database direction.

Kimi K3 is the day's shared object, but the useful posts disagree on what to measure: release specs, Arena scores, business viability, governance, and actual work completed.

Model release and evaluation

Kimi K3 arrives with a price and a warning label

Simon Willison, creator of Datasette and co-creator of Django, examined Moonshot AI's Kimi K3 in a July 16 post.
Moonshot describes K3 as a 2.8-trillion-parameter model, available through its website and API, with an open-weight release promised by July 27. 1
The release's self-reported benchmarks mostly beat Claude Opus 4.8 max and GPT-5.5 high; Simon also cites an Artificial Analysis long-horizon score of 1547 and a $0.94 cost per task. 1
His own SVG test used 13,241 reasoning tokens and cost $0.25, a useful reminder that long-horizon tool use matters more than a neat single-task demo. 1
Simon’s launch-and-testing post:
Cargando tarjeta de contenido…

An Arena win is not a general capability report

Ethan Mollick, a Wharton professor who studies AI, innovation, and startups, pushed back on how quickly Kimi K3's Arena result was being read.
He calls K3 a very good model but says people are overindexing on its Arena score, recalling the earlier Llama 4 episode. 2
His specific objection is that user-judged Elo is limited: front-end text chat is subjective and can be tuned toward what Arena users prefer. 2
The practical read is narrow but important: treat the score as one signal about a particular interaction setting, not as a complete model card. 2
Mollick’s benchmark caution:
Cargando tarjeta de contenido…

Open weights, business, and policy

The Chinese open-weight race is becoming a business problem

Mollick's second K3 post is less about a leaderboard than about whether several open-weight labs can support the economics of frontier-model competition.
He says Chinese open-weight models are now very good and gives a current ordering in his view: K3 ahead of GLM-5.2, which he places ahead of DeepSeek v4. 3
He also describes the companies behind them as becoming large and valuable businesses, making sustained competition the harder question. 3
That ordering is Mollick's comparison, not an independently verified ranking; the open question is whether all of these businesses can stay in the race. 3
The business angle:
Cargando tarjeta de contenido…

Open-weight models expose a governance gap

Mollick asks how pre-clearance could work for open-weight models, using Kimi K3 as the immediate case.
He notes that K3 did not yet have a model card in his post and that its weights were expected in a couple of weeks; he also points out that downloaded weights cannot simply be recalled. 4
His proposed policy lever is indirect: governments could restrict companies that serve their residents from using unvetted models, even if the weights remain freely downloadable. 4
He ends at international cooperation on model vetting, while acknowledging that the policy is still emergent and no one knows how it will work. 4
The governance question in full:
Cargando tarjeta de contenido…

Applied AI and developer infrastructure

A reported 6% throughput lift in Pakistani courts

Mollick also posted a compact case claim about a GPT-4-powered assistant used by Pakistani judges.
His post says the assistant increased the number of cases judges saw by 6%, with no impact on quality. 5
The result points to a throughput question rather than a replacement claim: what happens when a professional tool increases capacity without changing the measured quality of work? 5
The post includes no study link or method, so the 6% should be read as a reported result, not a validated effect size. 5
The source post:
Cargando tarjeta de contenido…

Turso is pitching beyond SQLite

Simon Willison says Turso is expanding beyond SQLite into a foundation on which multiple database-compatibility layers can be built.
That makes the project more than another SQLite deployment option, at least in his reading: the pitch is a portability substrate for different database interfaces. 6
The post does not include architecture, supported layers, or launch documentation, so the useful signal is the direction of travel rather than a confirmed feature set. 6
His short developer-tooling signal:
Cargando tarjeta de contenido…

Contenido relacionado

  • Inicia sesión para comentar.
More from this channel