
Best of your X follows: Kimi K3, benchmark skepticism, and open weights
Six original posts examine Kimi K3's release, the limits of Arena scores, open-weight business and policy, a reported court-workflow result, and Turso's broader database direction.
Kimi K3 is the day's shared object, but the useful posts disagree on what to measure: release specs, Arena scores, business viability, governance, and actual work completed.
Model release and evaluation
Kimi K3 arrives with a price and a warning label
Simon Willison, creator of Datasette and co-creator of Django, examined Moonshot AI's Kimi K3 in a July 16 post.
Moonshot describes K3 as a 2.8-trillion-parameter model, available through its website and API, with an open-weight release promised by July 27. 1
The release's self-reported benchmarks mostly beat Claude Opus 4.8 max and GPT-5.5 high; Simon also cites an Artificial Analysis long-horizon score of 1547 and a $0.94 cost per task. 1
His own SVG test used 13,241 reasoning tokens and cost $0.25, a useful reminder that long-horizon tool use matters more than a neat single-task demo. 1
Simon’s launch-and-testing post:
콘텐츠 카드를 불러오는 중…
An Arena win is not a general capability report
Ethan Mollick, a Wharton professor who studies AI, innovation, and startups, pushed back on how quickly Kimi K3's Arena result was being read.
He calls K3 a very good model but says people are overindexing on its Arena score, recalling the earlier Llama 4 episode. 2
His specific objection is that user-judged Elo is limited: front-end text chat is subjective and can be tuned toward what Arena users prefer. 2
The practical read is narrow but important: treat the score as one signal about a particular interaction setting, not as a complete model card. 2
Mollick’s benchmark caution:
콘텐츠 카드를 불러오는 중…
Open weights, business, and policy
The Chinese open-weight race is becoming a business problem
Mollick's second K3 post is less about a leaderboard than about whether several open-weight labs can support the economics of frontier-model competition.
He says Chinese open-weight models are now very good and gives a current ordering in his view: K3 ahead of GLM-5.2, which he places ahead of DeepSeek v4. 3
He also describes the companies behind them as becoming large and valuable businesses, making sustained competition the harder question. 3
That ordering is Mollick's comparison, not an independently verified ranking; the open question is whether all of these businesses can stay in the race. 3
The business angle:
콘텐츠 카드를 불러오는 중…
Open-weight models expose a governance gap
Mollick asks how pre-clearance could work for open-weight models, using Kimi K3 as the immediate case.
He notes that K3 did not yet have a model card in his post and that its weights were expected in a couple of weeks; he also points out that downloaded weights cannot simply be recalled. 4
His proposed policy lever is indirect: governments could restrict companies that serve their residents from using unvetted models, even if the weights remain freely downloadable. 4
He ends at international cooperation on model vetting, while acknowledging that the policy is still emergent and no one knows how it will work. 4
The governance question in full:
콘텐츠 카드를 불러오는 중…
Applied AI and developer infrastructure
A reported 6% throughput lift in Pakistani courts
Mollick also posted a compact case claim about a GPT-4-powered assistant used by Pakistani judges.
His post says the assistant increased the number of cases judges saw by 6%, with no impact on quality. 5
The result points to a throughput question rather than a replacement claim: what happens when a professional tool increases capacity without changing the measured quality of work? 5
The post includes no study link or method, so the 6% should be read as a reported result, not a validated effect size. 5
The source post:
콘텐츠 카드를 불러오는 중…
Turso is pitching beyond SQLite
Simon Willison says Turso is expanding beyond SQLite into a foundation on which multiple database-compatibility layers can be built.
That makes the project more than another SQLite deployment option, at least in his reading: the pitch is a portability substrate for different database interfaces. 6
The post does not include architecture, supported layers, or launch documentation, so the useful signal is the direction of travel rather than a confirmed feature set. 6
His short developer-tooling signal:
콘텐츠 카드를 불러오는 중…
관련 콘텐츠
- 로그인하면 댓글을 작성할 수 있습니다.
More from this channel›
- Gemini's cheaper agents, Claude Tag's 65% PRs, and a new prompting rule
- AI/tech signals: rare-disease grants, cloud agents, and model taste
- Thin X day: Kimi's language split, agent compilers, and open-source friction
- Best of your X follows: Codex in the wild, cyber defense, and new research bets
- Best of your X follows: live assistants, bioresilience, and frontendmaxxing
- Best of your X follows: agent habits, research bottlenecks, and anti-slop taste
- Best of your X follows: cache-friendly uvx, Canadian research, and agent UX
- Best of your X follows: Morpheus, Sol builds, and AI roadmaps
