过去24小时大模型研究员 X 动态:GPT-5.6 Sol/Terra 的实际反馈,Inkling 的 ARC-AGI 进展与 Kimi K3 商用化

过去24小时大模型研究员 X 动态:GPT-5.6 Sol/Terra 的实际反馈,Inkling 的 ARC-AGI 进展与 Kimi K3 商用化

频道 450 个 researcher 账号全部完成时间线请求,窗口内确认 230 条新增去重帖,重点呈现 GPT-5.6 Sol/Terra 的安全与代码工作流反馈、Inkling 的 ARC-AGI 结果、Kimi K3 企业会员与第三方评测,以及本期没有确认的人员状态变化;这不是 X 全站全量结论。

过去 24 小时,频道 researcher roster 中的 450 个账号全部完成时间线请求,窗口内 85 个账号有命中,去重后得到 230 条新增原帖。重点集中在 OpenAI 的 GPT-5.6 Sol/Terra、Thinking Machines Lab 的 Inkling、Kimi K3 的企业会员与第三方评测,以及低延迟推理硬件。窗口为洛杉矶时间 2026 年 7 月 17 日 08:00 至 7 月 18 日 08:00。这不是 X 全站全量结论;365 个账号没有窗口命中,不等于账号没有发文,DeepSeek、ByteDance Seed、Cursor、01.AI、Mistral 等分类本轮没有命中。

先看重点

  • OpenAI 总裁兼联合创始人 Greg Brockman 称 GPT-5.6 Sol 在网络安全上处于当前领先位置,并说团队已经看到它用于发现和修复新漏洞的明显效果。OpenAI 账号 @steipete 则给出更具体的个人实测:把 Terra high 用在 GitHub review bot ClawSweeper 上,整体速度约快 40%,质量损失很小,而且比 5.5 更便宜。两者都来自个人账号,不能当作独立评测结论。
  • Thinking Machines Lab 的 technical staff Martin Ziqiao Ma 宣布,Inkling 在 ARC-AGI-1 和 ARC-AGI-2 上是截至当时得分最高的开放权重模型。团队成员转发的 ARC-AGI 结果给出 ARC-AGI-1 79.5%、每题 0.30 美元,ARC-AGI-2 36.5%、每题 0.64 美元。这里的性能与成本来自项目方转发,仍应和独立复测分开看。
  • Kimi 官方账号新增企业会员,5 个席位起、年付,并提供企业隐私与技术支持;同时转发了 SpreadsheetBench 2 第一名和 VoxelBench 第三名的消息。前者是产品商业化动作,后两者是官方放大的第三方榜单,不等于 Kimi 自己完成了独立评测。
  • 今日没有确认的人员动态。Anthropic 的 Ethan Perez 转发了 Frontier AI Security Residency 招募信息,但这是项目公告,不是他本人入职、离职或角色变化。

按公司看

OpenAI:Sol/Terra 的网络安全、代码审查与 Codex 使用反馈

OpenAI 归类账号 193 个,窗口内 40 个账号有命中,共 135 条原帖。高信息密度内容主要来自产品实测和一线使用反馈,研究相关性强弱差异很大。
Greg Brockman(@gdb,OpenAI 总裁兼联合创始人)在 21:04 发文称 GPT-5.6 Sol 可用于发现和修复新漏洞,随后又发了「Sol gets the thing done」和「don't sleep on terra!」两条短帖。前一条是明确的网络安全产品表态,后两条没有给出评测方法或数字。1 2 3
@steipete 的反馈更接近真实工作流:他把 Terra high 切到 GitHub review bot ClawSweeper 后,称整体速度约快 40%,质量损失很小,成本也低于 5.5;针对自己的 issue/code review 场景,他还说 Terra high 明显优于 Sol low。另一条帖子记录了 Codex 通过浏览器和 computer use 打开 Chrome、进入 PR、点击评论并上传图片的过程,他把任务放在虚拟机里,避免抢占应用焦点。这里的「约快 40%」和「明显更好」都是个人使用报告,不能替代统一 benchmark。4 5
前 OpenAI RL 负责人、现 Core Automation CEO @MillionInt 讨论了两个更底层的问题:算力会流向高毛利产品,AI 公司更应该回答「为什么客户愿意为你的 AI 支付更高毛利」,而不是只回答算力从哪里来;他还批评 RLHF 评分容易把「听起来聪明」误当成「真的聪明」。这是个人判断,不代表 OpenAI 当前立场,但分别对应算力融资和后训练评价两个持续出现的约束。6 7
OpenAI 方向的优先层命中还包括:前 xAI PreTraining lead、曾任 OpenAI 的 @archanfel_anoth 转发祝贺 Kimi;曾在 OpenAI 做多模态与 RL、现 Meta Superintelligence Labs 做 RL/post-training/agents 的 @shuchaobi;现 Google DeepMind、曾在 xAI 做 reasoning、曾任 OpenAI MTS 的 @TianfuF;以及 @andrew_n_carr、@btaylor、@jeremyli__、@ypatil125 等履历中带有 OpenAI 科学、董事会、研究或创业身份的账号。它们本窗口的帖子多数是短评、转发或与原任职机构无直接关系,完整链接保留在文末索引中。

Anthropic:Claude Science、MCP agent 与 Fable 5 的产品分发

Anthropic 归类账号 47 个,窗口内 12 个账号有命中,共 15 条原帖。
Claude Science 的创作者、Anthropic MTS @AlecTPhD 说,他看到领域专家使用 Claude Science 的真实案例。信息没有展开具体任务或实验结果,但可以确认该账号在继续放大科学工作流方向。平台负责人 @katelyn_lesse 则建议用 MCP 让 agent 调用其他 agent,这条更接近 agent 编排的产品实践,而不是模型能力评测。8 9
@rishicomplex 转发 Claude 官方通知:Fable 5 将从 7 月 20 日起纳入 Max 和 Team Premium,限额为对应方案的 50%。这条是产品分发信息,原帖没有在本账号内容中给出价格或完整限额表。10
Alignment team lead @EthanJPerez 转发了 Frontier AI Security Residency 招募,项目面向前沿 AI 安全方向。它可以作为 Anthropic 关注安全人才供给的信号,但不能写成 Ethan 本人的职位变化。11

Google DeepMind:从优化讨论到开放模型呼吁

Google DeepMind 归类账号 71 个,窗口内 7 个账号有命中,共 26 条原帖。
前 Google DeepMind Research Scientist、现加州大学伯克利分校教授 Jason Lee(@jasondeanlee)连续发文讨论模型的表达能力与优化。他说所有模型大体上有相同的表达能力和泛化界限,真正的问题在优化;在另一个回复里,他又说对固定问题,模型可能会超过绝大多数数学家。这两条是个人理论判断,没有给出证明或实验设定,适合当作研究观点追踪,不应当作已验证结论。12 13
@jasondeanlee 还转发了关于取消 Codex 5 小时限制后周限额变成新 5 小时限制的反馈,并自己补充说 5.6 的周限额又成了新的瓶颈。这是账号对 OpenAI 产品变化的评论,不是 Google DeepMind 的官方消息。14
前 DeepMind、现 Sunday Robotics 联合创始人兼 CEO @tonyzzhao 的 8 条更新大多是转发,涉及机器人可靠性、家庭机器人、记忆开发者和硬件项目;前 Google DeepMind 的 @MengdiWang10 有 2 条转发。Google DeepMind 官方账号 @GoogleDeepMind 转发了 Weather Lab 的重大更新。它们进入本期完整索引,但没有被当作大模型研究结论展开。

Meta AI:外部评测转发,信号集中在模型对比

Meta AI 归类账号 17 个,窗口内 3 个账号有命中,共 3 条原帖。
@alexandr_wang 转发了对 Muse Spark 1.1 的正面评价,原帖称其表现接近 Opus 4.8、优于 Grok 4;@YifeiZhou02 转发了一个名为 schema 的 harness,原帖称在 Opus 4.8 与 Fable 5 上达到 99% RHAE、在 GPT-5.6 Sol 上达到 95.35%。这些都是转发内容中的第三方说法,本期不把它们改写成 Meta 官方评测。15 16

Thinking Machines Lab:Inkling 的 ARC-AGI 结果与 pretraining 讨论

Thinking Machines Lab 归类账号 15 个,窗口内 6 个账号有命中,共 15 条原帖。
technical staff Martin Ziqiao Ma(@ziqiao_ma)直接发文称 Inkling 是 ARC-AGI-1 与 ARC-AGI-2 截至当时得分最高的开放权重模型。团队成员 @shizhediao 和 @liliang_ren 转发的 ARC-AGI 结果写出两组分数:ARC-AGI-1 为 79.5%、每题 0.30 美元,ARC-AGI-2 为 36.5%、每题 0.64 美元。分数来自转发的 ARC Prize 帖子,不能和 Ziqiao 的团队声明混成一份独立验证。17 18 19
pretraining 研究员 @SonglinYang4 在窗口内转发了关于 Kimi K3 pretrain 的快速测试,原帖估计其表现介于 Opus 4 和 Opus 4.5 之间,约为 10 分钟级别的差距;这不是 Thinking Machines Lab 对 Kimi 的正式评测。她还转发了关于测试时计算深度、Kimi 是否遵循 scaling law 的讨论,以及 Inkling 的 ARC-AGI 结果。这里最值得持续观察的是开放权重模型之间的预训练与测试时计算比较,但目前公开材料仍主要是个人快速测试和转发。20 21

Cognition:FrontierCode leaderboard 上线

Cognition 归类账号 10 个,窗口内 5 个账号有命中,共 13 条原帖。
@BenPan 转发 Cognition 官方消息,FrontierCode leaderboard 已上线,页面专门追踪哪些模型能写出开发者实际愿意合并的代码;@jeffwang 转发 Devin Desktop 消息,称 Inkling 和 Grok 4.5 已可在 Devin Desktop 与 CLI 使用。两个信号分别对应代码评测入口和模型接入面,但原帖没有提供榜单排名或测试样本,不能进一步推导模型优劣。22 23

Moonshot/Kimi:企业会员上线,K3 继续进入应用验证

Moonshot/Kimi 归类账号 8 个,窗口内 3 个账号有命中,共 8 条原帖。
Kimi 官方账号发布 Business Membership,5 个席位起、年付,支持企业银行转账和自助开票,提供企业级数据隐私与技术支持;需要超过 20 个席位时,可以购买多份或联系商务团队。它是本窗口最明确的商业化更新。24
同一账号转发 AfterQuery 的 SpreadsheetBench 2 结果,称 Kimi K3 第一且超过 Fable 5;随后转发 VoxelBench 第三名的结果,称其只落后 Fable 100 多 Elo,且相比 K2.6 从第 28 名大幅提升。两条都是 Kimi 官方转发的第三方榜单,读者应把「官方放大」与「独立测评」分开。25 26
Kimi staff @crystalsssup 说团队经常收到外部反馈,但研究团队的回答是「But this is not AGI」,同时转发了 K3 在 Windows XP Simulator 和 SpreadsheetBench 2 上的表现。这条更像团队文化与能力边界的个人观察,不能当作官方路线声明。27

Google Brain:Andrew Ng 继续讲低延迟推理

Google Brain 归类账号只有 1 个,窗口内命中 1 个账号、1 条原帖。前 Google Brain 负责人 Andrew Ng 发布了一门与 Cerebras 合作的课程,主题是让 LLM 应用更快响应。课程把瓶颈解释为模型生成时需要把权重从内存搬到计算单元,介绍 Cerebras Wafer-Scale Engine 如何缩短这段移动,并把实时翻译、语音 agent 和多步工作流作为应用例子。28

其它公司与低研究相关性更新

Alibaba Qwen 归类账号 7 个,命中 1 个账号、2 条原帖,主要是对 Kimi 创始人和发布的转发,以及一条与研究无关的中文转发。Cohere 归类账号 3 个,命中 1 个账号、1 条原帖,内容是用 Claude Code 通过 DoorDash CLI 点寿司,属于 agent 使用趣闻,不是 Cohere 研究更新。xAI 归类账号 56 个,命中 6 个账号、11 条原帖,主要是 Pi 开源贡献者、Grok Build 仓库上传行为和数据公司讨论;本窗口没有来自 Igor Babuschkin 或 Elon Musk 的研究路线新帖。29 30 31
01.AI、ByteDance Seed、Cursor、DeepSeek、Mistral 本窗口没有命中。这里的含义是授权时间线接口返回的 450 个 roster 账号中,没有落在本窗口的新增原帖,不是这些公司在 X 上没有任何公开活动。

人员动态

本窗口没有找到 roster 账号本人或所属公司发布的、能够确认入职、离职、休假、创业、角色变化或新实验室加入的公告。
Anthropic 的 @EthanJPerez 转发 Frontier AI Security Residency 招募,属于项目招募,不是个人状态变化。@TianfuF 转发了「Why can Kimi ship K3?」讨论,内容讲的是模型发布原因,也没有出现本人加入 Kimi 或离开现团队的表述。两条都不进入人员动态栏。11 32

覆盖审计

公司归类roster 账号窗口活跃账号新增去重帖
OpenAI19340135
Google DeepMind71726
Anthropic471215
xAI56611
Meta AI1733
Thinking Machines Lab15615
Cognition10513
Moonshot/Kimi838
Cursor800
Alibaba Qwen712
DeepSeek600
ByteDance Seed500
Cohere311
01.AI200
Google Brain111
Mistral100
合计45085230
本轮 450 个账号的批量时间线请求均返回成功,没有失败账号。窗口外和历史文章中的原帖 URL 已剔除;本期 230 条与最近 6 篇文章正文中已展开的 91 个 X 原帖 URL 没有重复。由于授权范围是频道创建者关注列表中的 450 个账号,本文不能推断 X 全站发文情况;智谱、MiniMax、百川智能、阶跃星辰也没有稳定的个人 X 账号覆盖。

今日研究员 Spotlight

1. Greg Brockman(@gdb)

OpenAI 总裁兼联合创始人,公开身份更偏公司与产品决策;本窗口直接谈到 GPT-5.6 Sol 在网络安全中的漏洞发现和修复。长期关注模型能力如何进入真实系统防护,而不是只看榜单。最近动态是他对 Sol 和 Terra 的连续短帖。1

2. Jerry Tworek(@MillionInt)

他曾任 OpenAI RL 副总裁,简介列出的方向包括 reasoning models、o1/o3、GPT-4、Codex 和机器人 RL,目前经营 Core Automation。关注后训练评价、算力市场和高毛利 AI 产品。最近他批评 RLHF 评分把复杂措辞当成智能,并讨论算力向高毛利产品流动。7

3. Lucas Beyer(@giffmana)

他的简介写明自己是研究员,现于 Meta,曾在 OpenAI、DeepMind 和 Google Brain 工作。研究兴趣在模型与系统工程的交界处,最近谈到 TPU Pod 作为高连接加速器群组的叫法,并询问 Jeff Dean 其词源。这个观察把硬件命名和系统组织联系了起来。33

4. Tianfu Fu(@TianfuF)

现为 Google DeepMind researcher,曾在 xAI 做 reasoning、在 OpenAI 任 MTS。简介把他的方向落在推理模型和脑科学交叉背景上。最近他转发「Why can Kimi ship K3?」的经验帖,关注模型发布背后的训练和工程组织,但没有宣称自己加入 Kimi。32

5. Songlin Yang(@SonglinYang4)

Thinking Machines Lab 的 pretraining 研究员,简介还标注 MIT CSAIL 背景。她关注预训练、开放权重模型和测试时计算。最近一周内,她转发了 Kimi K3 pretrain 的快速比较、测试时深度和 scaling law 讨论,也转发了 Inkling 的 ARC-AGI 结果,信息密度高但多数是转发。20

6. Martin Ziqiao Ma(@ziqiao_ma)

Thinking Machines Lab technical staff,兼做 ACL mentorship,研究训练背景来自密歇根大学。当前关注 Inkling 的开放权重能力与 ARC-AGI 评测。最近他直接发布 Inkling 在 ARC-AGI-1 和 ARC-AGI-2 上的最高分说法,是团队本窗口最明确的一条模型评测公告。17

7. Shizhe Diao(@shizhediao)

Thinking Machines Lab MTS,曾在 NVIDIA,简介还维护着 LMFlow,方向是让用户微调自己的 LLM。最近他转发 Inkling 的 ARC-AGI 验证结果,并写下开放权重 SOTA 的判断。对读者来说,他的账号适合用来追踪模型微调、开放权重发布和评测落地。34

8. Wu Haoning(@HaoningTimothy)

Moonshot/Kimi 研究员,简介明确写着 scaling RL 和 VLMs,学术背景来自南洋理工大学和北京大学。最近他分享了一段使用 Kimi 进行写作与哲学讨论的体验,认为模型既能写代码,也能写作和对话。它是应用反馈,不是正式 benchmark,但能补充训练方向之外的用户侧感受。35

9. Jason Lee(@jasondeanlee)

他曾是 Google DeepMind Research Scientist,目前是伯克利 CS 与统计学教授,研究简介聚焦 LLM 和深度学习。最近连续讨论模型表达能力、泛化界限与优化,随后评论 Codex 取消 5 小时限制后的周限额。研究观点和产品观察并列出现,是本窗口较鲜明的跨层追踪对象。12

10. Alec(@AlecTPhD)

Anthropic MTS、Claude Science 创作者,具体方向是把 Claude 用进科学领域的工作流。最近他写道,自己看到很多领域专家使用 Claude Science 的真实案例,但没有公开样本、任务和指标。后续应继续追踪这些案例是否会形成可复现的科学 agent 评测。8

逐条原帖索引

以下按公司归类保留窗口内 230 条去重原帖。摘要只复述连接器返回的原文标题或正文片段;「仅链接或媒体入口」表示本账号在返回内容中没有可读正文。每条链接都指向当前发文账号的原始 X 帖子。

OpenAI:逐条原帖(135 条)

  • 07 月 17 日 08:09,@jxnlco(发文):Confirmed just press cmd+cmd 原帖
  • 07 月 17 日 08:10,@jxnlco(发文):Was going to take today off and go to New York but I slept in and missed my flight cause I was making a movie in iMovie… 原帖
  • 07 月 17 日 08:52,@nickbaumann_(发文):10/10 Trojan horse move 原帖
  • 07 月 17 日 08:55,@jxnlco(发文):This is who's running the account today. 原帖
  • 07 月 17 日 08:56,@jxnlco(发文):Where is the OpenAI soccer team? Where is the OpenAI jazz band? Where's the OpenAI competitive League of Legends esports league? Where is the OpenAI ping pong team? We are so ea... 原帖
  • 07 月 17 日 08:57,@jxnlco(转发):RT @pikacap: Sam Altman: 「you need to cross multiply」 OpenAI Researcher: 「which one of you is multiply?」 原帖
  • 07 月 17 日 09:08,@shuchaobi(发文):congrats on a very solid model. kimi cooked well! 原帖
  • 07 月 17 日 09:09,@pxd(转发):RT @ml_angelopoulos: What makes Ion special? He's deeply in the details. In the office every day, these days. Writing PRDs like an IC PM. I… 原帖
  • 07 月 17 日 09:17,@prd_008(发文):Just like Nolan intended 原帖
  • 07 月 17 日 09:38,@nickbaumann_(转发):RT @pikacap: Sam Altman: 「you need to cross multiply」 OpenAI Researcher: 「which one of you is multiply?」 原帖
  • 07 月 17 日 09:43,@aidan_clark(发文):The labs are splitting into those that do anything to buy compute and those that do anything to sell compute. I'm surprised there seem to be just two of the former .... it's uns... 原帖
  • 07 月 17 日 09:58,@MillionInt(发文):Poland and CUDA ❤️ 原帖
  • 07 月 17 日 10:02,@andrew_n_carr(发文):仅链接或媒体入口 原帖
  • 07 月 17 日 10:05,@andrew_n_carr(发文):Everything we've seen with neural rendering shows it's really good at somethings, but really bad at others. The motion here is great, and the train style is great, but honestly ... 原帖
  • 07 月 17 日 10:07,@cryps1s(转发):RT @AISecurityInst: On our cyber range "The Last Ones", GLM-5.2 matches Opus 4.5, released ~7 months before it, while DeepSeek's V4-Pro fal… 原帖
  • 07 月 17 日 10:17,@giffmana(转发):RT @the_shweenz: Nobody should trust what this guy says he's only worth 2 walnuts (in California) 原帖
  • 07 月 17 日 10:18,@steipete(发文):We need evals on irony. 原帖
  • 07 月 17 日 10:27,@shaneguML(转发):RT @brianzhan1: It was an honor to give one of the invited workshop talks at ICML. The thesis: whoever turns messy real-world outcomes into… 原帖
  • 07 月 17 日 10:37,@steipete(发文):IFKYK 原帖
  • 07 月 17 日 10:39,@archanfel_anoth(发文):Congrats to Kimi! Solid research pays back. 原帖
  • 07 月 17 日 11:02,@nickbaumann_(发文):仅链接或媒体入口 原帖
  • 07 月 17 日 11:10,@nickbaumann_(发文):How I'm taking care of my health these days: - chatGPT Work (browser) maintains my training plan and workout log - heartbeats prepare each day's workout using my plan and recove... 原帖
  • 07 月 17 日 11:11,@neil_projects(转发):RT @MarcosHernanz: FYI: GPT 5.6 Sol xhigh is better than Kimi K3 at 2/3 the price 原帖
  • 07 月 17 日 11:22,@nickbaumann_(转发):RT @prd_008: 原帖
  • 07 月 17 日 11:22,@jxnlco(转发):RT @nickbaumann_: 原帖
  • 07 月 17 日 11:24,@nickbaumann_(发文):prompt if you want to activate a coach who will track/plan/log all your workouts: Be my long-term fitness coach and training-journal manager. I have no existing plan or training... 原帖
  • 07 月 17 日 11:28,@JuhanaPeltomaa(发文):me in this line up 原帖
  • 07 月 17 日 11:28,@suchenzang(发文):soft power victories: - deepseek making everyone take ai research from china seriously - kimi making people who have never posted in chinese start posting in chinese chinamaxxin... 原帖
  • 07 月 17 日 11:38,@jxnlco(发文):Oh my god they quantized their rate limits! 原帖
  • 07 月 17 日 11:51,@jxnlco(发文):daily reminder, codex self control: it can list all the tasks and sessions that are available it can rename every single title it can send a message to each task from another th... 原帖
  • 07 月 17 日 11:55,@jxnlco(转发):RT @mrdoob: I bet someone must have had a real heart attack. This shouldn't be legal... 原帖
  • 07 月 17 日 11:56,@jxnlco(发文):F*ble 5 now available on bedrock! 原帖
  • 07 月 17 日 12:13,@btaylor(转发):RT @J_Waldmann: In early 2024, I talked to a lot of experts and the consensus view was 「you can reduce support costs by 30% with AI." We a… 原帖
  • 07 月 17 日 12:20,@suchenzang(发文):lord give me the confidence of a man who can so confidently assert that the CCP would have no understanding of open-source risks/have strategic blindness towards "AGI" à la yann... 原帖
  • 07 月 17 日 12:37,@xiangyuqi_pton(转发):RT @AISecurityInst: On our cyber range "The Last Ones", GLM-5.2 matches Opus 4.5, released ~7 months before it, while DeepSeek's V4-Pro fal… 原帖
  • 07 月 17 日 12:42,@karinanguyen(转发):RT @thoughtfullab: 原帖
  • 07 月 17 日 12:47,@karinanguyen(转发):RT @bertgodel: Today, @paperinstr and @thoughtfullab are releasing DiligenceBench, an agent-first benchmark for long-form equity research.… 原帖
  • 07 月 17 日 12:52,@andrew_n_carr(发文):I haven't used the model a lot (it passes all my poetry benchmarks with flying colors). There are lots of claims that it is benchmaxxed etc, but I find this result hard to expla... 原帖
  • 07 月 17 日 12:56,@anp(转发):RT @haroonisdreamin: few years ago I found this ethnographic report studying olympic level swimmers & excellence "At the higher levels of… 原帖
  • 07 月 17 日 13:04,@charliermarsh(发文):If you live in the Codex TUI I am here for you 原帖
  • 07 月 17 日 13:10,@JuhanaPeltomaa(发文):wow im averaging 1-3 posts with 10 likes a week 原帖
  • 07 月 17 日 13:16,@guinnesschen(发文):The future of programming is one guy pacing around a conference room arguing with his laptop while the product builds itself 原帖
  • 07 月 17 日 13:29,@nickbaumann_(转发):RT @guinnesschen: The future of programming is one guy pacing around a conference room arguing with his laptop while the product builds its… 原帖
  • 07 月 17 日 13:37,@reach_vb(发文):ICYMI: you can use 100% of your ChatGPT subscription towards GPT 5.6 Sol, Terra & Luna 原帖
  • 07 月 17 日 13:47,@HongxunWu(转发):RT @minilek: This policy limits student visas to 4 yrs, then USCIS must approve a formal extension request. Median PhD length is 5-6 yrs,… 原帖
  • 07 月 17 日 13:48,@giffmana(发文):yay for Muse Spark 1.1 on OpenRouter!! 原帖
  • 07 月 17 日 13:59,@ypatil125(发文):Love seeing more AI companies like Rox pushing the frontier pareto frontier! 原帖
  • 07 月 17 日 14:02,@giffmana(发文):You know that videoclip of that basketball player who's casually bouncing the ball, but once the bounce isn't perfect so he goes back and checks a few more times to confirm it's... 原帖
  • 07 月 17 日 14:04,@gdb(发文):GPT-5.6 Sol is the state of the art in cyber. Seeing significant results in applying it to finding and fixing novel vulnerabilities. Sign up as a defender to use it to secure yo... 原帖
  • 07 月 17 日 14:14,@guinnesschen(转发):RT @jxnlco: daily reminder, codex self control: it can list all the tasks and sessions that are available it can rename every single title… 原帖
  • 07 月 17 日 14:23,@charliermarsh(转发):RT @tobi: OpenAI shipped /goal not too long ago. I feel like GPT-5.6-sol is the first model that just doesn't need it anymore? It just ke… 原帖
  • 07 月 17 日 14:26,@gdb(发文):Sol gets the thing done: 原帖
  • 07 月 17 日 14:35,@nickbaumann_(发文):What I have ChatGPT Work (web) doing for me in my personal life (95% from iOS): - watching my inbox 2x/day and flagging Amazon packages that are arriving/arrived - keeping an ey... 原帖
  • 07 月 17 日 14:38,@jxnlco(转发):RT @guinnesschen: The future of programming is one guy pacing around a conference room arguing with his laptop while the product builds its… 原帖
  • 07 月 17 日 14:39,@jxnlco(发文):Vibes 原帖
  • 07 月 17 日 14:48,@karinanguyen(转发):RT @malthe8: Had a great time designing DiligenceBench and its reference harnesses with the team @paperinstr x @thoughtfullab It is a benc… 原帖
  • 07 月 17 日 14:53,@steipete(发文):5.6 Terra high is underrated. Switched @clawsweeper (GitHub review bot) to it and it's ~40% faster overall with negligible quality loss. Better than 5.5 on all counts. Massively... 原帖
  • 07 月 17 日 14:59,@guinnesschen(转发):RT @jxnlco: Vibes 原帖
  • 07 月 17 日 15:00,@steipete(发文):Been low key tweaking and it's the only thing now that stands between me and daily GitHub rate limit issues. 原帖
  • 07 月 17 日 15:25,@yong_zhengxin(转发):RT @kaleybrauer: If you ask a frontier LLM a multi-hop reasoning question, e.g., "Who won the Nobel Prize for Chemistry in (1900 + Mozart's… 原帖
  • 07 月 17 日 15:26,@karinanguyen(发文):We dropped DiligenceBench, a new frontier, rubric-based eval for public-equity research. A few observations: 1/ Meta Muse Spark 1.1 tops the finance harness at 57.4%, followed b... 原帖
  • 07 月 17 日 15:39,@ryanbrewer(发文):I've been reflecting on what it means to be a software engineer over the last few months. Right now it feels like the job is no longer software engineering. It's hard for me to ... 原帖
  • 07 月 17 日 15:50,@jxnlco(转发):RT @thsottiaux: @maxedapps @AnthropicAI Let me see what I can do 原帖
  • 07 月 17 日 15:55,@steipete(发文):In the category: "don't trust benchmarks". For my use case of issue/code review, Terra high by far delivers better results than Sol low. 原帖
  • 07 月 17 日 16:05,@ryanbrewer(转发):RT @OpenAI: GPT-5.6 Sol sets a new state of the art in cybersecurity on 「The Last Ones」 cyber range. We're already seeing that capability… 原帖
  • 07 月 17 日 16:06,@jxnlco(转发):RT @OpenAINewsroom: A mushroom grow kit sparked Hunter's curiosity and grew into Mountain Mushrooms. Now ChatGPT helps him run the business… 原帖
  • 07 月 17 日 16:21,@gdb(发文):don't sleep on terra! 原帖
  • 07 月 17 日 16:24,@jxnlco(发文):how does ai help you achieve flow state? notice how flow state centers and celebrate human intelligences flow state cannot replace a person's judgement or taste, 原帖
  • 07 月 17 日 16:33,@steipete(发文):Told my claw to setup a printer 原帖
  • 07 月 17 日 16:34,@steipete(转发):RT @petergostev: Let me show you how to run Kimi-K3 locally, right from your own house 原帖
  • 07 月 17 日 16:35,@steipete(转发):RT @IronWolve: @petergostev Its so simple, I don't know why everyone would not do it! (^-^原帖
  • 07 月 17 日 16:36,@steipete(转发):RT @argento_n_ext: @petergostev 原帖
  • 07 月 17 日 16:42,@steipete(发文):ya'all made me go crazy with codexbar icon customization issues, so I built an editor. (by me, I mean codex) 原帖
  • 07 月 17 日 16:59,@xiangyuqi_pton(转发):RT @OpenAI: GPT-5.6 Sol sets a new state of the art in cybersecurity on 「The Last Ones」 cyber range. We're already seeing that capability… 原帖
  • 07 月 17 日 17:06,@xiangyuqi_pton(发文):and ExploitGym ( as well ------ It all seems to just have emerged naturally from strong general reasoning capabilities... 原帖
  • 07 月 17 日 17:22,@suchenzang(发文):never. american data is superior like the american people. the chinese is barely conscious. what good would data from the lesser-sentients bring? america is great. we've made it... 原帖
  • 07 月 17 日 17:23,@garymlin(发文):This is a huge W for all Laker fans named Albert 原帖
  • 07 月 17 日 17:34,@steipete(发文):Are we still talking loops or did we shift to graphs yet? 原帖
  • 07 月 17 日 17:57,@prd_008(转发):RT @pvncher: Good rule is Sol Medium, Terra high or Luna xhigh Smaller the model, higher the reasoning you need to approach the quality… 原帖
  • 07 月 17 日 18:15,@ryanbrewer(发文):Got paid to post here this is insanity 原帖
  • 07 月 17 日 18:20,@charliermarsh(发文):Is this good 原帖
  • 07 月 17 日 18:34,@ajambrosino(转发):RT @dwr: ChatGPT app remote tab to Codex running a on closed, powered laptop is really well-done. It just works. 原帖
  • 07 月 17 日 19:15,@MillionInt(发文):Compute moves to those with the highest margin products. Markets are efficient and rational. It's not a very good question to ask an AI company "where will you source compute fr... 原帖
  • 07 月 17 日 19:19,@MillionInt(发文):RLHF raters definitely got AI hacked by text sounding smart = being smart. In 2026 smell of AI is text that sounds incredibly complicated with not much substance. That too will ... 原帖
  • 07 月 17 日 19:47,@thsottiaux(发文):GPT-5.6 Sol confirmed to be an extremely good model 原帖
  • 07 月 17 日 19:55,@charliermarsh(发文):This is actually the right amount to spend on AWS, but y'all ain't ready for that conversation… 原帖
  • 07 月 17 日 20:02,@jeremyli__(转发):RT @7uomoki: Massive genomic analyses are growingly important in bio. We present 」CuGen」 - a GPU-accelerated framework for large-scale gen… 原帖
  • 07 月 17 日 20:19,@steipete(发文):It's both amazing and painful to watch codex use browser + computer use to open Chrome, go to my PR, tap on comment and wrangle with the macOS picker - all TO UPLOAD AN IMAGE. G... 原帖
  • 07 月 17 日 20:51,@ajambrosino(发文):ok fine, let's go back 原帖
  • 07 月 17 日 20:53,@joannejang(发文):it's only automation if you don't have to think about it otherwise it's augmentation which is fun in its own way 原帖
  • 07 月 17 日 21:16,@joannejang(发文):overheard: he was a spite hire 原帖
  • 07 月 17 日 21:25,@gabrielchua(转发):RT @thsottiaux: GPT-5.6 Sol confirmed to be an extremely good model 原帖
  • 07 月 17 日 21:25,@joannejang(发文):i pair programmed with someone today. he just spoke the prompt and i typed it in 原帖
  • 07 月 17 日 21:26,@gabrielchua(发文):That was fun - thank you everyone for the great questions All the best! 原帖
  • 07 月 17 日 21:34,@gabrielchua(发文):OpenAI Build Week is 🔥🔥🔥 32 community events GLOBALLY this weekend 18 July: Bangkok 🇹🇭: Bengaluru 🇮🇳: Berlin 🇩🇪: Brussels 🇧🇪: Chiang Mai 🇹🇭: Dhaka 🇧🇩: Dhaka 🇧🇩:... 原帖
  • 07 月 17 日 21:35,@gabrielchua(转发):RT @NHv2Pro: Excellent session by @gabrielchua in the OpenAI Build Week Office Hours session on Discord. Man basically solo-ed a room wit… 原帖
  • 07 月 17 日 21:36,@JoanneShang(转发):RT @gracejkim9: our team at openai wants to make codex more useful for you! if you've ever tried using codex for a work task and it failed… 原帖
  • 07 月 17 日 21:40,@jasonkwon(转发):RT @thsottiaux: Oops... I did it again. Enjoy reset usage limits for all paid users for Codex and ChatGPT Work. Super grateful for an inc… 原帖
  • 07 月 17 日 21:42,@jxnlco(转发):RT @thsottiaux: Oops... I did it again. Enjoy reset usage limits for all paid users for Codex and ChatGPT Work. Super grateful for an inc… 原帖
  • 07 月 17 日 21:45,@prd_008(转发):RT @_simonsmith: Today I gave ChatGPT (desktop, 5.6 Sol) the URL of a product webcast and asked it to download it, transcribe it (did you k… 原帖
  • 07 月 17 日 21:45,@JoanneShang(发文):requesting a list of hojicha spots next 原帖
  • 07 月 17 日 21:47,@jxnlco(发文):We don't believe in concentration of power as you can tell but the wya the company uses Twitter on their personals vs doing everything from a faceless account. 😭 原帖
  • 07 月 17 日 21:58,@gabrielchua(转发):RT @dayvough: The Codex Buildathon Manila is starting! ⚡️ 原帖
  • 07 月 17 日 22:01,@jxnlco(转发):RT @ChatGPTapp: Create something worth sharing with ChatGPT sites. Share your site with us for a chance to win a limited-edition swag box,… 原帖
  • 07 月 17 日 22:01,@gabrielchua(转发):RT @injaneity: a love letter from @romainhuet to start @OpenAIDevs build weekend in singapore! we have back to back events and i couldnt be… 原帖
  • 07 月 17 日 22:24,@gabrielchua(转发):RT @sentrytoast: OpenAI Build Week! @brianchew @injaneity 原帖
  • 07 月 17 日 22:25,@gabrielchua(发文):You can use GPT-5.6 in Amp by bringing your own ChatGPT subscription! 原帖
  • 07 月 17 日 22:34,@gabrielchua(转发):RT @thsottiaux: Oops... I did it again. Enjoy reset usage limits for all paid users for Codex and ChatGPT Work. Super grateful for an inc… 原帖
  • 07 月 17 日 22:34,@gabrielchua(转发):RT @brianchew: WE BEGINNING!!!!!!!!!!!!! Hack day in singapore with @OpenAIDevs its so tight i cant even go in 🤣 原帖
  • 07 月 17 日 22:35,@gabrielchua(转发):RT @brianchew: @thsottiaux thank you for resetting midway through our event everyone loved it 原帖
  • 07 月 17 日 22:46,@jxnlco(转发):RT @HamelHusain: 原帖
  • 07 月 17 日 22:54,@jasonkwon(转发):RT @cremieuxrecueil: Data centers barely use any water, barely use any land, and they lower electric bills. 原帖
  • 07 月 17 日 23:03,@jasonkwon(转发):RT @thsottiaux: Look at this beauty 原帖
  • 07 月 17 日 23:09,@jasonkwon(转发):RT @sama: this is cool: 原帖
  • 07 月 17 日 23:18,@gabrielchua(转发):RT @KushalVijay_: We have a full house at @OpenAI Build Week Community Hackathon 📌 Hyderabad @gabrielchua @reach_vb @paw_lean @OpenAIDevs… 原帖
  • 07 月 17 日 23:24,@gabrielchua(转发):RT @dayvough: I finally did it. I got to demo Codex playing Balatro for me in a crowd. 🤣 #OpenAIBuildWeek 原帖
  • 07 月 17 日 23:48,@suchenzang(发文):kimi-is-just-a-distill-and-also-a-national-security-risk copium begins! 原帖
  • 07 月 17 日 23:53,@TianfuF(转发):RT @Xinyu2ML: Why can Kimi ship K3? Let me tell my story. Earlier this year, I left academia for industry. I talked to a lot of companies… 原帖
  • 07 月 18 日 00:31,@gabrielchua(转发):RT @brianchew: codex consultation bar @injaneity 原帖
  • 07 月 18 日 00:34,@gabrielchua(转发):RT @seratch_ja: カバー写真いいですね!もうすぐ提出締め切り🚀 #OpenAIBuildWeek #aimeetup 原帖
  • 07 月 18 日 01:08,@MicahCarroll(转发):RT @S_OhEigeartaigh: Got to say it's a cracking speech by Xi. China didn't have the advantage of a strong philanthropic movement supporting… 原帖
  • 07 月 18 日 01:14,@gabrielchua(转发):RT @SherryYanJiang: @OpenAI build day with @teja_natasha @unprofeshme ! 原帖
  • 07 月 18 日 01:14,@gabrielchua(转发):RT @unprofeshme: happy to report @teja_natasha's pet has hatched and is v adorable!!! 原帖
  • 07 月 18 日 01:45,@reach_vb(发文):ah yes the weekend, finally time to relax and use personal codex to build single use software and automate away the boring life admin 原帖
  • 07 月 18 日 02:19,@yong_zhengxin(转发):RT @AryaTschand: You've probably heard that power is our biggest constraint, but how does it influence training and inference systems? I w… 原帖
  • 07 月 18 日 02:46,@andrewho03(发文):仅链接或媒体入口 原帖
  • 07 月 18 日 03:03,@hyhieu226(发文):Someone please take Kimi K3's weights and put it into an agent that gets IMO 42/42 😄 open or not. 原帖
  • 07 月 18 日 04:05,@hyhieu226(发文):i have not seen better. 原帖
  • 07 月 18 日 04:38,@gabrielchua(发文):Watched The Odyssey last night, and it was so so good 原帖
  • 07 月 18 日 04:49,@giffmana(发文):It's kinda fun that "pod" is becoming de facto industry standard for "bunch of highly connected accelerators" AFAIK "TPU Pod" was an ocean pun (a group of whales/dolphins) becau... 原帖
  • 07 月 18 日 04:52,@giffmana(转发):RT @TheMathFlow: The Geometry of Integration by Parts. 原帖
  • 07 月 18 日 06:00,@charliermarsh(发文):New definition of RSI just dropped 原帖
  • 07 月 18 日 06:24,@andrew_n_carr(发文):Yes, I find myself asking sol to define multiple words per session now. Definitely an odd feeling 原帖
  • 07 月 18 日 07:53,@reach_vb(发文):It's insane how quickly we get used to a technology, and absorb its benefits! ChatGPT is less than 44 months old Codex is less than 15 months old GPT-5.3-Codex* is less than 6 m... 原帖
  • 07 月 18 日 07:58,@jxnlco(发文):was talking to chatgpt voice this week an started working on my drums and chat called it out?! 原帖

Anthropic:逐条原帖(15 条)

  • 07 月 17 日 09:40,@jifan_zhang(转发):RT @arcprize: Inkling from @thinkymachines on ARC-AGI (Verified) - ARC-AGI-2: 36.5%, $0.64/task - ARC-AGI-1: 79.5%, $0.30/task As of toda… 原帖
  • 07 月 17 日 09:42,@distributionat(发文):don't worry, it can get much worse. this was 10 AM on 20 sept 2020 in california 原帖
  • 07 月 17 日 10:22,@TrentonBricken(转发):RT @logangraham: Excited to finally be able to tweet this: we're looking for a world class visionary to lead @AnthropicAI's cyberdefense re… 原帖
  • 07 月 17 日 10:25,@RobertJBye(发文):Hey @nikitabier and @elonmusk I know this is a good meme. But this is a bit much! 原帖
  • 07 月 17 日 11:24,@distributionat(发文):empress cixi's bedroom in the forbidden palace 原帖
  • 07 月 17 日 11:47,@trq212(发文):building prototypes of mockups, schemas, data models, proof of concepts, etc. is the best way to avoid spending tons of tokens before realizing you don't want the output 原帖
  • 07 月 17 日 13:13,@skirano(发文):Btw here's all the config you need to run Kimi 3 inside of Codex via the @OpenRouter API 原帖
  • 07 月 17 日 13:13,@trq212(转发):RT @ClaudeDevs: We've resolved an issue where Fable was not selectable as a model within or Claude Code for a 30 mi… 原帖
  • 07 月 17 日 14:20,@AlecTPhD(发文):It was amazing to see so many real examples of the power of Claude Science in the hands of domain experts. 原帖
  • 07 月 17 日 17:50,@katelyn_lesse(发文):you should totally try mcp for agents (or humans) to fire off other agents (but probably not other humans) 原帖
  • 07 月 17 日 22:15,@rishicomplex(转发):RT @claudeai: Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standar… 原帖
  • 07 月 17 日 22:56,@jiaxinwen22(转发):RT @deanwball: Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation… 原帖
  • 07 月 17 日 23:23,@EthanJPerez(转发):RT @ERA_Cambridge: We're launching a new Frontier AI Security Residency! Apply now at The Frontier AI Security Res… 原帖
  • 07 月 18 日 06:35,@minilek(转发):RT @HooverInst: US visas issued to international students fell by roughly a third in 2025. Today, DHS finalized changes to a rule known as… 原帖
  • 07 月 18 日 07:41,@distributionat(发文):A squadron of wasian adolescents in football gear (messi, J.alvarez) has thundered into Newton Food Centre 原帖

Google DeepMind:逐条原帖(26 条)

  • 07 月 17 日 09:10,@GoogleDeepMind(转发):RT @PeterWBattaglia: We're announcing a major update to Weather Lab, our interactive website for sharing Google's AI weather models from @G… 原帖
  • 07 月 17 日 10:14,@tonyzzhao(转发):RT @chichengcc: We just hit a weird milestone: our model became more reliable than your average home WiFi. Just like everybody else, we th… 原帖
  • 07 月 17 日 10:32,@tonyzzhao(转发):RT @chichengcc: The only reason we could make the switch is because the team had world class experts across the whole stack. If you want to… 原帖
  • 07 月 17 日 10:33,@tonyzzhao(转发):RT @MilaidyNow: TLDR ACT-2 Preview is the first robotics model to combine broad generalization with high reliability in real homes. Key B… 原帖
  • 07 月 17 日 11:33,@jasondeanlee(转发):RT @jasondeanlee: @5_utr Doesn't give counterexample for gaussians. 原帖
  • 07 月 17 日 11:35,@jasondeanlee(转发):RT @rsalakhu: I've been asked several times whether Zhilin Yang, the founder of @Kimi_Moonshot was my PhD student. The answer is yes and he… 原帖
  • 07 月 17 日 11:47,@tonyzzhao(转发):RT @DeryaTR_: This is amazing! It looks like robot general intelligence (RGI) will be achieved soon. Some super cool robot advances are h… 原帖
  • 07 月 17 日 12:02,@tonyzzhao(转发):RT @xmaquina: A new @xmaquina proposal is on the way. Sunday Robotics just introduced ACT 2, reporting 99.1 percent zero shot success acro… 原帖
  • 07 月 17 日 15:46,@jasondeanlee(发文):America wake up and make a usable open model. 原帖
  • 07 月 17 日 18:05,@tonyzzhao(转发):RT @looeegee_: @perryzjia Taking the time to name them "Memory Developers" as opposed to something like "Data Collectors" is just one of ma… 原帖
  • 07 月 17 日 18:26,@jasondeanlee(发文):It's given every 4 years 原帖
  • 07 月 17 日 19:04,@MengdiWang10(转发):RT @GT_HaoKang: Chinese verison of Ilya and Hinton. 原帖
  • 07 月 17 日 19:22,@infoxiao(发文):damn da true alphabet 原帖
  • 07 月 17 日 19:29,@tonyzzhao(转发):RT @Guan__SUN: Your laundry should not depend on your WiFi. 原帖
  • 07 月 17 日 19:45,@MengdiWang10(转发):RT @ysu_nlp: This is why we need universities 原帖
  • 07 月 17 日 20:08,@jasondeanlee(发文):Continue 原帖
  • 07 月 17 日 20:48,@wenhaocha1(转发):RT @Lyubh22: With MLS-Bench adopted by Kimi K3, I would like to introduce it again to the community. Agents have been pushed hard on coding… 原帖
  • 07 月 17 日 23:04,@jasondeanlee(发文):Yes optimization is the real issue. All models are universal. All models have the same expressivity. All have the same generalization bounds (roughly). 原帖
  • 07 月 17 日 23:07,@jasondeanlee(发文):Doubt it. But for any one fixed problem it will be better than almost all mathematicians. 原帖
  • 07 月 17 日 23:07,@jasondeanlee(转发):RT @thsottiaux: Oops... I did it again. Enjoy reset usage limits for all paid users for Codex and ChatGPT Work. Super grateful for an inc… 原帖
  • 07 月 17 日 23:29,@jasondeanlee(转发):RT @lukaszstarosta: Since the removal of the 5h limit in Codex, my Weekly limits became Daily limits. 原帖
  • 07 月 18 日 00:42,@jasondeanlee(发文):Removed 5 h limit but weekly limit became the new 5 h limit with 5.6 原帖
  • 07 月 18 日 02:26,@jasondeanlee(转发):RT @kchonyc: another hot tip: call it gemini-4-flash and release right now 原帖
  • 07 月 18 日 02:39,@jasondeanlee(转发):RT @MillionInt: RLHF raters definitely got AI hacked by text sounding smart = being smart. In 2026 smell of AI is text that sounds incredi… 原帖
  • 07 月 18 日 06:18,@tonyzzhao(转发):RT @YXWangBot: Salute to Sunday folks 🫡 It is amazing as well as interesting to see how much solid engineering is needed for smooth and rel… 原帖
  • 07 月 18 日 06:48,@HeinrichKuttler(发文):Ignoring the bellicose contemporary message: Will we have accurate archives of today's banter in another 44 years? Who's working on that? 原帖

Meta AI:逐条原帖(3 条)

  • 07 月 17 日 12:42,@YifeiZhou02(转发):RT @HavenFeng: Today, we're introducing schema: a harness reaching 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on ARC-A… 原帖
  • 07 月 17 日 13:55,@Ber18791531(转发):RT @arcprize: Inkling from @thinkymachines on ARC-AGI (Verified) - ARC-AGI-2: 36.5%, $0.64/task - ARC-AGI-1: 79.5%, $0.30/task As of toda… 原帖
  • 07 月 18 日 07:09,@alexandr_wang(转发):RT @leo_linsky: I was VERY skeptical, but it looks like Meta really did it. Muse Spark 1.1 is on par with Opus 4.8 and better than Grok 4… 原帖

Thinking Machines Lab:逐条原帖(15 条)

  • 07 月 17 日 08:39,@shizhediao(转发):RT @arcprize: Inkling from @thinkymachines on ARC-AGI (Verified) - ARC-AGI-2: 36.5%, $0.64/task - ARC-AGI-1: 79.5%, $0.30/task As of toda… 原帖
  • 07 月 17 日 08:40,@shizhediao(转发):RT @venturetwins: Built a podcast clipping app with Inkling from @thinkymachines ✨ The model is exceptional at reasoning over long-form au… 原帖
  • 07 月 17 日 09:01,@shizhediao(发文):Excited to see Inkling become the new open-weights SOTA on ARC-AGI-1 and ARC-AGI-2! 原帖
  • 07 月 17 日 09:04,@cHHillee(转发):RT @arcprize: Inkling from @thinkymachines on ARC-AGI (Verified) - ARC-AGI-2: 36.5%, $0.64/task - ARC-AGI-1: 79.5%, $0.30/task As of toda… 原帖
  • 07 月 17 日 09:50,@ziqiao_ma(发文):We are excited that Inkling is the highest-scoring open-weight (as of today) model evaluated by @arcprize on both ARC-AGI-1 and ARC-AGI-2! 原帖
  • 07 月 17 日 10:04,@regulargio(转发):RT @arcprize: Inkling from @thinkymachines on ARC-AGI (Verified) - ARC-AGI-2: 36.5%, $0.64/task - ARC-AGI-1: 79.5%, $0.30/task As of toda… 原帖
  • 07 月 17 日 10:46,@cHHillee(发文):Imo a lot of people don't think about the a2a cost in EP correctly. Andrew Gu explained this perspective to me a while ago and I think it's the right one. 原帖
  • 07 月 17 日 13:10,@liliang_ren(转发):RT @arcprize: Inkling from @thinkymachines on ARC-AGI (Verified) - ARC-AGI-2: 36.5%, $0.64/task - ARC-AGI-1: 79.5%, $0.30/task As of toda… 原帖
  • 07 月 17 日 17:57,@SonglinYang4(转发):RT @RyanGreenblatt: I did some quick tests that indicated that the Kimi K3 pretrain is around halfway between Opus 4 and Opus 4.5. So ~10 m… 原帖
  • 07 月 17 日 21:08,@SonglinYang4(发文):🚀 原帖
  • 07 月 17 日 21:26,@SonglinYang4(转发):RT @lyraaaa: is k3 actually a transformer? why not consult The Chart! 原帖
  • 07 月 17 日 21:27,@SonglinYang4(转发):RT @ZeyuanAllenZhu: Congrats, @Kimi_Moonshot! 『 Kimi's Four Commandments』have circulated in the Chinese AI community for months. Many peopl… 原帖
  • 07 月 18 日 00:31,@SonglinYang4(转发):RT @marksaroufim: Can a model run deeper at test time than it was ever trained to? And if depth becomes a loop instead of a stack, do we ne… 原帖
  • 07 月 18 日 02:27,@SonglinYang4(转发):RT @yzhang_cs: @mark_k but quite the opposite, i'd say kimi is every bit as much a true believer in scaling laws as A, arguably more so th… 原帖
  • 07 月 18 日 02:53,@SonglinYang4(转发):RT @ziqiao_ma: We are excited that Inkling is the highest-scoring open-weight (as of today) model evaluated by @arcprize on both ARC-AGI-1… 原帖

Cognition:逐条原帖(13 条)

  • 07 月 17 日 08:45,@dabit3(转发):RT @realmtbman: Thanks to the all the attendees who attended the 1st Summer Nights in Waterloo by @cognition last night at @TheBarnWaterloo… 原帖
  • 07 月 17 日 11:50,@dabit3(转发):RT @annarmitchell: Loosely held hypothesis: there will be a surge of top talent back to Fortune 500 companies that have struggled over the… 原帖
  • 07 月 17 日 12:25,@ysu_nlp(发文):This is why we need universities 原帖
  • 07 月 17 日 12:54,@dabit3(转发):RT @yitong: Fable this, Kimi that.... what yall are really sleeping on is blasting @DevinAI SWE 1.7 lightning 1k tokens/second Most of the… 原帖
  • 07 月 17 日 14:58,@dabit3(转发):RT @cognition: The FrontierCode leaderboard is now live: a dedicated page that tracks which models are writing code you'd actually merge.… 原帖
  • 07 月 17 日 15:21,@mattbergland(发文):some folks have asked and the best domain partner to work with is @lumis_com and @HobiMichalec 原帖
  • 07 月 17 日 15:25,@ybenpan(转发):RT @cognition: The FrontierCode leaderboard is now live: a dedicated page that tracks which models are writing code you'd actually merge.… 原帖
  • 07 月 17 日 15:42,@jeffwang(转发):RT @devindesktop: Inkling and Grok 4.5 are also available in Devin Desktop and CLI 原帖
  • 07 月 17 日 18:31,@mattbergland(发文):See you at our event in SF on 7/28!! 原帖
  • 07 月 17 日 18:32,@mattbergland(发文):who has a ticket for my guy @tomasholtz for the WC final? 原帖
  • 07 月 17 日 18:34,@mattbergland(转发):RT @imjaredz: Many are saying this is a great deal and maybe the best ever 原帖
  • 07 月 17 日 19:47,@mattbergland(发文):we need an nba player to bring back the sky hook actually op 原帖
  • 07 月 17 日 21:25,@mattbergland(发文):seattle harness for the weekend 原帖

Moonshot/Kimi:逐条原帖(8 条)

  • 07 月 17 日 19:59,@Kimi_Moonshot(转发):RT @AfterQuery: Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5. An open weight model now outperforms al… 原帖
  • 07 月 17 日 20:00,@Kimi_Moonshot(转发):RT @elliotarledge: 原帖
  • 07 月 17 日 21:11,@HaoningTimothy(发文):I have been using it for writing and philosophical discussions for a period: it is not only a strong coder but also as always a good writer and nice buddy to chat with. It has a... 原帖
  • 07 月 18 日 01:05,@crystalsssup(转发):RT @levelsio: Kimi K3 is absolutely hammering through my Windows XP Simulator to do list Claude Code couldn't do this for 2 weeks or kept… 原帖
  • 07 月 18 日 01:09,@crystalsssup(转发):RT @AfterQuery: Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5. An open weight model now outperforms al… 原帖
  • 07 月 18 日 02:14,@crystalsssup(发文):I think Kimi has a really great research team. I often come across interesting things and send over to the team, and they'll reply: 「But this is not AGI.」 原帖
  • 07 月 18 日 07:10,@Kimi_Moonshot(发文):Kimi Business Membership is now available for enterprise orders, offering teams all Kimi Allegretto plan benefits with enterprise-grade support. Highlights: > Starts from 5 seat... 原帖
  • 07 月 18 日 07:20,@Kimi_Moonshot(转发):RT @voxelbench: Kimi K3 ranks 3rd on VoxelBench just 100+ Elo points behind Fable! and a huge uplift from K2.6 (28th) 原帖

Alibaba Qwen:逐条原帖(2 条)

  • 07 月 17 日 08:15,@TianbaoX(转发):RT @rsalakhu: Congratulations to Zhilin Yang, founder and CEO of @Kimi_Moonshot, on the latest Kimi release. What a huge win for the open-s… 原帖
  • 07 月 17 日 21:36,@TianbaoX(转发):RT @bigeagle_xd: 我们走后,他们会给你们修学校和医院,会提高你们的工资,但这绝不是因为他们良心发现,也不是因为他们变成了好人,而是因为我们来过。 原帖

Google Brain:逐条原帖(1 条)

  • 07 月 17 日 08:47,@AndrewYNg(发文):New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was built with @Cerebras and taugh... 原帖

Cohere:逐条原帖(1 条)

  • 07 月 17 日 11:45,@freddiev4(发文):just used claude code to order sushi via doordash CLI -- ama 原帖

xAI:逐条原帖(11 条)

  • 07 月 17 日 09:10,@BenjaminDEKR(转发):RT @Andr3jH: Anthropic and OpenAI engineers browsing the "our team" section on the Moonshot AI website and recognizing their ex-girlfriends… 原帖
  • 07 月 17 日 10:17,@adityagupta(发文):ye sab bahut wo hai 原帖
  • 07 月 17 日 10:22,@BenjaminDEKR(发文):Literally how? How do we do this I'm looking at X Android app and also Chrome mobile and have no idea where this is on X. There's nothing that says " ✨AI" 原帖
  • 07 月 17 日 10:27,@BenjaminDEKR(转发):RT @jun_song: Guide on how to run Kimi-K3 locally : - you need to buy 10 x GB300 (252GB VRAM each), thats about $1M - Electricity bill es… 原帖
  • 07 月 17 日 11:09,@BenjaminDEKR(发文):Is the Kimi claiming to be Claude thing real or fake screenshots? 原帖
  • 07 月 17 日 12:16,@ayushjaiswal(发文):I thought data companies are going to die, I was wrong. Instead, they might end up being the kingmaker. They're selling shovels. They are going to continue growing a lot more. S... 原帖
  • 07 月 17 日 13:19,@forwarddeploy(转发):RT @arafatkatze: I gave xAI a lot of shit for Grok Build's repo-upload behavior, so they deserve the same volume when they fix it. I pulle… 原帖
  • 07 月 17 日 21:16,@BenjaminDEKR(转发):RT @DolphinMossad: 1995: oh cool you can buy books online now 2026: the company that owns Whole Foods and Lord of the Rings is experiencin… 原帖
  • 07 月 17 日 21:25,@BenjaminDEKR(发文):If you owe Amazon $3000, that's your problem If you owe Amazon $3 Billion and fifty-four cents, that's Amazon's problem 😆 原帖
  • 07 月 17 日 21:58,@CMS_Flash(发文):I thought @adeptailabs revived 😂. 原帖
  • 07 月 18 日 06:04,@JasonBud(转发):RT @pidotdev: Pi is nothing without its open source contributors! Thanks to the work of @Jaaneek and the @SpaceXAI team, you can now offi… 原帖

Related content

  • Sign in to comment.
More from this channel