AI Product Updates, July 17: Kimi K3, Grok 4.5 and the Agent Stack Move Up

AI Product Updates, July 17: Kimi K3, Grok 4.5 and the Agent Stack Move Up

A July 16-17 sweep of Kimi K3, Grok 4.5, Gemini Notebook, Grok on Bedrock, Sierra Horizon, and DataRobot OpenCode, with the availability and pricing details that determine what teams can test now.

At a glance

Coverage window: July 16–17, 2026 (GMT+8)
ProductReleasing entityRelease dateWhat changedAvailability
Kimi K3Moonshot AIJuly 172.8T-parameter model with native vision, 1M-token context, long-horizon coding, and new API pricing.Live in Kimi.com, Kimi Work, Kimi Code, and the Kimi API. Full weights are promised for July 27, so self-hosting is not available yet. 1
Grok 4.5SpaceXAI / xAIJuly 16New coding, agentic-work, and knowledge-work model; 80 tokens per second and $2 per million input tokens / $6 per million output tokens.Live in Grok Build, Cursor on all plans, and the SpaceXAI console, with limited free usage in Grok Build and Cursor. 2
Grok 4.3 on BedrockAmazon Web Services and xAIJuly 17Enterprise distribution for Grok 4.3 with configurable reasoning effort, tool calling, strict structured output, image input, stateful conversations, and a 1M-token context window.Generally available on Amazon Bedrock through regional Mantle endpoints and Standard, Priority, or Flex service tiers. 3
Gemini NotebookGoogleJuly 16NotebookLM is now Gemini Notebook, with a secure cloud computer that can write and execute code inside notebooks, plus deeper Gemini and Search integration.The rename is live. Cloud-computer access is available to Google AI Ultra users and eligible Workspace customers; web access for Pro users is rolling out over the coming weeks. 4
HorizonSierraJuly 16Long-horizon agents that pursue business outcomes across interactions spread over days, weeks, or months, using a context engine and persistent learning.Announced for enterprise use. Sierra has not published self-serve access or standard pricing; it says customers pay for outcomes rather than tokens. 5
DataRobot OpenCodeDataRobotJuly 17A governed coding agent that lets teams switch among closed, open-weight, and bring-your-own models through the DataRobot LLM Gateway.Available through the DataRobot CLI or UI; DataRobot says installation takes three commands and offers a 30-day trial for new users. 6

Kimi K3 puts open weights on a timer

Moonshot AI launched Kimi K3 as a 2.8-trillion-parameter Mixture-of-Experts model with native vision and a 1-million-token context window. Its architecture uses Kimi Delta Attention and Attention Residuals, and the company positions it for long-horizon coding, knowledge work, and reasoning. The model is already selectable in Kimi.com, Kimi Work, Kimi Code, and the Kimi API. 1
The practical numbers are unusually clear for a launch of this scale. Moonshot lists API pricing at $0.30 per million cache-hit input tokens, $3 per million cache-miss input tokens, and $15 per million output tokens. The full model weights are promised by July 27, alongside a technical report. Until then, K3 is a hosted model with an open-weight release still pending, not a model developers can download and run locally. 1
That distinction matters more than the parameter count. Moonshot's own page reports frontier-level results and says K3 outperformed several tested models, while Quartz independently confirms the 2.8T size, 1M context, current availability, and July 27 weights date. The independent test of whether K3 is practical to serve, fine-tune, and maintain will begin when those weights arrive. 1 7
What to do: Use the API now if you want to test K3 against real coding or research workloads. Hold off on infrastructure decisions until the weights, serving guidance, and technical report land.

Grok 4.5 reaches the coding surface, while Grok 4.3 reaches Bedrock

Grok 4.5

SpaceXAI launched Grok 4.5 on July 16 as its new model for coding, agentic tasks, and knowledge work. The company reports 80 tokens per second and says the model uses roughly half as many output tokens as comparable leading models on its cited software-engineering task. It is priced at $2 per million input tokens and $6 per million output tokens. 2
The availability story is broader than a model-card launch. Grok 4.5 is live in Grok Build, in Cursor on all plans, and through the SpaceXAI console. SpaceXAI is also offering limited free usage in Grok Build and Cursor. That gives developers three different ways to test the same model: an app-building surface, an existing coding workflow, or a direct API. 2
What to do: Compare it on your own repository or agent loop, not on the headline benchmark table. The meaningful test is whether the lower token count survives your tool calls, retries, and review steps.

Grok 4.3 on Amazon Bedrock

AWS separately made Grok 4.3 generally available on Amazon Bedrock. The release adds a managed enterprise route to a 1-million-token, text-and-image model with configurable reasoning effort from none through high, tool calling, strict JSON-schema output, and stateful conversations through the Responses API. 3
The integration uses xAI's OpenAI-compatible Mantle endpoint rather than the standard Bedrock Runtime API. AWS says Grok 4.3 is available in regional, in-Region deployments and supports Standard, Priority, and Flex service tiers; geo and global cross-Region inference are not offered at launch. Exact token rates depend on the Bedrock pricing page and selected tier, so there is no new fixed price to carry over from the standalone xAI launch. 3
For AWS teams, this is the more consequential change than the model name itself: IAM-linked credentials, regional deployment, and familiar OpenAI SDK calls now sit around Grok workloads. That can shorten the path from evaluation to a governed production pilot, while the regional limitation still needs to be checked against data-residency requirements.

Google turns NotebookLM into Gemini Notebook

Google renamed NotebookLM to Gemini Notebook on July 16. The product remains a standalone, source-grounded research tool, but Google is adding a secure cloud computer that can write and execute code inside a notebook for data analysis. The cloud computer is available today to Google AI Ultra users and Workspace business customers with AI Ultra Access or AI Expanded Access, and is rolling out to all web Pro users over the coming weeks. 4
The second change is distribution. Notebooks can sync between the standalone experience and the Gemini app, while Google says notebooks will soon appear directly in AI Mode in Search. The rename is therefore more than a label change: the product is being positioned as a research workspace that can execute analysis and travel with the user across Google's AI surfaces. 4
Availability note: Pro users should treat the cloud computer as a staged rollout rather than a universal switch. Ultra and the specified Workspace access paths are the confirmed routes today; Google did not give a firm date for the broader Pro rollout.

Agents move from conversations to outcomes

Sierra Horizon

Sierra announced Horizon, a platform for agents that pursue long-horizon goals such as originating a loan or securing a healthcare referral. Instead of treating one chat as the unit of work, Horizon is designed to coordinate inbound and outbound interactions across days, weeks, or months, using a context engine to decide what to do between engagements. 5
Sierra also changes the commercial unit. The company says customers pay for business outcomes delivered rather than tokens, while the accumulated context from customer interactions becomes part of the enterprise's durable data advantage. The announcement references enterprise customers and design partners, but does not publish a self-serve signup path or standard price. 5

DataRobot OpenCode

DataRobot introduced OpenCode as a coding agent that can use closed, open-weight, or bring-your-own models through the company's LLM Gateway. The agent inherits DataRobot's approved gateway, agent skills, and Agent Assist layer, so changing the model is presented as a configuration change rather than a new coding-agent procurement cycle. 6
DataRobot says OpenCode can be installed through its CLI or launched from the DataRobot UI, with the model selected through /models. It advertises a 30-day trial for new users, but does not publish a standalone OpenCode price in the announcement. For teams already using DataRobot's gateway, the product's main change is control over model choice without giving up the existing governance path. 6

What to watch next

  • July 27: Moonshot's promised Kimi K3 weights release should turn a hosted model launch into a real deployment and fine-tuning test. 1
  • Google's Pro rollout: The Gemini Notebook cloud computer is confirmed for Ultra and specified Workspace customers now, with the broader web Pro rollout still staged. 4
  • Enterprise agent economics: Sierra's outcome-based pricing and DataRobot's model-agnostic gateway point in different ways toward buying AI work by completed result rather than by model seat. Both still need real customer metrics before the pricing shift can be judged.
  • Regional deployment: Grok 4.3 on Bedrock is GA, but its regional Mantle setup and lack of global cross-Region inference at launch are hard constraints for production architecture. 3
The useful pattern across this window is not simply that models are getting larger. Kimi K3 makes open deployment a near-term question, Grok spreads one model family across consumer, coding, API, and cloud surfaces, Gemini Notebook makes grounded research executable, and the agent platforms are selling control over longer workflows. For teams choosing what to test, the next step is concrete: benchmark the model, check the access boundary, and price the completed task rather than the demo.

相似内容

  • 登录后可发表评论。
More from this channel