
Grok 4.6 ties GPT-5.6 Sol on one composite score, but its edge is long-running agents
xAI's Grok 4.6 targets long-running agent work, matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, and is available through its API and several developer platforms.
Grok 4.6 is xAI's new flagship model for work that continues across many steps: research, codebase changes, interactive applications, and other artifacts that need repeated refinement. xAI released it on Aug. 12 as a follow-up to Grok 4.5, with the largest changes coming from supplemental training and agent-focused reinforcement learning rather than a newly disclosed model scale. 1
What changed
xAI says Grok 4.6 received a longer post-training run using model-generated reasoning data, engineering data, an improved optimizer, and regenerated supervised fine-tuning traces from Grok 4.5. Its reinforcement-learning tasks include knowledge work, general coding, kernel optimization, web development, and computer-aided design. The intended behavior is persistence: the model can keep a multi-step task organized, test parts of its own work, and iterate after feedback. 1
That makes the model more relevant to agent builders than to users looking for a simple chat upgrade. xAI's examples focus on turning a broad product idea into a working first version, then improving the structure, interactions, and visual treatment through several rounds. The announcement describes these capabilities and internal testing, but does not publish a general reliability rate for long-running tasks. 1
What the numbers say
In xAI's published comparison, Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol and trailing Fable 5 at 62. It leads GPT-5.6 Sol on CursorBench 3.2, 69.9% to 67.2%, and on FrontierCode 1.1, 61.3% to 60.6%. It trails GPT-5.6 Sol on DeepSWE v1.1, 65.9% to 73%, and Terminal-Bench 3.0, 26% to 34.6%. 1
These are useful signals, not a neutral league table. xAI says competing scores come from the developers' system cards or public leaderboards, while the Grok results are presented by xAI. GeekNews likewise notes that the comparison is largely based on company-reported results, even though the improvement over Grok 4.5 appears across most listed evaluations. 12
Access and practical significance
Grok 4.6 is available now in Cursor and Grok Build, as well as through xAI's API and partners including OpenRouter, Vercel, and Cloudflare. API pricing starts at $2 per million input tokens and $6 per million output tokens; a fast variant costs twice as much. Cursor and Grok Build are offering double included usage for the first week. 1
For developers, the immediate question is whether its long-horizon behavior survives real repositories and production tool calls. The release is accessible enough to test now, but the evidence supports a targeted trial, not a blanket claim that Grok 4.6 is the best general model.
References
- 1
- 2
AI Model & Product Launch Alerts
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
- Sign in to comment.