

Google's Gemini 4 Argon Tops Its Own Benchmark Table, and Only Cyber Defenders Can Touch It
On Wednesday, September 30, Google DeepMind announced Gemini 4 Argon, calling it its most capable model yet for real-world software engineering, enterprise knowledge work and cybersecurity defence. The announcement is signed by Koray Kavukcuoglu, the senior vice president who became the head of DeepMind in August. Google's last frontier-class release above its Flash line came more than seven months earlier, and investors had spent the summer worrying about delays. 123
Argon is not generally available. Google says it is rolling out to "a set of trusted cyber defenders" through its Fairwind Program — whose own page says it gives governments, healthcare providers and telecommunications services early access, and that it works with over 650 partners globally — while the company says it is "actively engaged in the U.S. government's voluntary process for pre-release model access." Broad availability is promised only "as soon as possible," starting with paid API customers and Google AI Ultra subscribers. The AI community had been passing around apparent leaked benchmarks for the model days before the announcement. 45
Google's own comparison covers 18 benchmarks. Argon leads outright on 12 and ties for first on one; GPT-6 Astra leads outright on three and Claude Opus 5.5 on two. Its clearest margins include DeepSWE v1.1 at 77.9%, AutomationBench at 51.3%, LVBench at 91.7% and Harvey's Legal Agent Benchmark at 19.6% against GPT-6 Astra's 5.4% and Claude Opus 5.5's 3.8%. The same table shows where it is behind: FrontierSWE v2 at 55.0% against Astra's 65.5%, Terminal-Bench Science 0.1 at 57.6% against 68.1%, Terminal-bench 4.0 at 57.4% against Opus 5.5's 66.4%, and PostTrainBench at 45.3% against 49.3%. Google says Argon also sets a new state of the art on DeepSWE and offers an industry-leading output limit of 1 million tokens, up from 64K. 6
The independent read comes from Artificial Analysis, which evaluated the model through the API. With high reasoning it scores 53 on its Intelligence Index, level with GPT-6 Astra (max) and one point ahead of GPT-6.1 Sol (max); at launch pricing it costs $1.99 per task, 60% of Astra's cost per task; and its 15% hallucination rate is the lowest among models scoring 45 or higher. Google's price is $2 per million input tokens and $10 per million output for an introductory period, with cached input 95% off, moving to $4 and $20 afterwards. Google has not said when that period ends. 7
Inside Google, the company says Argon agents freed more than 300 TiB of memory across its data centers, beat a published quantum baseline by 40% on qubits times gates in minutes, and are migrating C/C++ codebases to Rust up to 800K+ lines including the Fuchsia Zircon kernel — with 32,000 lines of SIMD replaced in the video decoder libgav1 for a port that runs 2.7x faster with identical output. Wiz, through its Scan for Good initiative, is already using the model; Google says it uncovered a critical vulnerability in healthcare software used by hospitals worldwide, and describes it without specifics. Tulsee Doshi, Google's Gemini model product lead, told CNBC that starting the rollout this way "gives us more confidence." 89
What is still open: there is no public release date; the public cannot run the model, so verification outside the test group rests on a handful of evaluators rather than on anyone who wants to check; the introductory price has no announced end; and Google's benchmark table, which carries every capability claim here, is Google's own. This episode is an original rap song and music video built from those facts.
References
- 1
- 2
- 3
- 4Google DeepMind — Fairwind Program
deepmind.google
- 5
- 6
- 7
- 8
- 9
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
