
Google launches Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber
Google's three new Gemini Flash models split the work between higher-quality general tasks, low-cost high-throughput processing, and restricted cybersecurity research.
Google splits the Gemini Flash tier by workload
Google DeepMind announced three Gemini models on July 21: Gemini 3.6 Flash for higher-quality general work, Gemini 3.5 Flash-Lite for low-latency scale, and Gemini 3.5 Flash Cyber for controlled vulnerability research. The announcement also says Gemini 3.5 Pro is still being tested with partners and that Gemini 4 pre-training has begun. 1
| Model | Intended job | What changes | Access and limits |
|---|---|---|---|
| Gemini 3.6 Flash | Coding, knowledge work, multimodal tasks, and computer use | Google says it uses 17% fewer output tokens than 3.5 Flash, with lower listed pricing of $1.50 per 1M input tokens and $7.50 per 1M output tokens. | Available through the Gemini API, Google AI Studio, Android Studio, Antigravity, the Gemini app, and enterprise products. The benchmark figures are vendor-reported. |
| Gemini 3.5 Flash-Lite | Agentic search, document processing, and other high-throughput tasks | Google cites 350 output tokens per second from Artificial Analysis, plus configurable thinking levels. Pricing is $0.30 per 1M input tokens and $2.50 per 1M output tokens. | Available through the Gemini API, Google AI Studio, Android Studio, and the Gemini app; a Google Search rollout is gradual. Its speed and cost focus make it a scaling model, not a universal replacement for a stronger reasoner. |
| Gemini 3.5 Flash Cyber | Finding and patching software vulnerabilities | A 3.5 Flash variant tuned for CodeMender, where multiple agents work together on one security report. Google describes its CyberGym result as frontier-level but does not publish a score in the announcement. | A limited-access pilot for governments and trusted partners because the model is dual-use. It is not a general public API model. |
The 3.6 Flash story is efficiency with a quality step up. Google reports 49% on DeepSWE versus 37% for 3.5 Flash, 63.9% versus 49.7% on MLE Bench, and 83.0% versus 78.4% on OSWorld-Verified. Those are useful signals for coding, machine-learning work, and computer use, but they remain claims from the launch material rather than independent leaderboard results. The official X announcement summarizes the same positioning: fewer tokens for higher-quality work, a cheaper high-speed option, and a security model for critical vulnerabilities. 2
For developers, the important change is portfolio design. A single Flash label now covers three different operating points: spend more for broader capability, spend less for throughput, or accept restricted access for specialized cyber defense. That gives teams a clearer way to match model choice to workload, but it also makes simple model rankings less useful. The next checks are real-world latency and cost, independent evaluations, and whether CodeMender's security results generalize beyond Google's pilot.
Google's forward-looking notes are deliberately thin: Gemini 3.5 Pro remains in partner testing, while Gemini 4 has entered pre-training with no release date disclosed. 1
Loading content card…
References
- 1
- 2
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.