
Google launches Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber
Google's three new Gemini Flash models split the work between higher-quality general tasks, low-cost high-throughput processing, and restricted cybersecurity research.
Google splits the Gemini Flash tier by workload
Google DeepMind announced three Gemini models on July 21: Gemini 3.6 Flash for higher-quality general work, Gemini 3.5 Flash-Lite for low-latency scale, and Gemini 3.5 Flash Cyber for controlled vulnerability research. The announcement also says Gemini 3.5 Pro is still being tested with partners and that Gemini 4 pre-training has begun. 1
| Model | Intended job | What changes | Access and limits |
|---|---|---|---|
| Gemini 3.6 Flash | Coding, knowledge work, multimodal tasks, and computer use | Google says it uses 17% fewer output tokens than 3.5 Flash, with lower listed pricing of $1.50 per 1M input tokens and $7.50 per 1M output tokens. | Available through the Gemini API, Google AI Studio, Android Studio, Antigravity, the Gemini app, and enterprise products. The benchmark figures are vendor-reported. |
| Gemini 3.5 Flash-Lite | Agentic search, document processing, and other high-throughput tasks | Google cites 350 output tokens per second from Artificial Analysis, plus configurable thinking levels. Pricing is $0.30 per 1M input tokens and $2.50 per 1M output tokens. | Available through the Gemini API, Google AI Studio, Android Studio, and the Gemini app; a Google Search rollout is gradual. Its speed and cost focus make it a scaling model, not a universal replacement for a stronger reasoner. |
| Gemini 3.5 Flash Cyber | Finding and patching software vulnerabilities | A 3.5 Flash variant tuned for CodeMender, where multiple agents work together on one security report. Google describes its CyberGym result as frontier-level but does not publish a score in the announcement. | A limited-access pilot for governments and trusted partners because the model is dual-use. It is not a general public API model. |
The 3.6 Flash story is efficiency with a quality step up. Google reports 49% on DeepSWE versus 37% for 3.5 Flash, 63.9% versus 49.7% on MLE Bench, and 83.0% versus 78.4% on OSWorld-Verified. Those are useful signals for coding, machine-learning work, and computer use, but they remain claims from the launch material rather than independent leaderboard results. The official X announcement summarizes the same positioning: fewer tokens for higher-quality work, a cheaper high-speed option, and a security model for critical vulnerabilities. 2
For developers, the important change is portfolio design. A single Flash label now covers three different operating points: spend more for broader capability, spend less for throughput, or accept restricted access for specialized cyber defense. That gives teams a clearer way to match model choice to workload, but it also makes simple model rankings less useful. The next checks are real-world latency and cost, independent evaluations, and whether CodeMender's security results generalize beyond Google's pilot.
Google's forward-looking notes are deliberately thin: Gemini 3.5 Pro remains in partner testing, while Gemini 4 has entered pre-training with no release date disclosed. 1
コンテンツカードを読み込んでいます…
関連コンテンツ
- ログインするとコメントできます。
More from this channel›
- Claude Opus 5 brings near-Fable performance to the Opus tier
- Kimi K3's second week turns a model launch into a compute and policy test
- Qwen-Audio-3.0-TTS splits the launch between speed and voice quality
- OpenAI Presence brings managed voice and chat agents to enterprise workflows
- Qwen3.8 Max preview lands with 2.4T parameters, but proof is still to come
- Thinking Machines launches Inkling, a 975B open-weight model built for customization
- Kimi K3 brings 2.8T parameters to the open-model frontier — but not yet to your servers
- Meta's Muse Spark 1.1 is live, and speed is the real headline