Developers and customers building production AI agents require higher token efficiency, lower latency, and more reliable performance. Our Flash series of models are built to meet the sweet spot of efficiency and quality, allowing you to scale your agent workflows. We are introducing a new Gemini model based on Gemini 3.5 Flash.
3.6 Flash: Our flagship model for better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, output token usage is reduced by 17% compared to 3.5 Flash, and some benchmarks such as Datacurve’s DeepSWE have observed lower costs per output token, up to 65%. 3.5 Flash-Lite: Our fastest and most cost-effective 3.5 class model, delivering 350 output tokens per second according to the Artificial Analysis Index, significantly exceeding previous generations of Flash-Lite. 3.5 CodeMender’s Flash Cyber: Successful cybersecurity applications require careful orchestration of models in parallel with agent infrastructure. We are introducing a new model of highly efficient and specialized cyber specialization combined with the CodeMender code security agent that provides cutting-edge and competitive performance.
In addition to today’s release, Gemini 3.5 Pro is currently being tested with partners and will be widely available when ready. In parallel, our team is already focused on building next-generation models. We have begun our most ambitious pre-training for Gemini 4 to date and are excited about our progress.
3.6 Flash: More efficient and higher quality than 3.5 Flash.
Gemini 3.6 Flash is built directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only provides a step-up in coding and knowledge work, it does this while significantly increasing token efficiency. For example, the Artificial Analysis Index shows that 3.6 flashes consume 17% fewer output tokens than 3.5 flashes. It also requires fewer inference steps and tool calls to accomplish multi-step workflows.
This enhanced efficiency also comes at a lower price than 3.5 Flash. 3.6 Flash reduces the overall cost per agent task at $150 per million input tokens and $750 per million output tokens, making it more cost-effective to build and run agents.

