Three new Flash models built for large-scale agents
Google launches Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, three new models optimized for running large-scale AI agents at lower costs.
Designed for running agents at scale rather than cutting-edge reasoning, Gemini's Flash series welcomes three models with distinct roles.
Gemini 3.6 Flash serves as the workhorse. It improves on code as well as document and multimodal tasks, while consuming, according to the Artificial Analysis Index, about 17% fewer output tokens than 3.5 Flash, with fewer reasoning steps and tool calls. Its pricing also drops below that of its predecessor, at $1.50 per million input tokens and $7.50 per million output tokens, which lowers the cost per agentic task.
Gemini 3.5 Flash-Lite targets volume. Billed as the fastest in the 3.5 class, it is clocked at 350 output tokens per second by the same source, for $0.30 per million input tokens and $2.50 per million output tokens. Its domain is agentic search and just-in-time document processing, with computer use now integrated as a native tool.
The third, Gemini 3.5 Flash Cyber, specializes in vulnerability detection and remediation, grafted onto the CodeMender agent where multiple instances collaborate to produce a single report. Given its dual-use nature, Google is restricting access to governments and trusted partners through a closed pilot program.
In parallel, version 3.5 Pro is in testing with partners, and Google indicates it has launched the training of Gemini 4, which it presents as its most ambitious pre-training cycle to date.