Google has released Gemini 3.6 Flash and 3.5 Flash-Lite, new flagship products designed to reduce latency and token costs for enterprise AI agents.
The economics of running autonomous software agents within a production environment comes down to a fixed equation: Few vendors advertise directly. The model must properly infer multi-step tasks, but each additional token generated adds cost and delay to a workflow that can run thousands of times per hour.
Teams building background agents rather than chat interfaces need throughput first, then parameter counts. Google’s answer, published this week, breaks down that tradeoff into three models. Gemini 3.6 Flash for coding and multimodal inference, Gemini 3.5 Flash-Lite for high-volume, low-latency work, and a limited Gemini 3.5 Flash Cyber variant built for vulnerability remediation.
The math behind Gemini 3.6 Flash
Google’s developer documentation for 3.6 Flash centers around one number: 17% fewer output tokens than previous 3.5 Flash versions, based on measurements from the Artificial Analysis Index.
In certain synthetic tests, including the Datacurve DeepSWE benchmark, Google reports up to a 65% reduction in token usage. Priced at $1.50 per million input tokens and $7.50 per million output tokens, it is positioned as a model for inference loops that run continuously rather than on demand.
With DeepSWE, the company has a success rate of 49 percent for 3.6 Flash compared to 37 percent for the previous version. MLE Bench saw its score rise from 49.7 percent to 63.9 percent, and Google’s GDPval-AA v2 test, which attempts to measure real-world knowledge work rather than coding puzzles, had a 3.6 Flash score of 1421, compared to 1349 for the older model.
Figma, Hebbia, and Harvey moved the model.
Figma has integrated 3.6 Flash into its prototyping infrastructure, and Matt Colyer, the company’s director of product engineering, says this model allows developers to iterate on designs faster without compromising output quality.
Legal technology platform Harvey and research tool Hebbia route data through a model of multimodal document work, including ingesting raw financial documents, parsing document structure, reading embedded charts, and creating draft reports for review.
Google also built client-side computational tools directly into the Gemini API and Gemini Enterprise platform, removing the custom intermediary software engineers previously built to make models work on the operating system.
The company reports that its OSWorld-Verified score has increased from 78.4% to 83.0%, and says its latest safeguards against chemical, biological, radiological, and nuclear abuse improve its resistance to jailbreaks without increasing the rejection rate of benign requests.
Cheap tier for high-volume background agents
Gemini 3.5 Flash-Lite targets a different job: high-volume document processing and agent search rather than deep inference. The Artificial Analysis Index measured the model at 350 output tokens per second, the fastest of the 3.5 series according to Google.
Priced at $0.3/$1 million for input tokens and $2.5 million/$1 million for output tokens, it is cheap enough for engineering teams to route simple, high-volume subagent requests to a minimal level of thinking and reserve a higher level of thinking for multi-step work.
In Google’s GDM-MRCR v2 long context test, Gemini 3.5 Flash-Lite recorded a success rate of 72.2 percent compared to 60.1 percent for the previous version, and its GDPval-AA v2 score nearly doubled from 642 to 1140. This model includes the same native computer usage tools as 3.6 Flash.
Separately, Google said Gemini 3.5 Pro continues to undergo partner testing ahead of its full release, and pre-training for the upcoming Gemini 4 architecture is already underway.
Gemini 3.5 Flash Cyber: Limited model for patching code
Automated vulnerability scanners are now surfacing flaws faster than most security teams can patch them, and that gap is where Google is positioning Gemini 3.5 Flash Cyber.
The model is built to examine and fix vulnerabilities in your code, and Google reports competitive performance with the Frontier model in the CyberGym benchmark (though it doesn’t publish the same detailed numbers as the consumer release).
Distribution will continue to be limited to governments and vetted partners through pilot programs, but Google puts this restriction in place as a safeguard against models that generate exploit code for offensive use.
Within Google’s CodeMender security agent, multiple instances of 3.5 Flash Cyber run in parallel and cross-check results with each other before creating a single remediation report that human reviewers approve.
Engineering teams looking to integrate these new models can access the models through the Gemini API through Google AI Studio, Android Studio, or the Gemini Enterprise Agent Platform. Consumers can also access the new model through the Gemini app, and 3.5 Flash-Lite will also roll out to Google Search.
SEE ALSO: Bristol-Myers Squibb buys Nvidia AI systems for drug discovery
Want to learn more about AI and big data from industry leaders? Check out the AI & Big Data Expos in Amsterdam, California, and London. This comprehensive event is part of TechEx and co-located with other major technology events such as Cyber Security & Cloud Expo. Click here for more information.
AI News is brought to you by TechForge Media. Learn about other upcoming enterprise technology events and webinars.

