Gemini 3.6 Flash improves efficiency by 17%

Google has launched Gemini 3.6 Flash together with Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber, each model aimed at different AI-agent application requirements. Gemini 3.6 Flash, the base model of the series, shows improved token efficiency: according to the Artificial Analysis Index it needs 17 % fewer output tokens than Gemini 3.5 Flash, and the DeepSWE benchmark from Datacurve reports a gap of up to 65 %. The model also records higher accuracy scores—49 % on DeepSWE versus 37 % on Gemini 3.5 Flash, 63,9 % on MLE Bench versus 49,7 %, and 83 % on OSWorld-Verified versus 78,4 %. Gemini 3.6 Flash includes built-in client-side compute capabilities via the Gemini API and Gemini Enterprise, and is priced at 1,50 USD per million input tokens and 7,50 USD per million output tokens, lowering overall task cost. It also ships with expanded Frontier Safety protections covering CBRN and cyber-attack scenarios, boosting resistance to jailbreaks and reducing unnecessary refusals for critical uses.
Gemini 3.5 Flash-Lite, the lightest and fastest model in the 3.5 line, delivers 350 output tokens per second according to the Artificial Analysis Index and is intended to accelerate low-latency agent workflows. Gemini 3.5 Flash Cyber, designed for cybersecurity, operates alongside CodeMender, Google’s code-security agent, to provide precise timing between the model and the agent infrastructure.
Google adds that Gemini 3.5 Pro is currently in testing with partners and will be released soon, while the largest pre-training run to date for Gemini 4 is already in progress.