Google has expanded its Gemini family with two new models aimed at businesses building AI agents at scale. The latest releases focus on reducing latency, lowering token costs, and improving performance for enterprise workloads, while also introducing a specialized cybersecurity model for vulnerability remediation.
As organizations increasingly rely on autonomous AI agents to automate workflows, efficiency has become just as important as raw intelligence. Google’s newest models are designed to strike that balance by delivering stronger reasoning while consuming fewer tokens.
Why Token Efficiency Matters
For enterprise AI deployments, every generated token carries a cost. Background agents that continuously process documents, write code, analyze data, or automate repetitive tasks may execute thousands of requests every hour. Even small reductions in token usage can translate into significant infrastructure savings.
Rather than simply increasing model size, Google has focused on improving reasoning efficiency. The result is a family of models tailored for different workloads instead of a one-size-fits-all approach.
Gemini 3.6 Flash Improves Performance While Cutting Costs
Gemini 3.6 Flash serves as Google’s primary model for coding, multimodal reasoning, and complex enterprise tasks. According to Google, the model generates approximately 17% fewer output tokens than Gemini 3.5 Flash while maintaining or improving response quality.
In certain software engineering benchmarks, token usage dropped by as much as 65%, making long-running reasoning tasks considerably less expensive. The model is priced at $1.50 per million input tokens and $7.50 per million output tokens, positioning it as an attractive option for organizations operating AI agents around the clock.
Google also reports meaningful benchmark improvements. On the DeepSWE coding benchmark, Gemini 3.6 Flash achieved a 49% success rate compared to 37% for the previous generation. Performance also improved substantially on machine learning engineering evaluations and knowledge-work benchmarks designed to better reflect real-world enterprise use cases.
Enterprise Adoption Is Already Underway
Several enterprise software companies have already begun integrating Gemini 3.6 Flash into their products.
Design platform Figma is using the model within its prototyping workflow to accelerate design iterations while maintaining output quality.
Meanwhile, AI-powered legal platform Harvey and research company Hebbia are leveraging Gemini for multimodal document analysis. Their systems process financial filings, interpret charts and diagrams, understand document structure, and generate draft reports that human reviewers can refine.
Google has also integrated computer-use capabilities directly into the Gemini API and Gemini Enterprise platform, eliminating the need for developers to build custom middleware that allows AI models to interact with operating systems.
Alongside these productivity improvements, Google says it has strengthened safeguards against chemical, biological, radiological, and nuclear misuse while maintaining usability for legitimate enterprise requests.
Gemini 3.5 Flash-Lite Prioritizes Speed
While Gemini 3.6 Flash targets more demanding reasoning tasks, Gemini 3.5 Flash-Lite is optimized for high-volume workloads where speed and affordability matter most.
Google positions Flash-Lite for document processing, search, data extraction, and lightweight AI agents that don’t require extensive reasoning.
The model delivers up to 350 output tokens per second, making it the fastest model in the Gemini 3.5 family. Pricing is significantly lower at $0.30 per million input tokens and $2.50 per million output tokens, allowing organizations to reserve larger models only for more complex requests.
Google also reports notable improvements in long-context comprehension and knowledge evaluation benchmarks, while Flash-Lite retains the same native computer-use functionality found in Gemini 3.6 Flash.
A Dedicated AI Model for Cybersecurity
Google also introduced Gemini 3.5 Flash Cyber, a specialized model built specifically for software vulnerability analysis and code remediation.
Unlike Google’s general-purpose models, Flash Cyber is designed to identify security flaws, validate vulnerabilities, and recommend code fixes. The company says the model performs competitively on cybersecurity benchmarks, although detailed benchmark results have not yet been publicly released.
Access to Flash Cyber remains restricted to government organizations and approved partners as part of a pilot program. Google says this limitation helps reduce the risk of the model being used to generate offensive exploit code.
Within Google’s internal CodeMender security agent, multiple instances of Flash Cyber analyze the same vulnerability simultaneously, compare their findings, and generate a consolidated remediation report that is ultimately reviewed by a human security expert.
Looking Ahead
Google’s latest Gemini releases highlight an important shift in enterprise AI development. Rather than focusing solely on building larger and more capable models, companies are increasingly optimizing for efficiency, throughput, and operational cost.
By offering separate models for advanced reasoning, high-volume automation, and cybersecurity, Google is giving organizations more flexibility to match AI capabilities with specific workloads while keeping infrastructure expenses under control.
As enterprise AI adoption continues to accelerate, improvements in token efficiency and specialized task performance may prove just as valuable as gains in raw model intelligence.


