Google DeepMind released three new proprietary models this week -- Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber -- engineered specifically to make AI agents faster, cheaper and more capable at enterprise scale. Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens through the API, while the leaner Flash-Lite variant runs at $0.30/$2.50 per million tokens, undercutting most premium frontier pricing by a wide margin.
The sticker price is only part of the story. Google designed the models to complete agentic tasks using meaningfully fewer tokens overall, which compounds the headline discount: effective cost per completed task drops roughly 31% on general workloads and as much as 71% on agentic coding tasks specifically, according to Google's own benchmarking. For enterprises running agents at volume, that token-efficiency gain matters more than the per-token sticker price, since agent workflows can burn through tokens quickly across multi-step reasoning chains.
The release lands in a pointed context: Gemini 3.5 Pro, Google's flagship model, is still only testing with partners rather than shipping broadly, months after its delayed June rollout wiped roughly $200 billion off Alphabet's market cap when the model missed internal coding and reasoning targets. Shipping three new Flash-tier models while the flagship Pro model remains in limited partner testing suggests Google is prioritizing the cost-and-efficiency layer of its lineup while it continues working through Pro's capability gap.
The pricing move also lands squarely inside the broader cost-war dynamic reshaping the AI market this year: DeepSeek's permanent 75% price cut and Moonshot's Kimi K3 have already forced Western labs to defend premium pricing on quality and reliability grounds rather than raw capability alone. Google shipping genuinely cheaper, more token-efficient agent models is a direct response to that pressure, competing on cost-per-completed-task rather than ceding that ground to open-weight Chinese alternatives.
What to watch: whether Gemini 3.5 Pro's continued partner-only testing signals a further delay beyond its already-slipped June target, and whether enterprise agent-cost benchmarks from independent evaluators confirm Google's internal efficiency claims.