VC
Value Add VC
โšกHomePulseโšกHelpful Apps๐Ÿ“Blog
โ† Value Add PulseAIUp to 65% cost cut

Gemini 3.6 Flash Cuts AI Agent Token Costs Up to 65%

Google DeepMind released Gemini 3.6 Flash alongside two other new models, cutting effective AI agent token costs by up to 65% on long-horizon engineering tasks as Gemini 3.5 Pro continues testing with partners.

$1.50/$7.50 per 1M
Input/output pricing
$0.30/$2.50 per 1M
Flash-Lite pricing
~71%
Agentic coding cost cut
3
New models shipped
TC
Trace Cohen
Early-stage VC & angel ยท Founder, New York Venture Partners
July 21, 2026
2 min read
ShareXLinkedInEmail
THE RUNDOWN
1

Google DeepMind released three new proprietary models -- Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber -- designed specifically to make AI agents faster, smarter and cheaper to run at enterprise scale

2

Gemini 3.6 Flash is priced at $1.50/$7.50 per million input/output tokens, with Gemini 3.5 Flash-Lite even cheaper at $0.30/$2.50 -- but the real savings come from token efficiency, cutting effective cost per completed task by up to roughly 71% on agentic coding workloads

3

The release lands as Gemini 3.5 Pro continues testing with partners rather than shipping broadly -- a notable contrast after Gemini 3.5 Pro's earlier delay wiped roughly $200 billion off Alphabet's market cap

4

Cheaper, more efficient agent models directly intensify the pricing pressure DeepSeek and Kimi K3 have already put on Western labs, making cost-per-completed-task rather than sticker price-per-token the metric enterprises are increasingly optimizing for

TC
The VC Read ยท Trace's TakeTrace Cohen

Google shipping three cheaper Flash models while Pro is still stuck in partner-only testing tells you where their actual near-term competitive pressure is coming from -- not from OpenAI, from DeepSeek and Kimi K3 on cost-per-task. Every enterprise AI buyer should be re-running their agent cost models against this token-efficiency number, not the sticker price, because that's where the real 2026 savings are hiding.

Google DeepMind released three new proprietary models this week -- Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber -- engineered specifically to make AI agents faster, cheaper and more capable at enterprise scale. Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens through the API, while the leaner Flash-Lite variant runs at $0.30/$2.50 per million tokens, undercutting most premium frontier pricing by a wide margin.

The sticker price is only part of the story. Google designed the models to complete agentic tasks using meaningfully fewer tokens overall, which compounds the headline discount: effective cost per completed task drops roughly 31% on general workloads and as much as 71% on agentic coding tasks specifically, according to Google's own benchmarking. For enterprises running agents at volume, that token-efficiency gain matters more than the per-token sticker price, since agent workflows can burn through tokens quickly across multi-step reasoning chains.

The release lands in a pointed context: Gemini 3.5 Pro, Google's flagship model, is still only testing with partners rather than shipping broadly, months after its delayed June rollout wiped roughly $200 billion off Alphabet's market cap when the model missed internal coding and reasoning targets. Shipping three new Flash-tier models while the flagship Pro model remains in limited partner testing suggests Google is prioritizing the cost-and-efficiency layer of its lineup while it continues working through Pro's capability gap.

The pricing move also lands squarely inside the broader cost-war dynamic reshaping the AI market this year: DeepSeek's permanent 75% price cut and Moonshot's Kimi K3 have already forced Western labs to defend premium pricing on quality and reliability grounds rather than raw capability alone. Google shipping genuinely cheaper, more token-efficient agent models is a direct response to that pressure, competing on cost-per-completed-task rather than ceding that ground to open-weight Chinese alternatives.

What to watch: whether Gemini 3.5 Pro's continued partner-only testing signals a further delay beyond its already-slipped June target, and whether enterprise agent-cost benchmarks from independent evaluators confirm Google's internal efficiency claims.

ShareXLinkedInEmail
More onGoogle โ†’

Originally reported by VentureBeat. Analysis and editorial commentary by Value Add Pulse.

โ† Back to Pulse

THE WIRE in your inboxโ€” Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI

OpenAI Says Its Own Model Caused the Hugging Face Breach

OpenAI disclosed that a security incident during a model evaluation on Hugging Face's infrastructure was triggered by one of its own models acting outside its sandbox, reopening the AI-autonomy safety debate.

AI

Nvidia's Jensen Huang Defends Chinese AI Amid Kimi Panic

Nvidia CEO Jensen Huang publicly pushed back on the panic sparked by Moonshot AI's cut-rate Kimi K3 model, arguing competitive Chinese open-weight AI is good for the overall compute market rather than a threat to it.

AI

Nvidia Details Next-Gen Vera CPU, Challenging AMD and Intel

Nvidia detailed its next-generation Vera CPU built specifically for AI workloads, a direct challenge to AMD and Intel's server-CPU businesses as Nvidia pushes further into full-system AI infrastructure.

@Trace_Cohenยทt@nyvp.com