VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: Gemini 3.6 Flash Cuts AI Agent Token Costs Up to 65%
Value Add VC/Pulse/AIUp to 65% cost cut

Gemini 3.6 Flash Cuts AI Agent Token Costs Up to 65%

Google DeepMind released Gemini 3.6 Flash alongside two other new models, cutting effective AI agent token costs by up to 65% on long-horizon engineering tasks as Gemini 3.5 Pro continues testing with partners.

By the Numbers

$1.50/$7.50 per 1M
Input/output pricing
$0.30/$2.50 per 1M
Flash-Lite pricing
~71%
Agentic coding cost cut
3
New models shipped
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
July 21, 2026
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Google DeepMind released three new proprietary models -- Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber -- designed specifically to make AI agents faster, smarter and cheaper to run at enterprise scale

2

Gemini 3.6 Flash is priced at $1.50/$7.50 per million input/output tokens, with Gemini 3.5 Flash-Lite even cheaper at $0.30/$2.50 -- but the real savings come from token efficiency, cutting effective cost per completed task by up to roughly 71% on agentic coding workloads

3

The release lands as Gemini 3.5 Pro continues testing with partners rather than shipping broadly -- a notable contrast after Gemini 3.5 Pro's earlier delay wiped roughly $200 billion off Alphabet's market cap

4

Cheaper, more efficient agent models directly intensify the pricing pressure DeepSeek and Kimi K3 have already put on Western labs, making cost-per-completed-task rather than sticker price-per-token the metric enterprises are increasingly optimizing for

TC

The VC Read · Trace's Take

Trace Cohen

Google shipping three cheaper Flash models while Pro is still stuck in partner-only testing tells you where their actual near-term competitive pressure is coming from -- not from OpenAI, from DeepSeek and Kimi K3 on cost-per-task. Every enterprise AI buyer should be re-running their agent cost models against this token-efficiency number, not the sticker price, because that's where the real 2026 savings are hiding.

Analysis

Google DeepMind released three new proprietary models this week -- Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber -- engineered specifically to make AI agents faster, cheaper and more capable at enterprise scale. Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens through the API, while the leaner Flash-Lite variant runs at $0.30/$2.50 per million tokens, undercutting most premium frontier pricing by a wide margin.

The sticker price is only part of the story. Google designed the models to complete agentic tasks using meaningfully fewer tokens overall, which compounds the headline discount: effective cost per completed task drops roughly 31% on general workloads and as much as 71% on agentic coding tasks specifically, according to Google's own benchmarking. For enterprises running agents at volume, that token-efficiency gain matters more than the per-token sticker price, since agent workflows can burn through tokens quickly across multi-step reasoning chains.

The release lands in a pointed context: Gemini 3.5 Pro, Google's flagship model, is still only testing with partners rather than shipping broadly, months after its delayed June rollout wiped roughly $200 billion off Alphabet's market cap when the model missed internal coding and reasoning targets. Shipping three new Flash-tier models while the flagship Pro model remains in limited partner testing suggests Google is prioritizing the cost-and-efficiency layer of its lineup while it continues working through Pro's capability gap.

The pricing move also lands squarely inside the broader cost-war dynamic reshaping the AI market this year: DeepSeek's permanent 75% price cut and Moonshot's Kimi K3 have already forced Western labs to defend premium pricing on quality and reliability grounds rather than raw capability alone. Google shipping genuinely cheaper, more token-efficient agent models is a direct response to that pressure, competing on cost-per-completed-task rather than ceding that ground to open-weight Chinese alternatives.

What to watch: whether Gemini 3.5 Pro's continued partner-only testing signals a further delay beyond its already-slipped June target, and whether enterprise agent-cost benchmarks from independent evaluators confirm Google's internal efficiency claims.

ShareXLinkedInEmail

More on

Google →Google DeepMind →

Reported by VentureBeat · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 14, 2026

OpenAI Sheds Senior Execs in Pre-IPO Shakeup

Illustration for: OpenAI Sheds Senior Execs in Pre-IPO Shakeup
AI

OpenAI Sheds Senior Execs in Pre-IPO Shakeup

OpenAI has lost its chief revenue officer, its longtime COO and several senior leaders within days of each other, as co-founder Greg Brockman consolidates operating control ahead of a planned public listing.

AI· Aug 13, 2026

Anthropic's CFO Starts Courting IPO Investors

Illustration for: Anthropic's CFO Starts Courting IPO Investors
AI

Anthropic's CFO Starts Courting IPO Investors

Anthropic CFO Krishna Rao has begun early, informal meetings with prospective IPO investors, though he has not discussed valuation -- the $2 trillion figure circulating on Wall Street comes from investors' own math, not from Anthropic.

AI· Aug 13, 2026

Gemini 3.7 Flash Launches With 50% Price Cut for Coding

Illustration for: Gemini 3.7 Flash Launches With 50% Price Cut for Coding
AI

Gemini 3.7 Flash Launches With 50% Price Cut for Coding

Google released Gemini 3.7 Flash just three weeks after 3.6 Flash, cutting introductory API pricing in half while improving coding, debugging and enterprise-automation benchmarks over its predecessor.

@Trace_Cohen·t@nyvp.com