The takeaway
Google isn't just shipping models — it's reshaping the economics of running AI agents in production, with token costs collapsing and Gemini 4 already in training.
Why it matters for builders
Google isn't just shipping models — it's reshaping the economics of running AI agents in production, with token costs collapsing and Gemini 4 already in training.
Google Launches Gemini 3.6 Flash, Slashes Token Costs by 65%, and Kicks Off Gemini 4 Training
In a late-night triple-drop, Google DeepMind just released three new Gemini models — including one that cuts token consumption by up to 65% on coding tasks — and simultaneously confirmed it has started its "most aggressive pre-training campaign in history" for Gemini 4. The message is unmistakable: Google is going all-in on making AI agents cheaper, faster, and smarter.
Three Models, One Night
On July 22, 2026, Google unveiled:
- Gemini 3.6 Flash — a next-gen workhorse that dramatically reduces cost and token usage
- Gemini 3.5 Flash-Lite — an ultra-fast, ultra-cheap variant hitting 350 tokens per second
- Gemini 3.5 Flash Cyber — a cybersecurity specialist built exclusively for vulnerability detection and remediation
Gemini 3.6 Flash: The Cost-Killer
The headline feature of 3.6 Flash is brute-force efficiency. According to the Artificial Analysis Index, it uses 17% fewer output tokens than its predecessor 3.5 Flash. On the DeepSWE coding benchmark, savings hit 65% — fewer detours, fewer redundant reasoning steps, fewer unnecessary tool calls per task.
Pricing reflects the efficiency gains: $1.50 per million input tokens and $7.50 per million output tokens — cheaper than 3.5 Flash despite stronger benchmark scores. Key performance jumps include:
| Benchmark | 3.5 Flash | 3.6 Flash |
|---|---|---|
| DeepSWE (coding) | 37% | 49% |
| MLE Bench (ML research) | 49.7% | 63.9% |
| OSWorld-Verified (computer use) | 78.4% | 83% |
| GDPVal-AA v2 (knowledge tasks) | — | +70 pts |
The model also shines at multi-agent orchestration, demonstrated through complex code migration and real-time 3D workflow generation via Gemini Canvas.
Flash-Lite: Small Model, Big Punch
Gemini 3.5 Flash-Lite hits 350 tokens per second at a staggering $0.30/M input, $2.50/M output — purpose-built for high-volume scenarios like bulk document processing and agentic search.
Most impressively, this lightweight model outperforms its larger sibling Gemini 3 Flash on both SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%). The intended architecture: 3.6 Flash as the "master brain" handling planning, with a fleet of Flash-Lite workers executing at scale.
Flash Cyber: Security-First AI
Gemini 3.5 Flash Cyber tackles a painful reality: AI finds vulnerabilities faster than existing systems can patch them. Integrated into Google's CodeMender platform, it has already set a new state-of-the-art on the CyberGym cybersecurity benchmark — using a lightweight model where competitors need heavyweight ones. It is not publicly available for now.
Gemini 4: The Real Bombshell
Beyond the three launches, Google confirmed it has internally kicked off its most aggressive pre-training run in history, targeting Gemini 4. Industry observers note that, following Google's typical ~6-month training cycle, Gemini 4 could arrive as early as late 2026.
Key takeaway: Google isn't just shipping models — it's reshaping the economics of running AI agents in production. With token costs collapsing, a lightweight model beating its larger predecessor, and Gemini 4 already in the oven, the cost of building on Google's AI stack is dropping faster than almost anyone predicted.
The Automation Brief
Read 5 AI stories instead of 50.
The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.
No noise. Unsubscribe anytime.
Editorial notes
Stefan Trbojevic
n8n Lab Editorial
22 July 2026
22 July 2026
Sources
AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.




