Google DeepMind Launches Gemini 3.7 Flash at Half-Price as Flagship Gemini 3.5 Pro Delay Continues
Google DeepMind is continuing its rapid-fire release cycle for its lightweight "workhorse" models, even as its highly anticipated flagship model continues to face severe bottlenecks. On Thursday, August 13, 2026, Google officially launched Gemini 3.7 Flash, its most intelligent and agent-focused lightweight model to date. Arriving just three weeks after the July 21 release of Gemini 3.6 Flash, Gemini 3.7 Flash cuts pricing by 50% (starting at $0.75 per million input tokens) while delivering massive gains in software engineering, web development, and autonomous agent tasks.
Gemini 3.7 Flash: The New Price-Performance Benchmark
Gemini 3.7 Flash has been engineered to directly target coding and agentic workflows, significantly outperforming previous Flash iterations:
- Blistering Agentic Coding: The model achieved a massive 65.3% score on DeepSWE (Google's software engineering benchmark) and a 43.6% code quality score, making it a highly competitive option for autonomous developer agents.
- Aggressive Pricing: Google has halved the pricing from Gemini 3.6 Flash, setting the rate at $0.75 per million input tokens and $3.00 per million output tokens, mounting intense pricing pressure on OpenAI's GPT-5.4/5.5 and Anthropic's Claude Sonnet 5.
- Long-Context Superiority: The model maintains Google's industry-leading 2-million-token context window, proving highly effective at long-context reasoning and multi-file code analysis.
The Flagship Problem: Gemini 3.5 Pro Remains Missing in Action
While Google's rapid iteration of the Flash series has kept it highly competitive at the edge and in developer APIs, the company's lack of a true frontier-class flagship has become a glaring strategic vulnerability. Gemini 3.5 Pro, originally promised at I/O 2026 for a June release, remains delayed with no clear launch date in sight:
- Internal Builds Scrapped: Reports indicate that Google DeepMind was forced to scrap its original Gemini 3.5 Pro build after it suffered from critical tool-calling failures and fell short of internal coding benchmarks.
- Organizational Restructuring: In response to the delays, Google reorganized engineering teams within DeepMind to consolidate its AI coding efforts under consolidated leadership.
- Competitive Disadvantage: With OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 actively dominating enterprise reasoning workloads, Google's inability to ship its top-tier model has forced it to rely entirely on lightweight, low-margin models to defend its developer market share.
Verbatim Quotes
"Google has rolled out Gemini 3.7 Flash, positioning it as the company's most capable workhorse model yet for coding and AI agent tasks, even as it continues to withhold a release date for its long-delayed flagship model, Gemini 3.5 Pro1... Bloomberg had earlier reported that Gemini 3.5 Pro's delay stemmed from an internal push to strengthen the model's coding capabilities, with Google reorganising engineering teams within DeepMind to consolidate its AI coding efforts."
— Free Press Journal, Gemini 3.7 Flash Launched; Google Still Silent On Delayed Flagship Model
"The launch comes as Gemini 3.5 Pro, Google's actual flagship promised at May's I/O, remains delayed after DeepMind reportedly scrapped its original build over tool-calling failures... The bigger promise, Gemini 3.5 Pro, still sits and waits - and if you're choosing a model today, nobody pays you for believing in a release that hasn't shipped."
— Startup Fortune, Google launches Gemini 3.7 Flash and halves its price as Gemini 3.5 Pro waits
-
An instance of The delay of a flagship frontier model forces a defensive flood of cheap, lightweight workhorse releases. — Google is actively flooding the market with cheap, fast workhorse models to defend its developer footprint while its flagship model sits delayed. ↩︎