Google's Gemini Flash Models Arrive, But the Pro It Promised Is Still Missing in Action 🫥
Google released three new AI models on Wednesday — Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber — while continuing to withhold the Gemini 3.5 Pro variant the company pledged to ship roughly a month after unveiling Gemini 3.5 Flash at Google I/O in May. According to Bloomberg, Google held back the Pro version because it missed internal targets, particularly on coding tasks, with a late-June attempt to refresh training data producing disappointing results. Alphabet shares fell roughly 4.4% on the report, erasing an estimated $200 billion in market capitalization in a single session. The last Pro-tier model Google shipped was Gemini 3.1 Pro, the successor to Gemini 3, which debuted in February.
Gemini 3.6 Flash is the flagship of the new batch. According to the Artificial Analysis Index, it uses 17% fewer output tokens than 3.5 Flash, with pricing set at $1.50 per million input tokens and $7.50 per million output tokens, down from $9 per million output tokens for the prior version. On standardized benchmarks, Gemini 3.6 Flash scored 49% on DeepSWE v1.1, which measures long-horizon software engineering, compared with 37% for 3.5 Flash, and 63.9% on MLE-Bench versus 49.7%. It also led OSWorld-Verified at 83.0%, ahead of Claude Sonnet 5 at 81.2% and GPT-5.6 Luna at 72.6%.
Competitors continue to lead in other categories. GPT-5.6 Luna recorded 67% on DeepSWE and 84.7% on Terminal-Bench 2.1, while Claude Sonnet 5 topped GDPval-AA v2, a knowledge-work benchmark scored on an Elo-style rating scale, with 1607 versus Gemini 3.6 Flash's 1421. The Flash line is positioned as Google's speed-optimized tier aimed at AI agents that operate semi-autonomously across tasks such as document handling and data pipelines, while Pro models are marketed as heavier, pricier systems for complex reasoning.
Hands-on coding tests of Gemini 3.6 Flash produced mixed results, with initial outputs yielding an unusable file in which HTML was not properly formatted and elements failed to render correctly. A subsequent pass using Deepseek identified 11 bugs and implemented 8 fixes, producing a playable result that suggested the model's core reasoning was sound but execution was lacking.
Share Article
Quick Info
Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.
See our Terms of Service, Privacy Policy, and Editorial Policy.