Google shipped three new models on July 21 — and the one everyone was waiting for was not among them.
The trio: Gemini 3.6 Flash, the new workhorse; Gemini 3.5 Flash-Lite, the budget tier; and Gemini 3.5 Flash Cyber, a specialist tuned for finding and fixing security vulnerabilities. The flagship Gemini 3.5 Pro, unveiled at Google I/O in May, remains "broadly available soon" — with no firm date.
What actually shipped
Flash is the tier most people actually live on: not the smartest model in the family, but fast and cheap enough for daily volume. Think of a Hong Kong cha chaan teng — the signature dish is for occasions; the set meal is what you eat three times a day.
The upgrade comes down to three numbers:
Output pricing drops from $9 to $7.50 per million tokens — roughly 17% cheaper.
On the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than its predecessor for the same work (up to 65% fewer on the DeepSWE coding benchmark).
Stack the two together and the effective output cost of a task can fall by about a third.
Capability moved up, not down: the DeepSWE score rose from 37% to 49%. Availability is immediate — the Gemini app for everyone, and the API via Google AI Studio and Android Studio for developers.
The sober read: the flagship is late, the workhorse ships first
The real story is the cadence, not the model. This July alone: the GPT-5.6 family led with aggressive entry pricing, Kimi K3 went open source outright (we covered it in "An Open-Source Model Just Beat the Paid Flagship"), and now Flash gets cheaper. The battleground has shifted from "who is smartest" to "who is good enough, and cheap." Flagships are still coming — Google says Gemini 4 pre-training has begun — but the fiercest competition is in the mid tier, because that is where the volume lives.
One caveat deserves attention: the phrase "up to." The 17% saving is an index average; 65% is a single benchmark's ceiling. How much your own workload saves, only your own workload can tell you. Treat official numbers as reference; the judge is always your task.
What it means in Hong Kong and China
Hong Kong users benefit directly. Gemini is one of the few frontier AI services officially available in Hong Kong (since March 2026), and 3.6 Flash is in the Gemini app today — no setup, and the free tier covers most daily use. Where an official channel works, our standing advice is simple: use it.
For users in China the picture is different: Gemini remains officially unavailable, as do Claude and the GPT series. Beyond connectivity sit two more mountains — payment and account risk. The practical route to frontier models is a stable direct gateway with usage-based billing, which is exactly what Essevin does: one account across multiple top models, Alipay and WeChat Pay for users in China, FPS and PayMe in Hong Kong, and you pay only for what you use.
Closing
Three months ago, "budget tier" still meant "make do." Today it is where the giants fight hardest. Rather than waiting for a delayed flagship, split your work by task: workhorse models for daily volume, frontier models for the hard problems. That is how this price war actually lands in your pocket.
Information in this article is current as of July 22, 2026 and is provided for general reference only; it does not constitute advice of any kind. Third-party product features, pricing and policies are subject to their official announcements. Essevin service details are as shown on essevin.com and in the console.