On July 9, in a single day: OpenAI shipped the GPT-5.6 family, Grok 4.5 opened to the public, and Meta released Muse Spark 1.1 — its first paid model ever. A week later, Kimi K3 arrived as an open-source release and jumped to first place on the frontend coding leaderboard within hours. Four days after that — today — Claude Fable 5's new subscription policy took effect.
In the same month, Google's Gemini 3.5 Pro missed its own launch date for the third time.
A year's worth of drama, compressed into thirty days. Which brings back the eternal question: which model should you actually use right now?
This is an opinionated answer: the mainstream models of July 20, 2026, ranked from GOAT to cooked, each with reasons and honest caveats. First, how we measure; then the tiers; then the four traps to avoid when reading any leaderboard — including this one.
Three yardsticks
Choosing a model is like choosing a tutor for your kids. Exam results matter, but no parent stops there — you also check the fees, and whether the tutor actually teaches in your district.
So this ranking uses three yardsticks:
Overall capability — anchored on the Artificial Analysis Intelligence Index (v4.1, a composite of nine evaluations including GDPval, Terminal-Bench and GPQA Diamond), supplemented by single-domain leaderboards.
Price — what the same work costs you. July was a price-war month, so this axis moved more than the first.
Availability — can you use it today, or is it perpetually "coming soon". For readers in Hong Kong and China, this axis also includes whether the vendor will take your business at all.
One fact before we start: the top ten models on the composite index sit within 8.5 points of each other. The summit is crowded. This ranking is a snapshot, not a verdict.
The tiers
GOAT: Claude Fable 5
Anthropic's Mythos-class flagship — released June 9, globally available July 1 — currently leads the Artificial Analysis composite at 59.9. Long-context reasoning, agentic work and software engineering are all front-of-pack. When you don't want to compromise, this is the default answer.
The caveats are equally clear: it sits in the most expensive bracket, and its subscription terms have shifted repeatedly within a single month — from July 20 it is included in Max and Team Premium plans, but at 50% of standard usage limits. And an honest note: second place is one point behind. GOAT is a fact, not a landslide.
Top tier: GPT-5.6 Sol and Kimi K3
GPT-5.6 Sol (OpenAI, July 9): 58.9 on the composite — practically breathing down the leader's neck. Its pitch isn't "smartest" but "same-tier results with fewer tokens at lower cost", plus inference on Cerebras at up to 750 tokens per second. Starting with GPT-5.6, OpenAI names its tiers Sol, Terra and Luna, each advancing on its own cadence.
Kimi K3 (Moonshot AI, July 16): the month's biggest variable. A ~2.8-trillion-parameter MoE on an open-source track with a 1M-token context window, third on the composite at 57.1 — and within hours of launch it jumped from 18th to 1st on Frontend Code Arena, taking six of seven sub-categories. Roughly one third the price of closed flagships, with a free tier at kimi.com.
Respective caveats: Sol ties you deep into the OpenAI ecosystem; K3 still trails the composite lead by more than two points — the single-domain crown is real, "beats the frontier across the board" is headline inflation.
Daily drivers
Not the top of the leaderboard — but price and reliability make these the workhorses most people actually run.
Model | In one line | Honest caveat |
|---|---|---|
Claude Opus 4.8 | Battle-tested workhorse with the most mature tooling | K3 has nearly caught its composite score; the halo moved to Fable 5 |
Claude Sonnet 5 | Near-Opus 4.8 performance at value pricing (June 30) | Under a month old; the intro pricing is time-limited |
GPT-5.6 Terra | The value pick for everyday work, ~80% of flagship power | Hard tasks still need Sol |
Grok 4.5 | Coding/agentic specialist (July 8), priced ~60% below its class | Composite around 54 — not an all-rounder |
DeepSeek V4 | The price butcher of the open-source camp | K3 stole its spotlight |
GLM-5.2 | Z.ai's 744B MoE, top open model for agentic coding | Little name recognition outside China |
A side note on Qwen3.6: the freest license and single-GPU deployability make it the local-deployment favourite — a different race from mainstream chat flagships.
The awkward tier
Muse Spark 1.1 (Meta, July 9): Meta's first paid model. 1M context, agentic focus, priced at roughly one tenth of flagship level. The problem is Meta's own comparison chart: GPT-5.5, Opus 4.8, Gemini 3.1 Pro — all previous-generation flagships. When you benchmark against yesterday's champions, good numbers deserve a discount. Genuinely cheap; whether it survives a month of real-world use, we'll see in August.
Gemini 3.5 Flash (Google): a decent model, and for Hong Kong readers it carries special weight — Gemini is one of the very few frontier channels that officially serves Hong Kong. That's the "hard to discard" half. The "hard to love" half: with the flagship missing, a mid-weight model is left alone to fight other vendors' Sols and Fables. Not a fair fight.
Cooked: Gemini 3.5 Pro
The only model in this ranking that ranks bottom for not existing.
Originally due in June; delayed three times since. Multiple reports say the original build was scrapped and pre-training restarted over structural problems in coding and recursive tool-calling. There is still no gemini-3.5-pro in the public API. Google is now the only major lab without a 2026 flagship in production.
In fairness: when it finally ships, it may jump straight back to top tier. "Cooked" describes its existence status on July 20, 2026 — not Google's ceiling. Which sets up trap number three below.
Four traps when reading leaderboards
1. Reading a single-domain board as the overall board. "K3 beat Fable 5" correctly reads "in frontend coding". Single-domain boards pick tools; the composite picks your default. Two boards, two different questions.
2. Inferring capability from price. July was a price war: Sol costs less than last generation's flagship and is stronger; Grok 4.5 being 60% cheaper doesn't mean 60% weaker. Price reflects strategy, not a capability scale.
3. Waiting for the perfect model. People waiting for Gemini 3.5 Pro have now waited three times. Your work won't wait. Use today's best combination for today's tasks and switch when a new king is actually crowned — with the workflow in trap four, switching costs you nothing.
4. Treating any tier list as gospel. Including this one. The most reliable test remains: take your own real task — a contract to revise, code to debug, minutes to summarise — feed the same prompt to two or three models, and compare outputs side by side. Fifteen minutes of your own testing beats ten tier lists, this one included.
The shelf life of this ranking is one month
If July taught anything, it's this: don't marry a model. A triple launch day, an open-source upset and a flagship no-show all happened within thirty days. Assign different subjects to different tutors and re-read the report card monthly — that's the baseline posture for using AI in 2026.
The August edition lands this time next month: K3 and Muse Spark will have a month of real-world mileage behind them, and Gemini 3.5 Pro — we'll see which tier it occupies. Who's your GOAT? Tell us in the comments; the August edition will include a reader vote.
Information in this article is current as of July 20, 2026 and is provided for general reference only; it does not constitute advice of any kind. Third-party product features, pricing and policies are subject to their official announcements. Essevin service details are as shown on essevin.com and in the console.