Blazarael
THE INDEXv0.2 — WK 32 / 2026 · UPDATED WITH EVERY LAUNCH

Frontier Model Tracker

The living dataset of frontier-class models: who ships, what it costs, how much context it takes, what the verified benchmarks actually say — and how far open weights sit behind the closed frontier. Every figure carries a primary source; every blank is deliberate.

6FRONTIER RELEASES IN 6 WEEKS
×2CHIP DEPLOYMENT EVERY 9 MONTHS
1MTOKENS — THE NEW CONTEXT FLOOR
2/8TRACKED MODELS ARE OPEN WEIGHTS
01 · THE TABLE⬡ = OPEN WEIGHTS · “—” = NOT PUBLISHED
FRONTIER MODELS — WK 32 / 2026SOURCES [1][2]
ModelLabReleasedContext$/M in · outWeightsTrajectory
CLAUDE OPUS 5ANTHROPIC2026-07-241M$5 / $25CLOSEDACCELERATING
GPT-5.6 SOLOPENAI2026-07CLOSEDACCELERATING
GEMINI 3.6 FLASHGOOGLE2026-07-21CLOSEDSTEADY
GEMINI 3.1 PROGOOGLE20261.05M$2 / $12CLOSEDSTEADY
GROK 4.5XAI2026-07CLOSEDSTEADY
DEEPSEEK-V4-FLASHDEEPSEEK2026-07-31⬡ OPENACCELERATING
QWEN 3.8 MAXALIBABA2026-08⬡ OPENACCELERATING
GPT-5.1OPENAI2025-11CLOSEDSTEADY
02 · VERIFIED AGENTIC BENCHMARKSTHE HANDS, NOT THE HEAD

Knowledge benchmarks are saturating; the spreads that matter in 2026 are agentic — real repository fixes, real desktops, real web navigation. These are the only numbers with a published, independent trail. Where a model is missing, no verifiable score exists yet.

SWE-BENCH PROREAL REPOSITORY FIXES · % — SOURCE [1]
CLAUDE OPUS 579.2
GPT-5.6 SOL64.6
OSWORLD 2.0OPERATING A REAL DESKTOP · % — SOURCE [1]
CLAUDE OPUS 570.6
BROWSECOMPWEB NAVIGATION · % — SOURCE [1]
GPT-5.6 SOL90.4

FIXED 0–100 SCALE ACROSS ALL BENCHMARKS. ONLY INDEPENDENTLY PUBLISHED NUMBERS ARE PLOTTED — NO VENDOR SLIDES, NO ESTIMATES.

03 · OPEN WEIGHTS2 OF 8 TRACKED MODELS

Two of the six frontier-class releases of the July wave shipped with downloadable weights [2]. That cadence — days behind the closed labs, not years — is the story: open weights now arrive inside the same news cycle as the frontier they chase.

⬡ DEEPSEEK-V4-FLASH · DEEPSEEK

Released Jul 31 — seven days after Opus 5

  • Open weights, released inside the July frontier wave [2]
  • Trajectory: accelerating — DeepSeek ships on a quarterly cadence
  • Independent agentic benchmark runs: not yet published — tracked
⬡ QWEN 3.8 MAX · ALIBABA

Released Aug 2 — the sixth release in six weeks

  • Open weights; closes the July wave [2]
  • Trajectory: accelerating — second open release of the wave
  • Independent agentic benchmark runs: not yet published — tracked
04 · THE GAP — OPEN VS CLOSEDMEASURED, NOT VIBES

What can be measured today: cadence and price. On cadence, the gap has collapsed — open-weights frontier models now release days after closed ones (Jul 24 → Jul 31 → Aug 2), where the 2023–24 lag was measured in quarters [2]. On price, open releases cap what closed labs can charge for the middle of the market: the $2/M-input floor set by Gemini 3.1 Pro [1] exists because a free alternative is always seven days away.

What cannot be measured yet: capability. No independent, like-for-like agentic benchmark run of the July open-weights releases has been published. The verified numbers in section 02 all belong to closed models. We will plot the open bars — hatched, on the same 0–100 scale — the day a source exists, and not before.

WHAT WE DON'T KNOW
  • Where DeepSeek-V4-Flash and Qwen 3.8 Max land on SWE-bench Pro and OSWorld
  • Real hosting cost per M tokens for optimized open-weights serving
  • Whether open cadence survives if the capex wave slows
05 · LAUNCH LOGRELEASES · BENCHMARK UPDATES · SPECIAL ANNOUNCEMENTS
JUL 21Gemini 3.6 Flash ships.
JUL 24Claude Opus 5: 1M context, 70.6% OSWorld, $5/$25 per M tokens.
JUL —Grok 4.5 and GPT-5.6 Sol land inside the same window.
JUL 31DeepSeek-V4-Flash releases as open weights.⬡ OPEN
AUG 02Qwen 3.8 Max releases as open weights — the sixth frontier-class model in six weeks.⬡ OPEN
AUG 02OpenAI's unreleased Astra reportedly solves 10 open math problems (~$2,000 compute).
06 · METHOD & SOURCESPRIMARY SOURCE OR IT DOESN'T RUN

The Index only records figures with a public, citable trail — vendor announcements count for existence and pricing, never for capability. Blanks mean “not published”, not “zero”. Corrections are made in public at this URL.