Frontier Model Tracker
The living dataset of frontier-class models: who ships, what it costs, how much context it takes, what the verified benchmarks actually say — and how far open weights sit behind the closed frontier. Every figure carries a primary source; every blank is deliberate.
| Model | Lab | Released | Context | $/M in · out | Weights | Trajectory |
|---|---|---|---|---|---|---|
| CLAUDE OPUS 5 | ANTHROPIC | 2026-07-24 | 1M | $5 / $25 | CLOSED | ACCELERATING |
| GPT-5.6 SOL | OPENAI | 2026-07 | — | — | CLOSED | ACCELERATING |
| GEMINI 3.6 FLASH | 2026-07-21 | — | — | CLOSED | STEADY | |
| GEMINI 3.1 PRO | 2026 | 1.05M | $2 / $12 | CLOSED | STEADY | |
| GROK 4.5 | XAI | 2026-07 | — | — | CLOSED | STEADY |
| DEEPSEEK-V4-FLASH | DEEPSEEK | 2026-07-31 | — | — | ⬡ OPEN | ACCELERATING |
| QWEN 3.8 MAX | ALIBABA | 2026-08 | — | — | ⬡ OPEN | ACCELERATING |
| GPT-5.1 | OPENAI | 2025-11 | — | — | CLOSED | STEADY |
Knowledge benchmarks are saturating; the spreads that matter in 2026 are agentic — real repository fixes, real desktops, real web navigation. These are the only numbers with a published, independent trail. Where a model is missing, no verifiable score exists yet.
FIXED 0–100 SCALE ACROSS ALL BENCHMARKS. ONLY INDEPENDENTLY PUBLISHED NUMBERS ARE PLOTTED — NO VENDOR SLIDES, NO ESTIMATES.
Two of the six frontier-class releases of the July wave shipped with downloadable weights [2]. That cadence — days behind the closed labs, not years — is the story: open weights now arrive inside the same news cycle as the frontier they chase.
Released Jul 31 — seven days after Opus 5
- Open weights, released inside the July frontier wave [2]
- Trajectory: accelerating — DeepSeek ships on a quarterly cadence
- Independent agentic benchmark runs: not yet published — tracked
Released Aug 2 — the sixth release in six weeks
- Open weights; closes the July wave [2]
- Trajectory: accelerating — second open release of the wave
- Independent agentic benchmark runs: not yet published — tracked
What can be measured today: cadence and price. On cadence, the gap has collapsed — open-weights frontier models now release days after closed ones (Jul 24 → Jul 31 → Aug 2), where the 2023–24 lag was measured in quarters [2]. On price, open releases cap what closed labs can charge for the middle of the market: the $2/M-input floor set by Gemini 3.1 Pro [1] exists because a free alternative is always seven days away.
What cannot be measured yet: capability. No independent, like-for-like agentic benchmark run of the July open-weights releases has been published. The verified numbers in section 02 all belong to closed models. We will plot the open bars — hatched, on the same 0–100 scale — the day a source exists, and not before.
- Where DeepSeek-V4-Flash and Qwen 3.8 Max land on SWE-bench Pro and OSWorld
- Real hosting cost per M tokens for optimized open-weights serving
- Whether open cadence survives if the capex wave slows
The Index only records figures with a public, citable trail — vendor announcements count for existence and pricing, never for capability. Blanks mean “not published”, not “zero”. Corrections are made in public at this URL.