Blazarael
CAPITAL & CODEISSUE Nº 0015 MIN READ

The Price of Cognition

In July the price sheet became the most-read document in AI. A million tokens of context is table stakes; the real question is what a unit of thinking costs — and who can afford to run it all day.


Every frontier release in July shipped with the same second page: the meter. Claude Opus 5 prices at $5 per million input tokens and $25 per million output, with batch processing at half price [1]. Gemini 3.1 Pro undercuts at $2 per million in and $12 out for prompts up to 200K tokens [1]. The capability race gets the headlines; the pricing table is where the industry actually competes.

The line item nobody budgeted

For a business running agents continuously, these numbers stop being API trivia and become cost of goods sold. An agent that reads a full million-token context pays $5 before it does anything; every long answer bills at output rates. The engineering conversation flips from ‘can the model do it’ to ‘what is worth the tokens’ — and procurement starts asking why cognition is the one input nobody forecasts.

The pressure on that floor is structural. Within days of the closed-lab releases, DeepSeek-V4-Flash and Qwen 3.8 Max shipped as open-weights alternatives [2] — and every open release resets the ceiling on what closed labs can charge for the middle of the market. The premium narrows to the frontier edge: agentic reliability, desktop hands, verified benchmarks.

SOURCES — PRIMARY OR IT DOESN'T RUN[1] Claude Opus 5 — specs, benchmarks, pricing (incl. Gemini figures)[2] LLM Stats — model releases tracker (July–August 2026)