Frontier models: intelligence vs cost per task

Lines connect measured and explicitly modeled effort points. Filled dots are measured by Artificial Analysis; hollow dots are formula-derived. Snapshot: .

Cost re-baselined 2026-08-03. AA revised its cost-per-task basis since the 2026-07-29 snapshot — Intelligence Index scores came back identical while $/task rose 12.5–24.3% (Anthropic 12.5–14.5%, Gemini 16.5%, Grok 16.8%, OpenAI 18.8–24.3%), with per-token list prices unchanged. Compare costs within this snapshot only.

measured (Artificial Analysis) formula-derived effort level
ModelEffortIndex$/taskPointSource / basis

Configurations. AA's Fable point is specifically Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback). OpenAI says max gives GPT-5.6 more reasoning time than xhigh, while ultra coordinates four agents in parallel by default. Ultra is a separate multi-agent mode without a directly comparable AA cost/task point, so it is not plotted. OpenAI cut Luna 80% ($1.00/$6.00 → $0.20/$1.20) and Terra 20% ($2.50/$15.00 → $2.00/$12.00) on 2026-07-30; Sol is unchanged at $5.00/$30.00. Grok 4.5 is plotted without a max rung: xAI's docs list reasoning_effort as low/medium/high only, and AA 404s on that slug.

Sources (2026-08-03): AA v4.1 methodology · AA GPT-5.6 Sol · AA GPT-5.6 Terra · AA GPT-5.6 Luna · AA GPT-5.5 · AA Opus 5 · AA Opus 4.8 · AA Fable 5 · AA Sonnet 5 · AA Grok 4.5 · AA Gemini 3.1 Pro Preview · OpenAI GPT-5.6 · xAI reasoning guide · Chart.js 4.4.1 license

Cost axis is AA Intelligence-Index cost per task, not coding-agent cost. Hollow points are modeled scenarios, not published measurements. To update, edit AA_SNAPSHOT and keep the companion Markdown in sync.