Frontier models: intelligence vs cost per task

Lines connect measured and explicitly modeled effort points. Filled dots are measured by Artificial Analysis; hollow dots are formula-derived. Snapshot: .

measured (Artificial Analysis) formula-derived effort level
ModelEffortIndex$/taskPointSource / basis

Configurations. AA's Fable point is specifically Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback). OpenAI says max gives GPT-5.6 more reasoning time than xhigh, while ultra coordinates four agents in parallel by default. Ultra is a separate multi-agent mode without a directly comparable AA cost/task point, so it is not plotted.

Sources (2026-07-18): AA v4.1 methodology · AA GPT-5.6 Sol medium · AA Opus 4.8 · AA Fable 5 · AA Sonnet 5 · AA Grok 4.5 · AA Gemini 3.1 Pro Preview · OpenAI GPT-5.6 · Chart.js 4.4.1 license

Cost axis is AA Intelligence-Index cost per task, not coding-agent cost. Hollow points are modeled scenarios, not published measurements. To update, edit AA_SNAPSHOT and keep the companion Markdown in sync.