Frontier models: intelligence vs cost per task

Lines connect measured and explicitly modeled effort points. Filled dots are measured by Artificial Analysis; hollow dots are formula-derived. Snapshot: (all 33 measured rows refreshed together across Google, xAI, OpenAI, and Anthropic).

Index v4.1.1 refresh. AA's current methodology still combines nine evaluations, but the v4.1.1 coordinates differ from the previous v4.1 chart. Every measured point was refreshed together so Grok 4.6 is compared on one consistent basis. Compare costs and scores within this snapshot only.

measured (Artificial Analysis) formula-derived effort level
ModelEffortIndex$/taskPointSource / basis

Configurations. AA's Fable point is specifically Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback). OpenAI says max gives GPT-5.6 more reasoning time than xhigh, while ultra coordinates four agents in parallel by default. Ultra is a separate multi-agent mode without a directly comparable AA cost/task point, so it is not plotted. The app ladder differs from the API ladder plotted here: the ChatGPT/Codex picker shows Light · Medium · High · Extra High · Ultra for Sol and Terra, but only Light · Medium · High · Extra High for Luna (no Ultra) — observed 2026-08-04. The names do not correspond to the API's lowmax, the app offers no "Max", and whether app "Ultra" maps to API max is undocumented. Every rung plotted here is an API configuration measured by AA. OpenAI cut Luna 80% ($1.00/$6.00 → $0.20/$1.20) and Terra 20% ($2.50/$15.00 → $2.00/$12.00) on 2026-07-30; Sol is unchanged at $5.00/$30.00. Grok 4.6 supports low, medium, high, and xhigh reasoning effort, with high as the default. AA currently measures only high, so Grok's other three rungs are visibly hollow M1 projections. xAI does not document a max rung for Grok 4.6, so none is plotted.

Sources (2026-08-13): AA Gemini 3.6 Flash · AA Gemini 3.5 Flash · AA Gemini 3.5 Flash-Lite · AA v4.1 methodology · AA GPT-5.6 Sol · AA GPT-5.6 Terra · AA GPT-5.6 Luna · AA GPT-5.5 · AA Opus 5 · AA Opus 4.8 · AA Fable 5 · AA Sonnet 5 · AA Grok 4.6 · AA Gemini 3.1 Pro Preview · OpenAI GPT-5.6 · xAI Grok 4.6 docs · Chart.js 4.4.1 license

Cost axis is AA Intelligence-Index cost per task, not coding-agent cost. Hollow points are modeled scenarios, not published measurements. To update, edit AA_SNAPSHOT and keep the companion Markdown in sync.