Lines connect measured and explicitly modeled effort points. Filled dots are measured by Artificial Analysis; hollow dots are formula-derived. Snapshot: (Gemini Flash rows captured that day; all other rows 2026-08-03, verified same cost basis across all four vendors).
Cost re-baselined 2026-08-03. AA revised its cost-per-task basis since the 2026-07-29 snapshot — Intelligence Index scores came back identical while $/task rose 12.5–24.3% (Anthropic 12.5–14.5%, Gemini 16.5%, Grok 16.8%, OpenAI 18.8–24.3%), with per-token list prices unchanged. Compare costs within this snapshot only.
| Model | Effort | Index | $/task | Point | Source / basis |
|---|
Configurations. AA's Fable point is specifically Claude Fable 5
(Adaptive Reasoning, Max Effort, Opus 4.8 Fallback). OpenAI says max gives GPT-5.6 more
reasoning time than xhigh, while ultra coordinates four agents in parallel by default.
Ultra is a separate multi-agent mode without a directly comparable AA cost/task point, so it is not plotted.
The app ladder differs from the API ladder plotted here: the ChatGPT/Codex picker shows
Light · Medium · High · Extra High · Ultra for Sol and Terra, but only
Light · Medium · High · Extra High for Luna (no Ultra) — observed 2026-08-04. The names do not
correspond to the API's low…max, the app offers no "Max", and whether app "Ultra" maps to
API max is undocumented. Every rung plotted here is an API configuration measured by AA.
OpenAI cut Luna 80% ($1.00/$6.00 → $0.20/$1.20) and Terra 20%
($2.50/$15.00 → $2.00/$12.00) on 2026-07-30; Sol is unchanged at $5.00/$30.00.
Grok 4.5 is plotted without a max rung: xAI's docs list
reasoning_effort as low/medium/high only, and AA 404s on that slug.
Cost axis is AA Intelligence-Index cost per task, not coding-agent cost. Hollow points are modeled scenarios, not published measurements. To update, edit AA_SNAPSHOT and keep the companion Markdown in sync.