Lines connect measured and explicitly modeled effort points. Filled dots are measured by Artificial Analysis; hollow dots are formula-derived. Snapshot: (all 33 measured rows refreshed together across Google, xAI, OpenAI, and Anthropic).
Index v4.1.1 refresh. AA's current methodology still combines nine evaluations, but the v4.1.1 coordinates differ from the previous v4.1 chart. Every measured point was refreshed together so Grok 4.6 is compared on one consistent basis. Compare costs and scores within this snapshot only.
| Model | Effort | Index | $/task | Point | Source / basis |
|---|
Configurations. AA's Fable point is specifically Claude Fable 5
(Adaptive Reasoning, Max Effort, Opus 4.8 Fallback). OpenAI says max gives GPT-5.6 more
reasoning time than xhigh, while ultra coordinates four agents in parallel by default.
Ultra is a separate multi-agent mode without a directly comparable AA cost/task point, so it is not plotted.
The app ladder differs from the API ladder plotted here: the ChatGPT/Codex picker shows
Light · Medium · High · Extra High · Ultra for Sol and Terra, but only
Light · Medium · High · Extra High for Luna (no Ultra) — observed 2026-08-04. The names do not
correspond to the API's low…max, the app offers no "Max", and whether app "Ultra" maps to
API max is undocumented. Every rung plotted here is an API configuration measured by AA.
OpenAI cut Luna 80% ($1.00/$6.00 → $0.20/$1.20) and Terra 20%
($2.50/$15.00 → $2.00/$12.00) on 2026-07-30; Sol is unchanged at $5.00/$30.00.
Grok 4.6 supports low, medium, high, and
xhigh reasoning effort, with high as the default. AA currently measures
only high, so Grok's other three rungs are visibly hollow M1 projections. xAI does not
document a max rung for Grok 4.6, so none is plotted.
Cost axis is AA Intelligence-Index cost per task, not coding-agent cost. Hollow points are modeled scenarios, not published measurements. To update, edit AA_SNAPSHOT and keep the companion Markdown in sync.