Frontier models: intelligence vs cost per task

Lines connect measured and explicitly modeled effort points. Filled dots are measured by Artificial Analysis; hollow dots are formula-derived. Snapshot: (26 frontier rows re-probed, plus one deployed-product comparator).

Gemini 3.7 Flash added and the chart pruned. Gemini's low, medium, and high points are all measured. Seven older or analytically redundant series were removed; each had a retained current point that was both smarter and cheaper. Compare costs and scores within this v4.1.1 snapshot only.

AI Health Export lens. Gemini 3.1 Flash-Lite remains as a single gray diamond because v2 currently uses it. The orange hollow-square line shows the entire Luna effort curve with Fast pricing. Fast is a per-request service tier, not a medium-only model, so every Luna point keeps its measured intelligence score and moves to exactly 2× its Standard cost. Luna medium therefore moves from $0.0113 to $0.0225/task. This is a pricing scenario, not a separate AA latency benchmark. Google Cloud billing and bulk prompt-test spend do not feed this axis.

measured (Artificial Analysis) formula-derived effort level official-rate pricing scenario
ModelEffortIndex$/taskPointSource / basis

Configurations. OpenAI says max gives GPT-5.6 more reasoning time than xhigh, while ultra coordinates four agents in parallel by default. Ultra is a separate multi-agent mode without a directly comparable AA cost/task point, so it is not plotted. The app ladder differs from the API ladder plotted here: the ChatGPT/Codex picker shows Light · Medium · High · Extra High · Ultra for Sol and Terra, but only Light · Medium · High · Extra High for Luna (no Ultra) — observed 2026-08-04. The names do not correspond to the API's lowmax, the app offers no "Max", and whether app "Ultra" maps to API max is undocumented. Every rung plotted here is an API configuration measured by AA. OpenAI cut Luna 80% ($1.00/$6.00 → $0.20/$1.20) and Terra 20% ($2.50/$15.00 → $2.00/$12.00) on 2026-07-30; Sol is unchanged at $5.00/$30.00. Gemini 3.7 Flash is measured at low, medium, and high. Grok 4.6 supports low, medium, high, and xhigh reasoning effort, with high as the default. AA currently measures only high, so Grok's other three rungs are visibly hollow M1 projections. xAI does not document a max rung for Grok 4.6, so none is plotted.

Sources (2026-08-14): AA Gemini 3.7 Flash low · medium · high · AA Gemini 3.1 Flash-Lite Preview · OpenAI Fast mode pricing · AA v4.1.1 methodology · AA GPT-5.6 Sol · AA GPT-5.6 Terra · AA GPT-5.6 Luna · AA Opus 5 · AA Sonnet 5 · AA Grok 4.6 · Google Gemini 3.7 Flash announcement · OpenAI GPT-5.6 · xAI Grok 4.6 docs · Chart.js 4.4.1 license

Cost axis is AA Intelligence-Index cost per task, not coding-agent cost. Hollow points are modeled scenarios, not published measurements. To update, edit AA_SNAPSHOT and keep the companion Markdown in sync.