Methodology · companion to the model chart

How the model chart is built

Where the numbers come from, what counts as measured versus estimated, and a running log of what has changed each time the chart was updated.

Updated September 5, 2026

This is the working notes behind the chart. It exists because a chart with no methodology is just a picture, and because when a number moves it is worth being able to say why it moved.

Where the numbers come from

Every plotted point is from Artificial Analysis, an independent benchmarking outfit. Two numbers per point:

  • Intelligence score is their Intelligence Index v4.1.1, a weighted average of nine separate evaluations.
  • Cost per task is what it cost them to actually run that benchmark on that model, at that effort level.

That second one matters more than it sounds. It is not the token price from the vendor's pricing page. A model that thinks longer burns more tokens to answer the same question, so two models at identical token prices can sit in very different places on the cost axis. The chart plots what the work costs, not what the menu says.

Measured versus estimated

Filled dots are measured. Artificial Analysis ran that exact configuration and published the result.

Hollow dots are estimated, and they are only ever drawn for effort levels a vendor genuinely offers but that nobody has benchmarked yet. Grok 4.6 is the current example: xAI documents four reasoning efforts, Artificial Analysis has published one. The other three are projected by taking the shape of GPT-5.6 Sol's measured effort curve and applying it from the known point.

That is an assumption, clearly marked as one. The rule is that the chart never invents a rung that does not exist. When Grok 4.5 appeared to have a max setting, checking xAI's own docs showed it did not, so the point was removed rather than relabelled.

The Fast line is a price, not a benchmark

The orange dashed line is GPT-5.6 Luna on OpenAI's Fast service tier. Fast is a delivery speed you pay extra for, not a smarter model, so every point keeps its measured intelligence score and moves to exactly twice its standard cost. It is a pricing scenario drawn onto measured intelligence, not a second benchmark run.

The one grey dot

Gemini 3.1 Flash-Lite is older than everything else on the chart. It stays because it is the model the AI Health Export iOS app runs today, which makes it a useful reference point for a real decision rather than an abstract one.

Two honest limits

Do not compare costs between snapshots. Artificial Analysis periodically revises how it calculates cost per task. When that happens every model's cost moves at once while the intelligence scores stay identical. Comparing today's dollar figure against one from a previous update can show a "price change" that never happened. Costs are comparable within a single snapshot only.

A benchmark is not your work. The Intelligence Index measures general capability. It does not measure whether a model writes in your voice, holds up under your load, or handles the specific thing you actually do. It is a starting point for a shortlist, not a verdict.

Most of these prices are introductory

Nearly every model on this chart launched with a pricing incentive. Sol's current rates are promotional. Terra and Luna both arrived the same way, and Gemini 3.7 Flash's introductory pricing is scheduled to double at the end of the year. Fable 5.1 is different: its $10 input and $50 output rates did not fall, but cache reads are 75% cheaper than Fable 5. The savings therefore grow with how much reusable context a workflow reads from cache.

The practical consequence is that the cheap window is the moment to test a model, not to wait it out. That is the whole argument for keeping systems portable between models: if switching is a config change rather than a rewrite, a pricing promotion becomes free capacity to get real work done.

What has changed

Newest first. Each entry is a full re-probe of every plotted model unless noted.

September 5, 2026: Astra added, existing comparison preserved

Added only Astra’s low, medium, high, xhigh and max thinking levels. All existing model scores, costs and levels remain exactly as recorded on September 2, including Gemini 3.1 Flash-Lite and the full Luna Fast pricing line.

Astra’s full-precision measurements come from Artificial Analysis’s page archived September 4 at 08:48 UTC, on Intelligence Index v4.1.1. The current AA pages use v4.2 and are not the source for this addition. This is an additive comparison with stated collection dates, not a fresh measurement of every model. Ultra has no comparable measured score/cost pair, so it has no plotted point.

September 2, 2026 — Fable 5.1 raises the ceiling

Claude Fable 5.1 launched on September 1 with five measured Artificial Analysis effort levels: low, medium, high, xhigh, and max. Max reaches 65.653 on the Intelligence Index, 2.60 points above Opus 5 max, but costs 57.9% more per task. Fable 5.1 high is the more economical comparison: it nearly ties Opus 5 xhigh while costing 20.6% less per task.

Anthropic's “up to approximately 45% cheaper” statement is not a 45% cut to every request. Input and output remain $10 and $50 per million tokens. Cache reads fell from $1.00 to $0.25, so Anthropic estimates about 25% lower cost for typical workloads and up to about 45% for highly agentic, cache-heavy work.

All 31 frontier rows plus the AI Health Export comparison point were re-probed on September 2. Existing intelligence scores and every non-Sol task cost were unchanged. Sol's task costs fell another 5.2% to 7.6% with no new OpenAI price change, so that movement is recorded as another benchmark-side cost recalculation rather than a second vendor cut.

August 26, 2026 — Sol got cheaper, and something else moved too

OpenAI cut GPT-5.6 Sol on August 21, from $5 to $4 per million input tokens and $30 to $20 per million output. Promotional, stated as available at least through November 21.

All 26 models were re-probed. Every intelligence score came back identical, so only the cost axis moved. Sol's five points fell between 15.7% and 18.3%.

Three other things moved with no price change announced at all:

  • GPT-5.6 Terra, up 2.4% to 4.3%
  • GPT-5.6 Luna, up 2.1% to 3.5%
  • Grok 4.6, up 12.0%, from $0.8367 to $0.9372 per task

Unchanged scores with moved costs and no vendor announcement is the signature of a benchmark-side recalculation, not a price rise. It is called out here rather than quietly absorbed, because a reader comparing against the previous snapshot would otherwise reasonably conclude xAI raised prices 12%.

That also means Sol's drop is not purely the price cut. A price-only change would have landed between 0.667× and 0.800× of the old figure; all five Sol points came in between 0.817× and 0.844×. Roughly 2% to 5% of the movement is the same recalculation, sitting on top of the genuine cut.

Two non-price corrections in the same pass, from OpenAI's current model documentation superseding the launch announcement: the GPT-5.6 context window is 1,050,000 tokens with a 128,000-token maximum output, and a long-context tier does exist. Prompts above 272,000 input tokens price the whole request at double input and 1.5× output. An earlier version of this page said there was no such tier, which was wrong.

August 14, 2026 — Gemini 3.7 Flash, and seven lines removed

Gemini 3.7 Flash launched August 13 and is plotted as a fully measured three-point line. Its introductory price is $0.75 per million input and $3.75 per million output through the end of 2026, after which Google says both double.

Seven series were removed to keep the chart readable. Each one had a model still on the chart that was both smarter and cheaper at every comparable point, so nothing was lost by dropping it:

RemovedSuperseded by
Claude Opus 4.8Claude Opus 5 medium
Claude Fable 5Claude Opus 5 xhigh
GPT-5.5GPT-5.6 Terra max
Gemini 3.1 Pro PreviewGemini 3.7 Flash low
Gemini 3.6 FlashGemini 3.7 Flash medium
Gemini 3.5 FlashGemini 3.7 Flash medium
Gemini 3.5 Flash-LiteGPT-5.6 Luna medium

August 13, 2026 — Grok 4.6, and a new index version

Grok 4.6 replaced Grok 4.5. xAI documents four reasoning efforts (low, medium, high, xhigh) with high as the default; Artificial Analysis publishes one measured point, Grok 4.6 high. The other three are drawn as estimates.

Artificial Analysis also moved to Intelligence Index v4.1.1, which changed coordinates across the board, so every measured row was re-pulled on the same day to keep all four vendors on one comparable basis.

The old charts

Every previous version is still here, and still interactive. These are the actual files, not screenshots, so you can hover the lines and read the numbers exactly as they stood on the day.

They are worth a look mostly because the shape of the field changed so fast. The chart started with seven models, grew to thirteen as new ones launched, then dropped back to nine once it became clear several were simply beaten on both axes at once.

SnapshotModelsWhat it shows
Aug 26, 20269Sol's promotional price cut and a full cost-basis refresh
Aug 14, 20269Gemini 3.7 Flash arrives; seven superseded lines removed
Aug 13, 202613Grok 4.6 replaces 4.5; everything re-measured on Index v4.1.1
Aug 4, 202613The widest it ever got, with three Gemini Flash models plotted
Aug 3, 202610GPT-5.6 Terra and Luna added after their launch pricing landed
Jul 29, 20268Claude Opus 5 added on launch day
Jul 18, 20267The first version

Two things to keep in mind when reading them. Do not compare costs between snapshots, for the reason above: the benchmark's cost basis was revised more than once in this window, so a dollar figure from July and one from August are not measuring quite the same thing. And older files carry older labels, because they are records rather than restatements. The grey comparison dot, for instance, is tagged as an internal shorthand in the August files and by its model name in the current one.

The current chart is on the main page.


The chart and this page are maintained by hand from the sources linked above. If a number here looks wrong, it may well be. Tell me and I will re-check it.