Loading
Loading
Intelligence has a price, and the price is convex. This page divides each frontier model's Artificial Analysis (AA) Intelligence Index score by its blended API price, weighting input 80% and output 20%. The last 3 index points cost 1.7× the money: Claude Opus 5 against Kimi K3 on blended price. The chart below is the same claim as the table: up and to the left is better, and the amber line is the Pareto frontier.
AA Intelligence Index and first-party list prices, retrieved 9 Aug 2026. Scores are mode-specific and measured at the effort level named in the table. Latency is a third axis this chart does not have: both OpenAI and Anthropic now sell a Fast tier at twice the standard rate.
Unverified. Our frontier is the Pareto set of the 17 models we hold, drawn from the 185 Artificial Analysis evaluates, and Artificial Analysis's own published line has not been pulled into the registry yet. Until it is, read the 5 names on the frontier above as optimal within our sample rather than optimal outright: a model we do not track could dominate any of them. Artificial Analysis's scatter is the comparison set.
Cost bases differ: cost per Intelligence Index task (not blended $/Mtok — orderings agree, units do not). Membership is compared, not coordinates.
Cost is not the only thing a frontier can be drawn against, and it is not always the binding one. Below, the horizontal axis is throughput once generation starts; further right is faster. The frontier here has almost no members in common with the one above, which is the finding: paying more does not buy speed, and the models that top the index are among the slowest on the page.
AA Intelligence Index and first-party list prices, retrieved 9 Aug 2026. Scores are mode-specific and measured at the effort level named in the table. Latency is a third axis this chart does not have: both OpenAI and Anthropic now sell a Fast tier at twice the standard rate.
Time to first token, log scale, lower is better. For a reasoning model this is mostly think time rather than network latency, which is why it spreads across three orders of magnitude while price spreads across two. Read it as the axis that decides whether a model can sit behind an interactive product at all.
AA Intelligence Index and first-party list prices, retrieved 9 Aug 2026. Scores are mode-specific and measured at the effort level named in the table. Latency is a third axis this chart does not have: both OpenAI and Anthropic now sell a Fast tier at twice the standard rate.
Holding every model Artificial Analysis scores would not make this chart better: a frontier is decided by the handful of rows along its edge, and the rest only add ink. What costs us is a missing lab, because that is the one error a frontier drawn over a subsample can make. So the labs worth tracking are listed rather than assumed, and the unpriced ones are named.
Named flagships are leads for the next pull, not data. Nothing enters the table above without a score and a price from the model's own source page.
Click any column to sort.
| DeepSeek V4-Flashopen weights scored at max · 0731 refresh moved this row; prior score predates the AA rebase | 52 | 0.14 | 0.28 | 0.17 | 309.5 | 132 | 1.15 |
| GPT-5.6 Luna scored at max · repriced 30 Jul, was $1.00 / $6.00 | 52 | 0.20 | 1.20 | 0.40 | 130.0 | 202 | 131.40 |
| MiniMax M3open weights | 45 | 0.30 | 1.20 | 0.48 | 93.8 | 112 | 1.37 |
| Muse Spark 1.2 scored at xhigh · standard tier; contributor tier $0.10 / $0.20 in exchange for training rights | 57 | 1.25 | 4.25 | 1.85 | 30.8 | n/r | n/r |
| Muse Spark 1.1 scored at xhigh · deprecated in favour of 1.2, which is 4 points up at the same list price | 53 | 1.25 | 4.25 | 1.85 | 28.6 | 234 | 2.36 |
| GLM-5.2open weights scored at max · repriced since 1 Aug, was $1.40 / $4.40 | 53 | 1.35 | 4.29 | 1.94 | 27.3 | 126 | 1.41 |
| Qwen3.8 Max weights trailed for 10 Aug 2026, not released at time of pull; rank #9 of 185 | 58 | 2.00 | 6.00 | 2.80 | 20.7 | 91 | 3.09 |
| Grok 4.5 scored at high · under-200K tier; a prompt >=200K reprices the whole request to $4.00 / $12.00 | 56 | 2.00 | 6.00 | 2.80 | 20.0 | 61 | 7.99 |
| Gemini 3.5 Flash scored at high · rank #25 of 185 | 52 | 1.50 | 9.00 | 3.00 | 17.3 | 189 | 53.43 |
| GPT-5.6 Terra scored at max · repriced 30 Jul, was $2.50 / $15.00; rank #13 of 185 | 57 | 2.00 | 12.00 | 4.00 | 14.3 | 147 | 167.74 |
| Qwen3.7 Max superseded by 3.8 Max; rank #34 of 185 | 47 | 2.50 | 7.50 | 3.50 | 13.4 | 210 | 2.29 |
| Kimi K3 scored at max · weights promised 27 Jul; top of AA open-weight cohort | 60 | 3.00 | 15.00 | 5.40 | 11.1 | 43 | 2.35 |
| Claude Opus 5 scored at max · rank #1 of 185 | 63 | 5.00 | 25.00 | 9.00 | 7.0 | 58 | 30.61 |
| Claude Opus 4.8 scored at max · deprecated in favour of Opus 5 | 57 | 5.00 | 25.00 | 9.00 | 6.3 | 60 | 15.81 |
| GPT-5.6 Sol scored at max · rank #5 of 185 | 61 | 5.00 | 30.00 | 10.00 | 6.1 | 70 | 122.00 |
| GPT-5.5 scored at xhigh · deprecated in favour of GPT-5.6 Sol | 56 | 5.00 | 30.00 | 10.00 | 5.6 | 82 | 106.73 |
| Claude Fable 5 scored at max · scored with Opus 4.8 fallback; rank #3 of 185 | 62 | 10.00 | 50.00 | 18.00 | 3.4 | 70 | 124.10 |
Blended cost per 1M tokens = 0.8 × input price + 0.2 × output price, the mix of a retrieval-heavy workload that reads far more than it writes. Scores are the AA Intelligence Index at the reasoning effort shown against each model, where AA names one; prices are first-party API list prices. The ratio rewards cheapness, so compare models within a capability tier rather than across tiers · a model that fails the task is expensive at any price.
DeepSeek V4-Flash, GLM-5.2 and MiniMax M3 are open-weight models; the rest are closed.
Source: artificialanalysis.ai per-model pages, retrieved 9 Aug 2026. Prices and scores move; the retrieval date is the claim.