Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism
| Source: THE DECODER
Tags: Artificial Analysis, GPT-6 Astra, Claude Fable, benchmarks, leaderboard, model evaluation
Artificial Analysis revised its Intelligence Index to v4.2 after criticism that its scoring underrepresented GPT-6 Astra's capabilities — Astra now sits 4 points above predecessor Sol, but Claude Fable 5.1 still leads. GPQA-Diamond retired as frontier models saturated it; private test data now 40% of weighting to curb benchmark gaming.
Details
Artificial Analysis published v4.2 of its Intelligence Index, a methodology revision prompted by skepticism from the evaluation community. Earlier independent evaluations had placed GPT-6 Astra well ahead of other models — Epoch AI ranked it first among 267 models with 169 points across 50+ benchmarks, and ARC-AGI-3 showed a large jump — but Artificial Analysis had scored Astra on par with its predecessor Sol. The updated index corrects this, giving Astra a four-point gain over Sol and placing it second overall. Anthropic's Claude Fable 5.1 holds the top spot in v4.2, with Meta in third. On a separate efficiency metric, GPT-6 Astra uses fewer tokens per task than any other frontier model — relevant for teams running high-volume production workloads. On cost-to-performance ratio, Anthropic, OpenAI, Meta, and Zhipu AI all cluster at the top. The revision adds two new benchmarks: AA-Briefcase (real-world knowledge work) and GDP.pdf from Surge AI (PDF document analysis). GPQA-Diamond was dropped because models have effectively saturated it. Private test data now makes up 40% of weighting, a direct effort to make benchmark gaming harder. Artificial Analysis says Version 5 of the index has been in development for eight months and will roll out in stages. The scramble to revise Astra's scores mid-cycle reveals how quickly benchmark credibility can erode when independent evaluations diverge sharply from one another.