Models
Artificial Analysis Reworks Index After Astra Scoring Doubts
Version 4.2 of the Intelligence Index adds two benchmarks, drops a saturated one and raises private test data to 40 percent of the weighting. Astra gains four points; Claude Fable 5.1 keeps the lead.

When a benchmark disagrees with everyone else, which one is wrong? Last week Artificial Analysis scored GPT-6 Astra roughly level with its predecessor, GPT-5.6 Sol, while Epoch AI ranked Astra first out of 267 models and ARC-AGI-3 showed a large jump. On September 5 Artificial Analysis shipped version 4.2 of its Intelligence Index, and Astra now sits four points ahead of Sol. The provider does not say the update was a response to the criticism, but the timing speaks for itself.
What changed in the index
Two benchmarks were added: AA-Briefcase, a private evaluation of real-world knowledge work, and GDP.pdf from Surge AI, which tests PDF document analysis. GPQA-Diamond was dropped because current models have effectively solved it. Private test data now accounts for 40 percent of the weighting, which makes training on the test set harder. The provider also says it corrected scoring errors on several benchmarks and adjusted grading for more stable results.
The leaderboard after the update: Anthropic's Claude Fable 5.1 still first, GPT-6 Astra second, Meta third. Astra uses fewer tokens per task than any other frontier model, according to the same data. On cost against performance, Anthropic, OpenAI, Meta and Zhipu AI share the lead.
Why the provider moved now
Artificial Analysis says it normally freezes methodology around major launches to keep scores comparable. This time the top of the table moved fast enough that an interim update was, in its words, necessary. A full version 5 has been in development for eight months and will roll out in stages.
That explanation is plausible and also convenient. The uncomfortable fact is that a widely cited independent index and OpenAI's own numbers pointed in different directions for a week, and the index changed rather than the model. Whether that is a correction or a capitulation depends on details the provider has not published, such as how much of Astra's four-point gain comes from the new benchmarks versus the scoring fixes.
Reading leaderboards after this
Three habits are worth keeping. Check the version number of any index you quote, because v4.1 and v4.2 scores are not comparable. Weight private-data benchmarks more heavily than public ones, since the whole point of the 40 percent shift is that public tests leak into training data. And treat the ranking itself as the least useful number; for a business decision, the token-efficiency and cost-to-performance figures in the same release say more about what a model will cost you to run than a single composite score does.
Sources: The Decoder, Artificial Analysis

Written by
Muhammet Fatih Batman
Founder & Editor
Founder of YZ Uzman, with 20+ years of experience in web design and software development.