38 versus 58. The first number is Mistral Large 4's score on Artificial Analysis' Intelligence Index; the second is Claude Opus 5.5's. That 20-point gap is the context for everything Mistral said on October 6 when it released Large 4, nicknamed "le Chonk", as a public preview and called it the strongest open-weight model outside China "by a substantial margin."

The model

Large 4 is a mixture-of-experts model with about 1 trillion total parameters and 52 billion active per request, according to Mistral's announcement (the model card cited by third parties says 1.05 trillion and 49 billion). It accepts text and images, was trained on data covering more than 160 languages including every official EU language, and was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral's own European data centres. It is served from that same European infrastructure. The context window is not stated in the announcement; third-party listings disagree between 524K and 1M tokens, so treat it as unconfirmed.

Pricing is $1.36 per million input tokens and $4.18 per million output. Access is through the API on Mistral Studio only; the announcement does not mention Le Chat, Azure or Bedrock. Open weights are promised "by the end of the month." No licence has been named yet.

Mistral's numbers

The pitch leads with cybersecurity. Mistral reports 82% on a reproduce-and-patch vulnerability test, which it calls the highest of any model, and 93% of Cybench tasks solved. The announcement says Claude Opus 5.5 and GPT-6 Astra "score near zero" on the patch test, and gives the reason: those models refuse the task. The comparison measures policy as much as capability.

Coding is weaker. DeepSWE v1.1 comes in at 61.7%, Terminal-Bench 4 at 28.3%. In Surge AI's blind human evaluation of coding output, Large 4 Preview scores 3.74 out of 5, second to Claude Opus 5 at 4.22. Even in Mistral's own charts, Kimi K3, GLM-5.3 and DeepSeek sit ahead on coding. Elsewhere Mistral cites AutomationBench 59.9%, an AA-Briefcase Elo of 1,393 and a 93.3% resistance score on the Lakera B3 attack benchmark.

Independent numbers

Artificial Analysis puts Large 4 Preview at 38 on its Intelligence Index: just under GPT-6 Luna and DeepSeek V4.1 Flash at 39, and well below Claude Opus 5.5 (58), Claude Sonnet 5.5 (56) and GPT-6 Astra (53). On the same provider's Cyber Index it scores 50, level with GLM-5.3-Flash and behind MiMo-V2.6-Pro at 56.

One more figure matters for anyone budgeting: running the Intelligence Index took the model about 200 million tokens against a median of 81 million. A verbose model turns a low list price into a higher cost per task. By one estimate that is roughly $1.13 per index task, several times what DeepSeek V4.1 Flash costs.

The fairest framing comes from Simon Willison: Mistral Large 3 scored 9 on the same index in December 2025, so 38 is a very large jump for the company. Critics point to the distance between the marketing line and a score that puts Large 4 level with, not ahead of, the leading Chinese open models, and to the dual-use nature of a model sold on finding and exploiting vulnerabilities.

Where the model fits

For a business, the interesting line is not the benchmark but the infrastructure. A model trained and served in Europe under European law, soon runnable on your own servers, answers the objection we hear most often from regulated customers: the data must not leave our jurisdiction. On that criterion Large 4 competes in a much smaller field than the index suggests.

Our advice is to scope it accordingly. Large 4 is not a drop-in replacement for Opus 5.5 or Astra on general reasoning. It is a candidate for a specific job: internal document search, security analysis, or classification where the data cannot leave your perimeter. Test the API preview on your own workload now, and compute cost from tokens consumed per task rather than list price. When the weights and licence arrive at the end of the month, you will have real measurements for the self-hosting decision.

Sources: Mistral, Large 4 announcement, Simon Willison, le Chonk, Slashdot