Mistral's Le Chonk model reproduces security exploits that Claude and GPT refuse to touch
Mistral AI released Mistral Large 4, internally nicknamed Le Chonk, on October 6, and the headline number is not the parameter count. On one benchmark that asks a model to reproduce a real software vulnerability and then patch it, ML4 scored 82%, the highest of any model tested, while Claude Opus 5.5 and GPT-6 Astra scored near zero because they refused to attempt the task at all, per Mistral's announcement.
What
ML4 is a 1-trillion-parameter natively multimodal model with 49 billion active parameters, trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral's own European data centers, according to Mistral. A public preview API is live today on Mistral Studio; full weights are due at the end of October, after a red-teaming period in which cybersecurity partners and government authorities get expanded access with reduced moderation, per the same post. On Cybench, a 40-challenge security competition set, the model scores 93%, and it ranks among the top five models globally on the Artificial Analysis Cyber Index, ahead of every open-weight model built outside China, per Mistral's post. VP of Science Pierre Stock told TechCrunch the model was trained on a fraction of the compute rival labs use, calling it "two to three times less" than Chinese competitors and "significantly less" than closed-source rivals.
Mistral turns rivals' safety refusals into a selling point
The framing inverts the usual safety story: Mistral argues that a closed model's refusal to reproduce a known exploit is itself a weakness for defenders, who often need to prove a flaw is real before they can patch it. That matters for any security team evaluating which model to put in an incident-response or vulnerability-research pipeline right now, since Claude and GPT's caution on this specific task is not a hypothetical, it is a measured benchmark gap. ML4 itself is pitched as a sovereign, self-hostable alternative for that exact workflow, pairing open weights with a claim of top-tier cyber performance. Government access to a less-moderated version during red-teaming has not been detailed publicly, and the full weights are not yet available for independent teams to verify any of these numbers themselves.
What to watch next
Whether Mistral ships the promised weights on schedule at the end of October, and whether independent benchmarks from Artificial Analysis or LMSYS corroborate Mistral's self-reported cybersecurity and coding scores once outside researchers can run the model directly. Also watch for detail on which government authorities get the expanded, less-moderated access during the red-teaming window, and under what terms.
Sources
- Introducing Mistral Large 4: Mistral blog, October 6, 2026
- Mistral's new 1T model aims to leapfrog closed and open rivals: TechCrunch, Anna Heim, October 6, 2026
- The next hot open AI model is named Le Chonk: The Verge, Robert Hart, October 6, 2026
