Le Chonk pushes Mistral to the trillion-parameter mark
Mistral Large 4, nicknamed Le Chonk, scales to one trillion parameters with 49 billion active. The natively multimodal model was trained on 3,800 NVIDIA Grace Blackwell GPUs in Europe, with an open-weight release scheduled for the end of the month.
One trillion parameters, 49 billion active
Nicknamed "Le Chonk," Mistral Large 4 represents a substantial increase in scale. The model contains one trillion parameters, with 49 billion active, and natively processes both text and visual inputs.
The initial version is available through the API in Mistral Studio. Its weights are not public yet, with Mistral planning an open-weight release at the end of the month. In the meantime, the company is testing the model with cybersecurity specialists, vetted partners and public authorities using a version with reduced moderation and expanded cyber capabilities.
Mistral describes ML4 as the highest-performing open-weight model developed in the US or Europe on its aggregated evaluations. The claim is based on the company's benchmark suite and does not mean ML4 leads every individual test. Trained on 3,800 Grace Blackwell GPUs in Europe
ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs located in Mistral's European data centers. The public preview is served from the same infrastructure.
A significant share of the training data spans more than 160 languages, including every official language of the European Union. Mistral also plans a European deployment operated directly by the company, independently of other digital service providers and under European law.
Once the weights are released, ML4 will also support private-cloud and on-premise deployment. Cybersecurity becomes a major benchmark
Cybersecurity accounts for a significant part of the published evaluations. On the Artificial Analysis Cyber Index, Mistral says ML4 ranks among the five strongest models globally.
On a task requiring models to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82%, the highest reported result on that test. It also solves 93% of the 40 Cybench challenges.
That comparison requires additional context. Mistral notes that some closed models score near zero on the vulnerability reproduction task because they refuse certain requests. The benchmark therefore captures both technical capability and the refusal policies applied to each model. Coding and professional agents
For software development tasks, ML4 reaches 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4. Its combined Coding Agent Index score reaches 49.8% in Mistral's results.
A blind human evaluation conducted with Surge AI gives the preview a coding quality score of 3.74 out of 5, behind Claude Opus 5 at 4.22 and ahead of the other models included in that comparison.
Agentic workflows are evaluated separately. ML4 scores 59.9% on AutomationBench, covering 657 business workflows across services including Gmail, Google Sheets, Slack and Salesforce. On AA-Briefcase, which evaluates long-horizon knowledge work involving documents, presentations and spreadsheets, the model reaches 1,393 Elo. Visual grounding against closed models
ML4's multimodal capabilities cover documents, charts, natural images, engineering drawings and geospatial imagery. The model can combine visual perception with agentic behavior to progressively inspect an image or locate specific elements across very large resolutions.
On Dense 200, Mistral reports 42% for ML4 versus 41% for GPT-6 Astra. This specific result supports the company's claim that ML4 can outperform closed frontier models on some visual grounding tasks. It is not a general comparison between the two models. Billions of RL tokens generated every day
Post-training combines supervised fine-tuning with reinforcement learning. At its current scale, Mistral runs RL across roughly 3,000 GPUs. The company says a single run generates around 33 billion tokens per day, including approximately 16 billion trainable completion tokens after filtering.
Tens of thousands of rollouts are generated in parallel, with long trajectories supporting budgets of millions of tokens. The RL run behind the current preview is still underway, meaning this version of ML4 is not presented as the final checkpoint.
Mistral Large 4 is available through the API now. Its weights, additional architecture details, more benchmarks and the post-training methodology are scheduled for release at the end of the month.