Home / News / Mistral's 'Le Chonk' Packs 1 Trillion Parameters, With Open Weights Due October 27
Mistral

Mistral's 'Le Chonk' Packs 1 Trillion Parameters, With Open Weights Due October 27

Oct 7, 20265 min read
Mistral's 'Le Chonk' Packs 1 Trillion Parameters, With Open Weights Due October 27

News Summary

Mistral AI, the Paris-based AI lab, opened a public preview of Mistral Large 4 (ML4), nicknamed "Le Chonk," on October 6, 2026 (Central European Time, as reported by outlets covering the launch). The company describes the roughly one-trillion-parameter model as the strongest open-weight system developed outside China, and says its full weights will be published on October 27, 2026 after a testing period of about three weeks. Wired's original report could not be retrieved for this write-up, so the details below are cross-checked against VentureBeat, Euronews, Tech Insider, and other coverage. Most performance figures are Mistral's own claims and have not been independently reproduced.

What Mistral Announced

Le Chonk is the company's newest flagship. Mistral says it ranks among the top open-weight models in the world on aggregate benchmarks and leads all open-weight models trained outside China by a substantial margin. Co-founder and chief scientist Guillaume Lample said ML4 "is at the frontier of open weight models."

The preview is available through Mistral's API and developer studio. Early access has been aimed at developers, cybersecurity professionals, and public-sector users. The weights are planned for release on October 27, 2026, though some outlets describe the window loosely as late October (October 27 to 31). Mistral says the delay allows red-teaming with security experts and government agencies. Reinforcement-learning training was reported to still be wrapping up during the preview.

Architecture and Specifications

ML4 uses a sparse mixture-of-experts design, an approach Chief Executive Arthur Mensch earlier summarized as "fat but sparse." Instead of running every parameter for each token, the model routes each token to a small subset of specialized expert networks. That keeps inference cost closer to that of a much smaller model.

Reported specifications, which vary slightly by outlet:

  • Total parameters: about 1 trillion (one outlet reports 1.05 trillion).
  • Active parameters per token: about 49 billion.
  • Inputs and outputs: multimodal input (text and images) with text-only output; one outlet cites a 1.6-billion-parameter vision encoder.
  • Context window: 1 million tokens, according to Tech Insider; not confirmed in every report.
  • Languages: more than 160, including all official EU languages and non-Latin scripts.
  • License: a custom Mistral license, per VentureBeat; full terms had not been detailed in other coverage.

Training Setup

Mistral says it trained the model from scratch in about two months on roughly 4,000 NVIDIA Grace Blackwell GPUs, in European data centers that it owns. Reports put the range at 3,800 to 4,000 GPUs. The company stresses that the whole pipeline ran in Europe, which it presents as a key part of its sovereignty pitch to governments and regulated industries.

Benchmark Claims

Mistral highlighted results in software engineering, cybersecurity, finance, manufacturing, and visual grounding. Figures reported by the outlets:

  • DeepSWE v1.1 (coding): 62%.
  • FinWorkBench (finance): 67%.
  • Harvey Legal Agent: 15% task-pass rate.
  • Dense200 (visual grounding): 42%; DIOR-RSVG (satellite and aerial grounding): 73%.
  • Cybench: 93%; CyberGym-E2E: 82%.

Mistral also said several closed frontier models score near zero on these cybersecurity tasks because of safety filtering rather than a lack of ability. This is a claim about refusal behavior, not about underlying capability, and readers should treat it as a vendor comparison until third parties test it. Mistral further reports top results on semiconductor-design benchmarks and says it is narrowing the gap with leading closed models on coding.

Pricing and Access

Tech Insider reports preview API pricing of $1.36 per million input tokens and $4.18 per million output tokens. Other outlets did not disclose pricing, so this should be verified on Mistral's official page. Once the weights are public, organizations can run the model on their own infrastructure, which matters for teams that need data to stay in-house.

Company Context

The launch follows a €3 billion Series D round in September 2026 at a valuation above €21 billion, reported as the largest equity round in European tech. Mistral says it serves more than 125 enterprise customers, including Airbus, ASML, and HSBC, and that its science team has grown from about three people to roughly 300. Earlier in 2026, annual recurring revenue was reported above $400 million.

Why It Matters for Learners and Builders

For students and engineers, ML4 is a useful case study in three ideas: sparse mixture-of-experts models, which decouple total size from per-token compute; open-weight releases, where the trained parameters can be downloaded and studied even if the training data and code are not fully open; and staged release practices, where a preview and safety testing precede publication. Open weights also let researchers fine-tune, evaluate, and audit the model independently.

What to Watch Next

The key milestone is the October 27 weight release. Independent evaluations should show whether the benchmark claims hold up, how the custom license treats commercial use, and what hardware is needed to run a trillion-parameter model locally. Final pricing and the context-window details should also be confirmed once Mistral publishes full documentation.

MistralOpen-Weight