Mistral Large 4 Preview: 1T MoE, Weights Due in October
Mistral Large 4 enters public preview as a ~1T-parameter multimodal MoE with 1M context. Weights are promised for October; scores are vendor-run.
Mistral Large 4 enters public preview as a ~1T-parameter multimodal MoE with 1M context. Weights are promised for October; scores are vendor-run.
Introduction
On October 6, 2026, Mistral AI launched a public preview of Mistral Large 4, which the company calls "le Chonk" and describes as "unofficially ML4." The preview is available through the Mistral Studio API today. Mistral says it will release the weights by the end of October, so the model is not open-weight yet; for now it is an API-only preview.
Mistral's blog describes a natively multimodal model that is the largest and most capable the company has built. Parameter counts differ slightly between Mistral's own pages. The blog rounds the model to 1 trillion parameters with 49 billion active, while the model page lists 1.05 trillion total, 52 billion active, plus a 1.6-billion-parameter vision encoder. Every benchmark figure below is reported by Mistral, including scores on third-party benchmarks, and none has been independently reproduced yet.
Feature Overview
Architecture and scale. The docs page describes a general-purpose multimodal model with a granular Mixture-of-Experts (MoE) design. In an MoE model only a fraction of the parameters run for each token, which keeps per-token compute far below what the total parameter count suggests. The context window is 1 million tokens. Mistral says architecture details, more benchmarks and the post-training methodology will arrive together with the weights.
Training infrastructure. According to Mistral, the model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters, and the preview is served from the same infrastructure. Mistral says the European deployment is operated end to end by the company under European law, and that training data covers more than 160 languages, including every official EU language. Pierre Stock, Mistral's VP of Science, told TechCrunch the model was trained on "only 4,000 Nvidia GPUs," which he described as two to three times fewer than Chinese competitors use.
Cybersecurity. Mistral reports that on the Artificial Analysis Cyber Index the model "ranks among the top five models globally" and leads open-weight models developed outside China. On one test in that index, reproducing a real vulnerability in open-source software and then patching it, Mistral reports 82%, which it calls the highest of any model. It also reports 93% on Cybench, a set of 40 security-competition exercises. Mistral claims that Claude Opus 5.5 and GPT-6 Astra score near zero on that test because they refuse the task. That is Mistral's claim, and the refusal behavior of those models was not verified for this article.
Coding and agents. Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4, and a Coding Agent Index of 49.8%, which it says is ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. In a blind human coding evaluation run with Surge AI on a 1-5 scale, Mistral says the preview placed second of five at 3.74, behind only Claude Opus 5 at 4.22 and ahead of GLM-5.3 (3.60) and Kimi K3 (3.59). On AutomationBench, which covers 657 business workflows across Gmail, Google Sheets, Slack and Salesforce, it reports 59.9%.
Other claims. Mistral reports 42% versus 41% for GPT-6-Astra on the Dense 200 visual-grounding test, and says the model exceeds GPT-6-Astra on vals.ai legal and finance tests. For safety it reports that the model resists 93.3% of attacks on the Lakera B3 AI Security Benchmark.
Usability Analysis
Pricing is published on the Mistral docs page, per million tokens: $0.68 for input, $0.07 for cached input and $2.09 for output. The page shows each of these next to a struck-through higher figure ($1.36, $0.14 and $4.18). It does not explain the strike-through or say how long the lower price lasts, so teams budgeting beyond the preview should confirm it with Mistral. Supported endpoints include chat completions, conversations, agents, function calling, structured outputs, document question answering, batching and built-in tools.
The stated optimization targets are cybersecurity, finance and chip design, according to Stock. Chip design matters for Mistral's backers: TechCrunch notes that ASML led the Series C and Samsung led the Series D. For European organizations that need an EU-operated model, the preview offers an API path now. Self-hosting is a separate matter. A model of roughly one trillion parameters needs substantial multi-accelerator hardware, and the weights are not yet available.
| Item | Mistral blog | Mistral docs page |
|---|---|---|
| Total parameters | 1 trillion | 1.05 trillion |
| Active parameters | 49 billion | 52 billion |
| Vision encoder | Not stated | 1.6 billion |
| Context window | Not stated | 1M tokens |
Pros and Cons
The strengths are a long context window, published per-token pricing, European operation and the promise of open weights. The reported cybersecurity and agentic-coding scores are strong on paper. The limits are mostly about verification. Scores come from the vendor, weights are not out, and key technical details are deferred. Mistral also plans reduced-moderation access for vetted partners during red-teaming, which is a policy choice outsiders cannot audit yet.
Outlook
Mistral says the model is the first milestone of a roadmap funded by its 3 billion euro Series D and the base for a new generation of specialized models. It also says it trains with the same environment it sells as Mistral Forge. On reinforcement learning, Mistral says a run at current scale (about 3,000 GPUs) produces roughly 33 billion tokens per day, and that the RL run is still in flight with "no signs of saturation." Stock told TechCrunch the weights should arrive in about three weeks, after safety testing, and that Mistral will work with trusted partners and governments so the weights can be used to defend rather than attack. Whether that release plan works in practice will be visible only after the weights ship.
Conclusion
Mistral Large 4 is a credible large-scale entry from a European lab, with usable pricing and a clear release timeline. The benchmark story is Mistral's own for now. Developers who want an EU-operated API can test the preview today. Those waiting for open weights, independent evaluations or architecture details should hold their judgment until the end-of-October release and third-party testing arrive.
Editor's Verdict
Mistral Large 4 Preview: 1T MoE, Weights Due in October earns a solid recommendation within the Other LLM space.
The strongest case for paying attention: a 1M-token context window and multimodal input cover long documents and mixed-media workflows. That alone raises the bar for what readers should expect in this space. Reinforcing that, pricing is published on the docs page, so teams can estimate costs during the preview — practical value rather than just headline appeal. The broader signal worth registering is straightforward: the gap between 1T/49B on the blog and 1.05T/52B on the docs page shows how rounded marketing figures can differ from the model card. On the other side of the ledger, one constraint is real rather than a marketing footnote: the weights are not yet released, so today the model is reachable only through Mistral's API. It should factor into any serious decision. Layered on top of that, benchmarks are reported by the vendor and have not been independently reproduced — which narrows the set of teams for whom this is an obvious yes.
For multi-model deployment teams, cost-conscious operators, and developers willing to evaluate beyond the major labs, this is a serious evaluation candidate, not just a curiosity to bookmark. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.
Pros
- A 1M-token context window and multimodal input cover long documents and mixed-media workflows.
- Pricing is published on the docs page, so teams can estimate costs during the preview.
- European operation under European law is a clear fit for customers with data-residency requirements.
- The promised open weights would let organizations inspect and self-host the model after release.
- Reported agentic-coding and cyber results are competitive on paper against several named rivals.
Cons
- The weights are not yet released, so today the model is reachable only through Mistral's API.
- Benchmarks are reported by the vendor and have not been independently reproduced.
- Reduced-moderation access for vetted partners is a policy choice that outsiders cannot audit yet.
- Architecture details and post-training methodology are deferred until the weights ship, and self-hosting a trillion-parameter model will need serious hardware.
References
Comments0
Key Features
1. Public preview via Mistral Studio API, launched October 6, 2026; weights promised by end of October 2. Natively multimodal granular MoE: about 1T total / 49B active (blog) or 1.05T / 52B plus 1.6B vision encoder (docs) 3. 1M-token context window; $0.68 input / $0.07 cached / $2.09 output per million tokens (docs page) 4. Trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's European datacenters 5. Vendor-reported results in cybersecurity (82% on one Cyber Index test, 93% Cybench) and agentic coding (61.7% DeepSWE v1.1)
Key Insights
- The gap between 1T/49B on the blog and 1.05T/52B on the docs page shows how rounded marketing figures can differ from the model card.
- An API-only preview ahead of the weights lets Mistral collect usage data and red-team results before opening the model.
- Cybersecurity is the main differentiator Mistral chose, and its refusal-behavior comparison with rivals is a vendor claim that needs outside testing.
- Training on about 3,800 GPUs in European datacenters supports a sovereignty pitch for EU customers and governments.
- Struck-through list prices on the docs page show a lower current price, but the page does not say why or for how long.
- A blind human coding evaluation that places the model second of five is more informative than self-selected benchmarks, but it was commissioned and reported by Mistral.
- The shared training and RL environment with Mistral Forge links this release to the company's enterprise customization business.
Was this review helpful?
Share
Related AI Reviews
Reflection Beam: 501B Open-Weight MoE, Weights Due Soon
Reflection previews Beam, a 501B MoE with 23B active parameters. Apache 2.0 weights are promised later this month; benchmarks are unverified.
Meta Muse Gadget SDK: Build Your Own ESP32 and Pi Devices
Meta open-sourced Apache 2.0 ESP32 and Linux SDKs for DIY Muse gadgets, with candid security caveats and a Home Link giveaway of 5,000 units.
Meta Expands Muse Agent With Video Chat, Glasses, Charm
Meta expanded its Muse AI agent at Connect 2026 with video chat, smart glasses, Mac control, and a Muse Charm wearable shipping in December.
Grok 4.7 Review: Same Price, Trails Fable 5.1 and GPT-6
xAI's Grok 4.7 keeps Grok 4.6's $2/$6 pricing on a larger base model, but Artificial Analysis scores it 46 versus 53 for Fable 5.1 and GPT-6.
