Mistral AI launched the public preview of Mistral Large 4 (ML4, nicknamed le Chonk) today. The headline says "1 trillion parameters," but the right place to start is a different number.

Total:  1.05T parameters
Active: 49B parameters per token
Vision: 1.6B encoder
Context: 1M tokens

Each token does not touch all one trillion parameters. A granular MoE activates roughly 49B of expert paths per token. (Official docs also show 52B active including embeddings and output layers. Read 49B as the routed per-token scale.)

The weights are not public yet and are due by the end of October. What you can use now is the preview API on Mistral Studio. Treat "open-weight" as a promise, not the current state.

The specs, collected in one place

From the announcement and docs:

Architecture  granular MoE, natively multimodal
Total         1.05T parameters
Active        49B per token
Vision        1.6B vision encoder
Context       1M tokens
Inputs        multimodal (text + images/documents)
Features      function calling, Agents & Conversations,
              built-in tools, structured outputs
Training      European datacenters, 3,800 Grace Blackwell GPUs,
              from scratch over ~2 months
Languages     160+, covering every official EU language

Preview pricing is discounted:

input         $0.68 / M tokens
cached input  $0.07 / M tokens
output        $2.09 / M tokens

The list prices ($1.36 / $0.14 / $4.18) are struck through, so expect them to change at general availability.

Why 49B matters more than 1T

This is exactly the structure our MoE active-parameter guide describes:

1.05T → total capacity, the weight footprint to store
49B   → the compute path that actually runs per token

The trillion number says how many patterns the model can store across experts. The 49B number decides per-token cost and latency. The reason Mistral chose this shape is straightforward: coding and agent workloads repeat dozens or hundreds of turns of reasoning and tool calls. Without capping per-token compute, agent workloads become unaffordable.

Read the benchmarks as specialist workloads

Mistral emphasizes job-shaped results over generic chatbot scores:

Coding      DeepSWE v1.1 61.7%, Terminal-Bench 4.0 28.3%,
            Coding Agent Index 49.8%
            (ahead of DeepSeek V4 Pro, Qwen3.8 Max — vendor figures)
Human eval  Surge blind coding quality 3.7/5, 2nd place
            (after Claude Opus 5 at 4.2; ahead of Kimi K3, GLM-5.3 at 3.6)
Agents      AutomationBench 59.9% (657 business flows across
            Gmail, Sheets, Slack, Salesforce), AA-Briefcase 1,393 Elo
Security    Top open-weight AA Cyber Index ranks, Cybench 93%,
            82% on reproduce-then-patch (best on that task)
Legal/fin   Above GPT-6-Astra on Harvey Legal Agent and
            Finance Agent v2 (third-party evals by vals.ai)
Vision      Dense200 grounding 42% (vs 41% for GPT-6-Astra)

The security point has a practical edge. Proving a flaw is real through reproduction is where defense starts, and strongly refusing closed models can block exactly that step. Mistral answers with the open-weight argument: ship the capability, let the customer control deployment location and policy. That is why European sovereign deployment and on-premise runs feature so prominently.

The charts make the picture concrete. On DeepSWE 1.1, ML4 Preview (orange, 62) leads Beam's self-reported figure (44), Qwen3.8 Max (51), DeepSeek V4 Pro (57), and GLM-5.3 (61), with Kimi K3 (68) on top.

Artificial Analysis DeepSWE 1.1 bar chart. Mistral Large 4 Preview 62, Beam 44, Qwen3.8 Max 51, DeepSeek V4 Pro 57, GLM-5.3 61, Kimi K3 68

Source: Mistral AI announcement "Introducing Mistral Large 4" (Oct 6, 2026). Measured by Artificial Analysis under differing harness conditions. Beam's figure is labeled self-reported in the original.

The security chart shows the same shape. On the AA Cyber Index, ML4 Preview and GLM-5.3-Flash jointly lead at 50, ahead of Kimi K3 and DeepSeek V4.1-Flash at 41 and GLM-5.3 at 36.

Artificial Analysis Cyber Index bar chart. Mistral Large 4 Preview 50, GLM-5.3 36, DeepSeek V4.1-Flash 41, Kimi K3 41, GLM-5.3-Flash 50

Source: Mistral AI announcement "Introducing Mistral Large 4" (Oct 6, 2026). Artificial Analysis Cyber Index.

Human judges point the same way. In the Surge blind eval (1–5 scale), ML4 Preview scores 3.7, second only to Claude Opus 5 (4.2) and ahead of Kimi K3 and GLM-5.3 (3.6) and GLM-5.2 (3.4).

Surge human evaluation bar chart. Mistral Large 4 Preview 3.7, GLM-5.2 3.4, Kimi K3 3.6, GLM-5.3 3.6, Opus 5 4.2

Source: Mistral AI announcement "Introducing Mistral Large 4" (Oct 6, 2026). Surge AI blind human evaluation.

These are still vendor-reported numbers. Read them as direction until the weights land and independent reproductions run.

CodeBridge mini lab: measure five things on one workload

Turn the briefing's action into an experiment. Skip scoreboard comparisons and compare task success on your own workload instead.

1. Fix one workload (e.g. fixing 10 repo issues)
2. Run it with the same harness, tools, and time limits
3. Swap only the model and record these five:

   [ ] task success rate (test-pass criterion)
   [ ] wall-clock latency (minutes per task)
   [ ] token usage (input / output / cache hits)
   [ ] tool-call success rate (valid calls share)
   [ ] cost with 1M context (include 3+ long inputs)
4. Compute cost per success (success-normalized active compute)

Once that table is filled, the "is it a 1T model" question disappears. One question remains:

How little active compute solves the same agent task?

Conclusion: judge by work per kilowatt, not by totals

Mistral Large 4 points in a clear direction:

Old question: who holds the biggest model?
New question: who solves the same agent task
              with the least active compute?

1.05T of capacity worked at 49B at a time, specialist workloads like coding, security, finance, and manufacturing inside one model, sovereign serving from European infrastructure. Less a "big model" than a way of working sold as a model.

One short, sharp next step: run just five of your repo issues through the preview API and write down tokens, time, and cost per success. That single row beats ten benchmark tables. When the weights land at the end of October, add local serving cost on the same harness and you will see the true open-weight price.

Further reading

References

Go deeper with a course

To practice picking models and wiring evaluation metrics into agent loops and graph structures, train with guided, hands-on examples just like this article's mini lab.