# Mistral Large 4: Inside “Le Chonk,” the 1.05T-Parameter Open-Weight Model

Mistral Large 4 explained: 1.05T parameters, 49B active, 1M context, pricing, benchmarks, API access, and the Oct. 27 open-weight release.

Canonical URL: https://aiidelist.com/blog/mistral-large-4

Language: en

Published: 2026-10-06

Updated: 2026-10-06



## Key Takeaways

- **Mistral Large 4, nicknamed “Le Chonk,” entered public preview on October 6, 2026.** It is Mistral AI’s new flagship general-purpose multimodal model.
- The model uses a **granular Mixture-of-Experts architecture with 1.05T total parameters, 49B active parameters, and a 1.6B vision encoder**.
- Its context window is **1 million tokens**, compared with 256K for Mistral Large 3.
- Mistral currently displays preview pricing of **$0.68 per million input tokens, $0.07 per million cached input tokens, and $2.09 per million output tokens**.
- Launch materials report **62% on DeepSWE v1.1, 67% on Finch/FinWorkBench, about 15% on Harvey Legal Agent, 73% on DIOR-RSVG, and 42% on Dense200**. These results are preliminary and should not be read as a final independent ranking.
- The API is available now, while **public model weights are planned for October 27, 2026** following a staged testing period.

![Image](https://cdn.aiidelist.com/api/image/Ihs1LJwZKRyngf35HBwV4.webp)

## What Is Mistral Large 4?

Mistral Large 4 is the successor to Mistral Large 3 and one of Mistral AI’s most ambitious releases to date. Mistral calls the model **ML4** and has embraced the community nickname **Le Chonk**, a reference to the earlier “Le Chaton Fat” meme about an imagined giant French AI model.

The important change is architectural. Large 4 is not a trillion-parameter dense transformer. It is a sparse **Mixture-of-Experts, or MoE, model**. The network contains 1.05 trillion parameters in total, but only about 49 billion are active during an inference path.

That design gives Mistral much more total model capacity without requiring every parameter to participate in every token calculation. It also explains why the headline “1T parameters” should not be interpreted as 1T parameters of active compute.

At the same time, **49B active does not mean Large 4 has the storage footprint of a 49B model**. Self-hosters still need to store and route across the full expert weight set.

![Image](https://cdn.aiidelist.com/api/image/UXbE7heGtboTDhDTv7whm.webp)

## Mistral Large 4 Specs

| Specification | Mistral Large 4 |
| --- | --- |
| Status | Public Preview |
| Preview date | October 6, 2026 |
| API model ID | `mistral-large-4` |
| Architecture | Granular Mixture-of-Experts |
| Total parameters | 1.05T |
| Active parameters | 49B |
| Vision encoder | 1.6B |
| Context window | 1M tokens |
| Input | Text + images |
| Output | Text |
| Structured outputs | Yes |
| Function calling | Yes |
| Document QnA | Yes |
| Batching | Yes |
| Agents / conversations | Yes |
| Built-in tools | Yes |
| Public weights | Planned for October 27, 2026 |

Mistral’s official model card confirms the core architecture, context size, pricing, API identifier, and agent-oriented features.

## Mistral Large 4 vs Mistral Large 3

Large 4 is a substantial upgrade rather than a minor refresh.

| Specification | Mistral Large 3 | Mistral Large 4 |
| --- | ---: | ---: |
| Total parameters | 675B | 1.05T |
| Active parameters | 41B | 49B |
| Context | 256K | 1M |
| Architecture | Granular MoE | Granular MoE |
| Multimodal | Yes | Yes |
| Input price | $0.50/M | $0.68/M current preview price |
| Output price | $1.50/M | $2.09/M current preview price |
| Weight status | Released | Planned Oct. 27 |

Large 3 is officially documented as a 675B-total, 41B-active model with a 256K context window and Apache 2.0 weights.

Large 4 therefore increases total parameter capacity by roughly **56%**, while active parameters rise by only about **20%**. The context window expands nearly fourfold.

That ratio is revealing. Mistral is scaling expert capacity much faster than per-token active parameters, suggesting that specialization and sparse routing are central to the Large 4 strategy.

## Why the 1.05T / 49B MoE Architecture Matters

A useful approximation is:

`49B / 1.05T ≈ 4.7%`

This does not equal an exact FLOP ratio, but it shows how sparse the model is.

The potential advantages are clear:

- more expert capacity without dense-model inference costs;
- better specialization across coding, finance, vision, cyber, and industrial workloads;
- a more attractive serving profile than an equally large dense model;
- room to increase total model knowledge without increasing active compute at the same rate.

The disadvantage is infrastructure complexity. Expert routing creates demanding requirements around GPU memory, memory bandwidth, interconnects, sharding, and expert parallelism.

A simple weight-storage estimate illustrates the challenge:

| Precision | Approximate raw storage for 1.05T parameters |
| --- | ---: |
| BF16 | ~2.1 TB |
| 8-bit | ~1.05 TB |
| 4-bit | ~525 GB |

These are rough calculations and exclude KV cache, the vision encoder, runtime buffers, quantization metadata, and distributed serving overhead.

So although Large 4 is positioned as open-weight, **it is not a normal single-GPU local model**. Practical self-hosting will likely require high-memory multi-GPU systems or distributed clusters unless future quantization work changes the deployment picture.

## The 1M Context Window Is a Major Upgrade

Large 4’s **1 million-token context window** may be more useful to many enterprise teams than the parameter count.

A 1M context can support:

- large software repositories;
- long collections of contracts;
- annual reports and supporting financial documents;
- extensive research dossiers;
- enterprise knowledge packs;
- multimodal document collections;
- long-running agent trajectories.

However, a million-token maximum does not mean applications should always send a million tokens. Long prompts increase latency and cost, and maximum context length does not guarantee perfect recall of information buried deep in the prompt.

Good implementations should still retrieve relevant data, remove duplication, cache stable context, and test long-context recall under realistic workloads.

## Mistral Large 4 Pricing

Mistral currently displays the following prices for the preview:

| Token type | Price |
| --- | ---: |
| Input | $0.68 / 1M tokens |
| Cached input | $0.07 / 1M tokens |
| Output | $2.09 / 1M tokens |

Mistral AI 资料

The official page also displays crossed-out higher figures of $1.36 input, $0.14 cached input, and $4.18 output. Mistral has not clearly explained on the model card whether the lower values are temporary launch rates, so long-term cost models should be rechecked before production deployment.

Cached input is particularly important for 1M-context applications. An agent repeatedly working over the same repository, policy library, or knowledge base may reduce costs substantially if stable context can be reused.

For example:

```text
20M uncached input × $0.68 = $13.60
50M cached input   × $0.07 =  $3.50
 5M output         × $2.09 = $10.45

Estimated total = $27.55
```

## Mistral Large 4 Benchmarks

Mistral is emphasizing coding, finance, legal agents, visual grounding, cybersecurity, and industrial work.

The headline launch scores are:

| Benchmark | Area | Mistral Large 4 Preview |
| --- | --- | ---: |
| DeepSWE v1.1 | Agentic coding | 62% |
| Finch / FinWorkBench | Finance workflows | 67% |
| Harvey Legal Agent | Legal agents | ~15% |
| DIOR-RSVG | Remote-sensing grounding | 73% |
| Dense200 | Dense visual grounding | 42% |

Venturebeat

### Coding: strong, but benchmark configuration matters

Mistral reports 62% on DeepSWE v1.1, compared in its launch material with 61% for GLM-5.3, 57% for DeepSeek V4 Pro 0813, 51% for Qwen 3.8 Max, and 44% for Reflection Beam.

That does not prove Large 4 is the best coding model overall. Agentic coding benchmarks are highly sensitive to the harness, tools, retry strategy, context management, and inference settings. VentureBeat noted that stronger published configurations exist for some competitors than the specific numbers shown in Mistral’s chart.

The safer conclusion is that Large 4 is **competitive with leading open-weight coding systems**, but independent standardized evaluations are still needed.

### Finance: a notable 67% result

Mistral reports 67% on Finch/FinWorkBench, tied with DeepSeek V4 Pro in its launch comparison and ahead of GLM-5.3 at 65%.

This benchmark is useful because it focuses on multi-step finance workflows involving documents, spreadsheets, search, modeling, and reporting rather than simple knowledge questions.

### Legal: strong relative ranking, weak absolute reliability

Large 4 is reported at roughly 15% task pass rate on the Harvey Legal Agent benchmark.

That may be competitive, but **15% is still a low absolute success rate**. It supports using the model inside supervised legal workflows, not treating it as an autonomous legal professional.

### Vision grounding may be the hidden strength

Mistral reports 73% on DIOR-RSVG and 42% on Dense200.

These benchmarks matter because they test locating and reasoning about objects and regions in images. That maps naturally to:

- satellite and aerial imagery;
- industrial inspection;
- technical diagrams;
- manufacturing;
- geospatial analysis;
- CAD-like documents;
- semiconductor workflows.

This helps explain why Mistral is positioning Large 4 as an enterprise and industrial model rather than merely another chatbot.

## Why Cybersecurity Is Central to the Launch

Cybersecurity is one of Mistral’s most strategic Large 4 narratives.

The argument is that defensive security teams may need capabilities that closed API providers can restrict or modify through policy changes. Open weights give organizations more control over deployment, moderation, access, and data handling.

Reuters reports that Mistral is working with cybersecurity experts and government authorities during the staged period before the October 27 weight release.

However, the public evidence is not yet complete. Broad cybersecurity claims were part of the launch, but a comprehensive independently reproducible cyber benchmark package was not yet available. Buyers should wait for the final checkpoint, red-team details, and independent testing before treating cyber leadership as established.

## Trained in Europe on About 4,000 Blackwell GPUs

Mistral says Large 4 was trained from scratch over roughly **two months on about 4,000 Nvidia Grace Blackwell GPUs** in its European data centers.

This is central to the company’s sovereign-AI positioning.

Mistral is not only selling model quality. It is selling the idea that governments and enterprises can use a frontier-class model trained and served from European infrastructure and later deploy the weights on infrastructure they control.

For regulated organizations, jurisdiction, operational continuity, customization, and provider independence can matter almost as much as benchmark scores.

## How to Use Mistral Large 4 Through the API

The official API model identifier is:

`mistral-large-4`

A typical Python integration looks like this:

```python
import os
from mistralai import Mistral

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])

response = client.chat.complete(
    model="mistral-large-4",
    messages=[
        {
            "role": "user",
            "content": "Explain sparse MoE inference in practical terms."
        }
    ]
)

print(response.choices[0].message.content)
```

The official model card also lists structured outputs, function calling, Document QnA, batching, agents and conversations, and built-in tools.

Because the model is still marked **Public Preview**, teams should test output consistency, function-calling reliability, long-context latency, caching behavior, multimodal limits, and preview-version changes before committing critical production workloads.

## Can You Download Mistral Large 4?

Not yet at the beginning of the preview.

The API became available on **October 6, 2026**, while Mistral plans to release the public model weights on **October 27, 2026**.

That distinction matters because search results may describe Large 4 as open-weight even though the downloadable checkpoint has not yet been released.

As of the preview launch:

- API access: **available**;
- public preview: **available**;
- downloadable public weights: **planned for October 27**;
- final repository details: **pending**;
- final weight license: **verify at release**.

Do not automatically assume Large 4 uses the same Apache 2.0 license as Large 3. Large 3 is explicitly labeled Apache 2.0, while the final Large 4 weight terms should be checked when the actual checkpoint is published.

## What to Watch When the Weights Arrive

The October 27 release should answer several questions that the API preview cannot:

- exact checkpoint and repository;
- file size and precision;
- final license terms;
- official quantization formats;
- vLLM and SGLang support;
- expert-parallel recommendations;
- practical GPU memory requirements;
- tokenizer and chat-template details;
- final post-RL benchmarks;
- independent coding and intelligence evaluations;
- multimodal serving requirements;
- cybersecurity evaluation details.

Mistral has said reinforcement learning is still ongoing during the preview period, so the final public checkpoint may differ from the model represented by the initial October 6 benchmark results.

## Conclusion

Mistral Large 4 is a significant return to the frontier-model race for Mistral AI. Its combination of **1.05T total parameters, 49B active parameters, a 1.6B vision encoder, 1M context, native multimodality, agent features, and low preview API pricing** gives it a distinctive position among open-weight models.

The early benchmark story is encouraging, especially in coding, finance, and visual grounding, but the evidence is still preliminary. Vendor-reported scores, harness differences, continuing reinforcement learning, and unreleased weights make definitive rankings premature.

The most important milestone is therefore **October 27, 2026**. If the released weights reproduce the preview’s performance while offering workable licensing and deployment economics, Large 4 could become one of the most important enterprise open-weight models of the year.

Developers can begin evaluating the `mistral-large-4` API now, but self-hosting and long-term production decisions should be revisited once the final weights, license, runtime requirements, and independent benchmarks are public.
