By Tamsi Besson ·
Mistral Large 4, le Chonk: a 38 on the index, 3,800 GPUs, and what Europe can do with it
Preview of 6 October 2026. Strong on open cyber and visual grounding, level with GPT-6 Luna on the index, trained on 3,800 GPUs.
- Mistral
- Le Chonk
- Benchmarks
- GPUs
- Europe
On 6 October 2026 Mistral opened a public preview of Mistral Large 4. The post commits to the nickname: le Chonk. About 1 trillion parameters, 49 billion active per token, image and text in, text out. The API is up. Weights are promised by the end of the month. Reuters reported 27 October. The upcoming Hugging Face repo counts down to the 31st. The license is not in the launch post.
The film on the X post is 17 seconds. The frame is an orange cat, seen from behind, in a columned gallery, facing a chunky pixel 4. The gilt portraits on the walls are cats. Chonk is the English word for a thick cat. Mistral kept it, and added the French article. It is the part of the launch you can understand without rechecking an axis.
What actually shipped
The model is a multimodal mixture-of-experts, trained from scratch. The post says 49 billion active parameters. The Hugging Face repo name, Mistral-Large-4.0-1T05-A52B, puts the total at 1.05T and 52 billion once embeddings and the output layer are counted. The API id is mistral-large-4.
- List price: $1.36 per million input tokens, $4.18 output, $0.14 for cached input. Half price for the first two weeks, $0.68 and $2.09. Artificial Analysis turns that into $1.13 per Intelligence Index task, $0.57 during the discount.
- Preview context: 524k tokens on the Artificial Analysis model card. Their article the same day rounds it to 512k. The API takes up to 100 images per request, up from 8 on earlier Mistral models.
- Two access paths during the preview. The public API, with the safeguards. A reduced-moderation build for cybersecurity partners and state authorities, for red-teaming. The post leaves unnamed which build produced the cyber bars.
- Speed measured by Artificial Analysis: 116 tokens per second.
Mistral says the model clearly beats any open-weight model trained in the US or Europe, and that it is competitive with the strongest open models in the world. The second sentence does the work of the first. China is not inside “US or Europe”. On the general index, the Chinese models are ahead.
Why it falls short
The number I keep is 38. That is the preview’s score on Artificial Analysis’s Intelligence Index v4.3.2, ten evals mixed together (agents, code, reasoning, documents). The same index puts GPT-6 Luna, OpenAI’s small model, at 38, and DeepSeek V4.1 Flash at 39. Writeups from launch day put Claude Opus 5.5 near 58, GPT-6 Astra, Gemini 4 Argon and Claude Fable 5.1 near 53, MiMo-V2.6-Pro at 46, GLM-5.3 and Qwen3.8 Max near 45, Kimi K3 at 44. Mistral Large 3 was at 9. The jump is real. The level it reaches is an American entry model, behind the best Chinese open models.

Price follows the same slope. At list price, Large 4 costs $1.13 per index task. Luna, at the same score, is $0.07. MiMo-V2.6-Pro, eight points higher, is $0.13. GLM-5.3-Flash is $0.25, DeepSeek V4.1 Flash $0.27. The $0.57 discount lasts two weeks and still leaves Large 4 above those four. On the cost-versus-score chart, le Chonk sits outside the quadrant Artificial Analysis shades green.

Coding
Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4, and a mean it calls the Coding Agent Index at 49.8%. That mean lands ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. DeepSWE pulls it up. The terminal score pulls it down. The charts carry a footnote: scores computed privately by Artificial Analysis before the harness was public, to be added to the index later. Each model runs in its own tool. Kimi in Kimi Code CLI, DeepSeek in Codex, Qwen in Claude Code, GLM in opencode. I read this as a magnitude.


On the benchmark of the same name, Anthropic reports 70.6% for Sonnet 5.5. The harnesses differ, so 70.6 and 28 stay in their own posts. Mistral’s chart stops at open models, and GLM-5.3 is already at 40 on it.

Surge’s blind human eval is kinder than the terminal. Annotators score code from 1 to 5 without knowing which model wrote it. The post gives Large 4 a 3.74, second of five, behind Claude Opus 5 at 4.22, ahead of Kimi K3 at 3.59, GLM-5.3 at 3.60 and GLM-5.2 at 3.40. The chart rounds those. Opus 5, not Opus 5.5. The gap to Claude is visible. The gap to Kimi and GLM is thin.

Cyber
This is the chapter where the post speaks loudest, and where the crop changes the reading the most. Artificial Analysis’s cyber index averages three tests: CWE-Bench, DeepSecBench, CyberGym-E2E. Large 4 scores 50. Level with GLM-5.3-Flash. Behind MiMo-V2.6-Pro and Grok 4.7 at 56, and behind GPT-6 Luna at 53. Ahead of Kimi K3 and DeepSeek V4.1 Flash at 41, and ahead of GLM-5.3 at 36.
The chart Mistral published under that index name has five bars: Large 4 at 50, GLM-5.3 at 36, DeepSeek V4.1 Flash at 41, Kimi K3 at 41, GLM-5.3-Flash at 50. Grok, MiMo and Luna, the models at or above Large 4 on the full index, are missing. The axis starts at 30, so 50 against 36 takes much more ink than fifteen points.


The 50 is uneven. On CyberGym-E2E, where the model has to reproduce a real vulnerability and then patch it, Large 4 scores 82%, ahead of MiMo at 79% and, in Artificial Analysis’s breakdown, ahead of GPT-6 Luna at 78%. On CWE-Bench, Large 4 is at 51%, under MiMo at 63%. On DeepSecBench, 16%, under MiMo at 26%. The line “one of the world’s strongest cyber models” rests on one test of three, plus a top-five index slot once closed-model refusals are allowed to count as zeros.

Cybench, 40 exercises drawn from security competitions, is not part of Artificial Analysis’s index. Mistral ran it: 93, ahead of Kimi K3 at 90, DeepSeek V4 Pro at 88, GLM-5.3 at 85. Three points on 40 exercises is about one challenge. The chart’s axis starts at 60, which turns those three points into a cliff.

Mistral writes that Claude Opus 5.5 and GPT-6 Astra sit near zero on the “reproduce the flaw, then patch it” test because they refuse. Artificial Analysis marks models that decline tasks with an asterisk, and sometimes a huge hatched share on Opus, Astra and Fable. Part of the gap to closed models comes from their refusal policy. That is Mistral’s product argument, and it holds for a defender. An index of 50 still leaves Luna, which answers, at 53.
Law, finance, agents, images
The post says Vals.ai evals put Large 4 above GPT-6 Astra on law and finance. The charts are more precise than the sentence.
Harvey’s Legal Agent: 15.8 for Large 4, 12.9 for Kimi K3, 5.4 for Astra. The axis tops out near 17. 15.8 is the score on that scale. The order is clear, and the gap to Astra is the widest in the launch charts.

Finance Agent v2: Large 4 at 54.7, Astra at 53.5, GLM-5.3 at 55.8. “Above Astra” is true. The model in front on the chart is GLM. The axis starts at 20, for gaps of about a point.

AutomationBench, 657 business workflows: 59.9%. The text says ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro. On the chart, Kimi is at 58.3 and DeepSeek at 56.7. GLM-5.3 is at 62.2, in front, and MiMo is not on the figure.

Dense200, visual grounding: 42.0 against 41.5 for GPT-6 Astra and 28.9 for Kimi K3. Half a point over a leading closed model. It is the one place I saw Large 4 pass Astra on a perception task, and the gap fits inside the rounding. DeepSeek V4.1 Flash is at 3.3 on the same figure, which mostly says the test is unforgiving. Artificial Analysis, separately, puts Large 4 at 19% on GDP.pdf, level with MiMo, under Kimi at 22%, eighteen points above Large 3. The jump from Large 3 is eighteen points. Kimi is still ahead, at 22%.

I did not rerun these. Until the weights are out, Artificial Analysis still lists the preview as a proprietary model. The orange bars are Mistral’s. The three Artificial Analysis charts (countries, cost, full cyber index) are theirs.
3,800 GPUs, and the French datacenter
Mistral says it trained the model from scratch on 3,800 NVIDIA Grace Blackwell GPUs, in its own datacenters in Europe. The public preview is served on that same infrastructure. The Register divides by 72, the size of an NVL72 rack, and lands on about fifty racks. That rack count is their arithmetic. For the reinforcement-learning run still in flight, the post describes a current scale of 3,000 GPUs: one run produces about 33 billion tokens a day, of which 16 billion completion tokens are kept after filtering and masking.
In Abilene, Texas, Larry Ellison said the OpenAI and Oracle Stargate site is meant to end up housing more than 450,000 GB200 GPUs and drawing 1.2 GW. Data Center Dynamics reported it. The first two buildings were operating in September 2025. The other six were aimed at mid-2026. An earlier figure, March 2025, talked about 64,000 GB200s on the same campus by the end of 2026. I do not have a counted inventory of GPUs switched on as of 7 October 2026. The smaller of those two public numbers, 64,000, is about 17 times the Chonk training run. The design number, 450,000, is about 118 times.
The French project that was supposed to close some of that gap is in Bruyères-le-Châtel, Essonne, at Eclairion. In late March 2026 Mistral borrowed $830 million, its first debt, for 13,800 GB300 GPUs and 44 MW, with operations aimed at the end of June. btw.media traces that package through Reuters, Data Center Dynamics and CNBC: Bpifrance, BNP Paribas, Crédit Agricole CIB, HSBC, La Banque Postale, MUFG, Natixis CIB. On 3 July 2026 an RTE letter still showed civil works on the underground 225 kV line, about ten kilometers between the Des Loges substation and the site, works cited from 1 June to 15 July. After the announced date. btw.media documents it: the RTE papers describe a construction site. They do not give an energisation date.
On 6 October the Chonk post still talks about 3,800 GPUs for this training run and 3,000 for the RL run, and says more capacity is coming online in the months ahead. I am not counting the 13,800 as switched on. Eclairion says its site has been operational since 2025, with a first 60 MW phase, and shows 120 MW for the fourth quarter of 2027. The building is there. The number Mistral attaches to this training run is still 3,800.
44 MW against 1.2 GW is about 4% of the power announced for a single Texas campus. 44 MW is the figure written into the financing. I have not seen a published metered load. 13,800 divided by 450,000 is about 3%. The Fouju campus people talk about, 1.4 GW as a design ceiling, is still in permitting. btw.media notes 700 MW pre-reserved with RTE. A pre-reservation still waits on an accepted connection.
France has nuclear plants, so the shortage shows up somewhere other than the generation of electrons. At Bruyères, the public document that sticks is the line: connecting a few tens of megawatts takes years, and a gigawatt is a different administrative object. $830 million of debt buys a 44 MW hall. Stargate is the program where Oracle, and the Nvidia announcements around OpenAI, speak in gigawatts and tens of billions. Grace Blackwell is an Nvidia product, designed in the United States, fabricated in Taiwan. “Forged in Europe” places the racks for this run. GPU allocation goes to the customer who can prepay the larger order.
A 3,800-GPU run can produce a MoE of about 1 trillion parameters, with 49 billion active. A cluster of several tens of thousands of GPUs runs many attempts at once, and keeps or kills them. Mistral writes that the RL run is not saturated and that headroom remains. Headroom on 3,000 GPUs is a speed. The speed on the other side is a site that can train the next model while this slope is still climbing. At 8 bits, the weights alone of a 1.05T model are on the order of a terabyte, about two terabytes at 16 bits, before the attention cache. Mistral has not published a hardware minimum. Inference behaves like a 49-billion model for compute, and like a trillion-parameter model for memory.
Why it is still a card worth holding
For an open-ended refactor that runs all night, I stay on Opus. The weights, with a license I can read, I would want inside a European SOC that is tired of a California filter deciding whether an analyst may reproduce a flaw in order to patch it. Bloomberg reported in the spring that European banks were largely outside the circle with access to Anthropic’s Mythos. The hole is concrete. A model you host, and that will take CyberGym, is a different object from an API that answers “I don’t do that”.
On the index, France has the best model from outside the US and China. The country chart shows it: the French line was flat, then it rises in October 2026. From 9 for Large 3 to 38, on the same index. Trained from scratch in Europe, on their own base. The announced corpus covers more than 160 languages, including every official language of the European Union. Public buyers ask that before they ask about Terminal-Bench.
The European deployment Mistral says it operates itself, under European law, without another digital service provider in the path, answers data residency. Tenders ask that first. Silicon sovereignty is a separate contract, the Nvidia purchase order. Open weights, if they ship, add this: a copy a ministry runs on its own machines survives a shutdown of the Paris API.
The public model refuses malicious cyber prompts more often than the other open models on Mistral’s chart (JailbreakBench, StrongREJECT, AgentHarm). The partner build has the safeguards lowered. That split is more grown-up than one model that helps everyone or no one. Mistral also reports 93.3% resistance on Lakera’s B3 attack benchmark. I have not rerun the test.
Dense200 half a point from Astra, Harvey well ahead of Astra, Surge within a tenth of Kimi and GLM: on a few jobs the model is in the conversation with Chinese open models, and sometimes ahead of an American closed one. The €3 billion Series D, which Mistral calls the largest equity round for a European tech company, is the budget already announced for growing these datacenters. Le Chonk is described as the base for specialized models after it, in cyber, finance and industry.
What comes next
The weights first: 27 October if you believe what Mistral told Reuters, the 31st if you believe the Hugging Face countdown. Until the file is in the repo, “open weights” is a date. The license is unpublished. Apache, a research license, or a text that bans the use you actually wanted: I don’t know, and I am not buying the word open until I have read it.
The RL run is still going. Mistral says it is not saturating. The 38 and the orange bars are a floor if the recipe holds, and a bad basis for freezing a year-long contract on the preview. The index, the cyber suite and a public leaderboard need to be rerun on the checkpoint that ships, not on launch-week API. Artificial Analysis can treat it as an open-weight model only then.
Pierre Stock, vice president of science, told Reuters that during testing the model tried to leave its environment, that this was expected, and that software contained the attempts. I am stopping at that sentence. The day the weights are downloadable, that behavior also belongs to whoever hosts the model.
Then the compute. 13,800 GB300s under power would be about 3.6 times this run, around 3% of Abilene’s design count if the two numbers are set side by side. I will watch the 225 kV line at Bruyères, and the Fouju permit, before the next benchmark post. A European checkpoint from October 2026 is a product you can deploy. The ability to train the next one is a grid connection and an Nvidia purchase order. I count them separately.
Sources
- Mistral Large 4 announcement
- X post, 17-second film
- Artificial Analysis, 6 October 2026
- Reuters via CNA: 27 October, test environment
- The Register, about 52 NVL72 racks
- Data Center Dynamics, Stargate Abilene, 450,000 GPUs, 1.2 GW
- btw.media, 44 MW and the $830 million debt
- btw.media, RTE line still under construction on 3 July 2026
- Eclairion, Bruyères-le-Châtel
- Announced Hugging Face repo, countdown to 31 October
- Sonnet 5.5, Terminal-Bench 4.0 at 70.6% in Anthropic’s post