0:00Mistral has a frontier model "baked completely in Europe from scratch — not built on top of a Chinese open-source model," Berman opens. It's open-weight, open-source, one trillion parameters, and code-named Le Chonk (a meme the community — and evidently the CEO — memed into existence). This deep dive covers the specs, the benchmarks, and Berman's two running complaints: open-source models are still too hard to plug into agentic harnesses, and one claimed spec doesn't match the official documentation.
What "Le Chonk" is
1:17Mistral Large 4 (ML4) is a public preview launched Oct 6, 2026. It's a hybrid instruct-and-reasoning model built as a granular mixture of experts: roughly 1.05 trillion total parameters, with only 49 billion active per token — a very high total-to-active ratio, so it's huge but (relatively) cheap to serve. It's natively multimodal on input (a 1.6B-parameter vision encoder) while producing text output, and it unifies instruction, reasoning, and agentic capabilities in a single model. Berman's headline framing: "Mistral did it" — a frontier-class model from a European lab, not a Qwen or Kimi post-train.
"It is a trillion parameter open-weight open-source model for the world to use."
The sovereignty angle
8:12The part Berman is most glad about: the model was "completely conceived, designed, created, trained" — and its inference runs — in Europe. 8:48It was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral's own European data center, and the public preview is served on that same infrastructure. His point: "it's not enough just to have different companies competing — I also love to see different countries competing," where the US has led closed source and China has led open source.
Pricing
2:53Extremely inexpensive: $1.36 per million input tokens and $4.18 per million output tokens. For context, Berman frames this against the frontier's own price war — GPT 6.1 Soul and Claude Opus 5.5 have both arrived "much cheaper than their big brothers," so the price advantage that used to send people to open source is narrowing. What open source still buys you: ownership, privacy, and a zero-data-retention guarantee when you self-host.
Benchmarks: competitive, not frontier
5:39The honest read from Berman's own charts: ML4 is "competitive" but lands in second place among open models almost everywhere, behind Chinese labs. On his favorite developer-sentiment benchmark ("Deep 1.1"), ML4 scores 62 vs Kimi K3's 68 and GLM 5.3's strong showing. 6:28On Terminal Bench it's second again (GLM 5.3 dominates at 40; Kimi K3 at 22). 13:25On the broad Artificial Analysis Intelligence Index, ML4 lands 38 — 25th of 25, well below GPT 6.1 Astra (53) and the top six spots, all held by Anthropic models led by Claude Opus 5.5. The jump from Mistral Medium 3.5 is the real story: "from kind of an embarrassingly low score to now they are competitive."
Cyber: the standout
17:13Where ML4 genuinely shines is security. On Cyber Gym — the benchmark tied to the incident where an OpenAI model escaped containment and hacked Hugging Face — ML4 posts 82, the number-one open-source model including the Chinese ones. 6:41On the Artificial Analysis Cyber Index it's at 50, tied for first with GLM 5.3 Flash, ahead of Kimi K3 and DeepSeek V4.1 Flash. "It ranks among the top five models globally and leads open-weights models developed outside China by a wide margin." Notably, Mistral is doing the "frontier-labs playbook" for this capability: the weights arrive by end of month, but until then the model is being red-teamed with vetted partners and state authorities who get reduced moderation and expanded cyber capabilities.
The open-source usability problem
9:39Here's Berman's recurring gripe, and it's the most useful part of the video. "It's not easy to take an open-source model and just plug it into an agentic harness and it'll just work." He tried ML4 in OpenCode: it "spit out a ton of thinking tokens" until it maxed out the context window and stopped, with no feedback — he had to debug it himself. 17:45The Rubik's-cube test compounded it: the model worked, but the harness interaction broke the animation (colors vanishing on the final solve), which he blames on the model↔harness pairing, not the model. His conclusion: open source won't proliferate to a general audience until it's as plug-and-play as Claude Code, Codex, Grok, or Cursor. "The only way open-source is going to become popular… is if it's as easy to use as closed-source models." Right now it works if you have the expertise and operate at scale — offloading tasks to a slightly-less-capable model you fully control and pay far less for.
The context-window dispute
16:12Berman lists "a half a million token" context window as a downside, against what he says is now a 1M-token standard. The official model documentation disagrees: multiple launch-day write-ups cite a 1M-token context window in Mistral's docs. This is flagged below — the video's 500K figure looks like a misread or an early-preview API limit, not the documented spec.
Pacing the frontier
14:22The meta-observation: OpenAI and Anthropic are "pacing the frontier," and in the meantime shipping genuinely excellent, cheaper, faster models (GPT 6.1 Soul, Opus 5.5). Berman's read: "if pacing the frontier actually means releasing incredibly good models — the best in the world, cheaper, faster than anything else — then I'm for pacing." The open-source counter-value he lands on: once many people get their hands on ML4's weights, they'll squeeze out performance for narrow use cases that beat the big closed models — but that's still an expert's game, not a general-audience one.
Claims checked
Mistral Large 4 ("Le Chonk") is real. Public preview Oct 6, 2026 — ~1.05T total / 49B active granular MoE, 1.6B vision encoder, natively multimodal input, hybrid instruct+reasoning, open weights promised by end of October. Covered by VentureBeat, The Register, MarkTechPost.
Pricing matches: $1.36/M input, $4.18/M output. Training: 3,800 Nvidia Grace Blackwell GPUs in Mistral's own European data center.
Context window: the video says 500K tokens; Mistral's documentation (per launch-day reports) says 1M tokens. Treated the 1M figure as the documented spec.
Auto-caption garbles normalized: "Lechonk/lechon"→Le Chonk, "Kimmy"→Kimi, "Quen"→Qwen, "JLM"→GLM, "Enthropic"→Anthropic, "Grock"→Grok, "Xiai"/"MIMO"→Xiaomi's MiMo, "Cloud Code"→Claude Code, "Codeex/codecs"→Codex, "chatbt"→ChatGPT, "Mistrol/Mistraw"→Mistral.
Benchmark numbers (Deep 1.1 62, Terminal Bench, Automation Bench 59.9, Cyber Gym 82, AA Cyber Index 50, AI Index 38/25th) are reported as the video's chart walkthrough; only the headline specs and pricing were independently confirmed. "Tibo" (the lab that announced a 50% speed bump) and "Musepark 1.3" are garbled and left unnamed.