DeepSeek V4 video thumbnail
Agent News · Analysis

DeepSeek V4: The Open-Source Threat to America's AI Lead

Source: Matthew Berman 17:21 · Published Apr 25, 2026 Subject: DeepSeek V4 (Apr 24, 2026)

0:00Matthew Berman opens by declining to do a normal model review. DeepSeek's new flagship, V4, is "massive, powerful, open-source, and a fraction of the cost" — but the story he wants to tell is bigger: America has the best chips and the most money, "yet China was able to release a frontier-level model that matches the best of them, completely open-source, completely open weights, at a fraction of the cost and resources — literally working with nerfed Nvidia GPUs. That's not supposed to be possible."

Who DeepSeek is

1:27The context: roughly 18 months earlier, DeepSeek released R1, an open-weights model that "could think" at a time when thinking models were a closed-lab monopoly — and "the stock market dropped 20% pretty much overnight." R1's real significance was efficiency: it was trained at a fraction of the cost of the hundreds of billions the US frontier labs were spending, which briefly convinced people Nvidia GPUs were overvalued. Berman notes the lesson people got wrong: "when things get cheaper in price, we actually use a lot more of it — that's Jevons paradox."

The specs: Pro and Flash

2:52The V4 preview shipped in two variants, both open-weights:

  • 3:00V4 Pro — 1.6 trillion total parameters, 49 billion active (mixture-of-experts), and a 1-million-token context length, "the frontier" of context limits.
  • 3:33V4 Flash — the workhorse: 284 billion total, 13 billion active; smaller, faster, much cheaper.

3:56Both were trained on roughly 33 trillion tokens. Berman highlights the "enhanced agentic capabilities," which he places "comparable to the state-of-the-art agentic coding models like Opus 4.7 and GPT-5.5 — literally the models that were just released in the last week from Anthropic and OpenAI."

"Behind, but just a little"

4:22Across MMLU Pro and GPQA Diamond, the chart tells one story: DeepSeek V4 Pro sits "slightly behind" the newest closed frontier models — "but just a little bit." That is the whole point, in Berman's framing. "The vast majority of use cases do not require the absolute frontier level of intelligence. And the fact that DeepSeek is so much more efficient and so much cheaper is actually the problem for the United States."

The cost story

5:19On the intelligence-vs-price chart, GPT-5.5 and Opus 4.7 sit at the very top; DeepSeek V4 Pro is a little lower on intelligence but "much, much cheaper," and Flash is an "absolute workhorse at pennies per million tokens." The historical Elo line is the kicker: the US opened a massive gap with GPT-4 (May 2023), then Qwen, GLM-4, and — right after o1 preview — DeepSeek R1 "closed the gap almost completely." Since then it's been "an ebb and flow: every time the US shoots ahead, Chinese open source catches up."

6:30Berman's refrain: "They have always been behind, but that might not always be the case."

Are export controls working?

7:20Export controls mean Nvidia can't sell its top chips (the GB300-class parts) to China directly — with widespread rumors of smuggling around them. Berman's answer is "kind of yes and kind of no": yes, because China demonstrably has less compute than the US; no, because "they are innovating on the algorithm side" — coming up with "incredible algorithmic unlocks" that make training and inference efficient enough to reach the frontier on nerfed or domestic GPUs. He notes Jensen Huang's argument that China will build its own chips and "might as well be built on American technology" — which is precisely why the V4 moment is economically consequential.

The distillation attacks

9:24A few weeks before this video, Anthropic published a report alleging "distillation attacks" — Chinese labs extracting Claude's capabilities by harvesting question-answer pairs to train their own models. "Those question-answer pairs are everything. That's the IP of companies like Anthropic and OpenAI."

10:07The US government followed: OSTP Director Misha Kratsios said the US "has evidence that foreign entities, primarily in China, are running industrial-scale distillation campaigns to steal American AI." The new part wasn't the accusation (Anthropic had already made it) but the government confirming it.

11:19But Berman reads the report against the grain. The scale of DeepSeek's alleged extraction — 150,000 exchanges — is small next to Moonshot/Kimi (3.4 million) and MiniMax (13 million). "150,000 exchanges is not really enough to explain the level of quality that DeepSeek has been able to achieve." Pair that with fully open weights and an unusually detailed, self-critical white paper, and "it just doesn't mesh" — a lot of what looks like distillation could be ordinary benchmark comparison. He leans on the white paper's own admission: "Due to constraints in high-end compute capacity, the current service capacity for Pro is very limited," with 950 supernodes coming online in the second half of the year to cut prices. "They are very compute-constrained — they were able to bake this model, but they can't even serve it in the most optimized way."

The real threat

12:51Here is Berman's thesis, stated plainly: DeepSeek V4 "doesn't need to be as good. Just being nearly as good is good enough for almost everybody — including enterprise companies in the United States."

"Why would you pay so much more for a US frontier lab to serve you their model, over an open-source Chinese model?"

13:24The CEO calculus: GPT-5.5 and Opus 4.7 run around $30 per million output tokens. DeepSeek is "literally a fraction," and because it's open-source you can fine-tune it, self-host it, and control it — at a fraction of the bill. "The calculus that these CEOs are making becomes very obvious."

15:12The stakes he sketches are large: trillions pouring into US AI infrastructure that requires a return, the potential for economic strain if that return evaporates, and — culturally — the inverse of the social-media era: if "we're all built on Chinese models," they dictate "what the models are able to say and what they're not able to say." His prescription: the US must go "much harder on open source" (only Google is really trying, and not at V4 scale), and it must drive down serving costs so building on US AI "makes sense cost-wise."

Claims checked

Verified

DeepSeek V4 is real. Released April 24, 2026. Two open-weights MoE variants — V4 Pro (1.6T total / 49B active) and V4 Flash (284B / 13B) — both with a 1M-token context, MIT-licensed. Specs match the video.

Verified

Distillation report. Anthropic's Feb 23, 2026 report names DeepSeek, Moonshot, and MiniMax, citing ~24,000 fraudulent accounts and 16M+ exchanges. The video's per-lab figures (DeepSeek 150K vs Moonshot 3.4M vs MiniMax 13M) are consistent with it. A later CISA advisory (Sep 2026) formalized the US government position.

Corrected

R1 market drop. The video says the stock market "dropped 20% overnight" on DeepSeek R1 (Jan 2025). The sharpest single-day move was Nvidia falling ~17% (~$589B), not the broad market dropping 20%.

Opinion

The geopolitical thesis (US enterprises migrating to Chinese open-source, economic-collapse risk, "they dictate the narrative") is Berman's argument, not established fact — treat it as analysis.

Sources & further reading

☰ View all