Did Kimi K3 Really Beat Fable? thumbnail

Did Kimi K3 Really Beat Fable?

Matthew Berman · ~12 min
Did Kimi K3 Really Beat Fable? — Matthew Berman
⏱ ~12 min 🎤 Matthew Berman 🏷 Kimi K3 🏷 Moonshot AI 🏷 Open Source 🏷 Fable 🏷 GPT 5.6 🏷 Benchmarks

1 The Headline: Kimi K3 Tops Front-End Benchmark

▶ 0:00

Moonshot AI just dropped Kimi K3 and Matthew calls it potentially "the next DeepSeek moment." This Chinese AI lab released what they claim is the best open-source/open-weights model on the planet. It's competitive with Fable 5 and GPT 5.6 — at least on some measurements.

The benchmark everyone points at: Arena AI's front-end development benchmark shows Kimi K3 at #1 — 76% vs Fable 5 at 63%. That's not a small margin.

▶ 0:44 — Arena AI front-end benchmark breakdown: K3 at #1.

2 Kimi K3 by the Numbers

▶ 1:26

2.8 trillion parameters — the biggest open-source model to date. You can't run this on a home computer; it needs data center serving. It sports a million token context window, designed for long-horizon coding, knowledge work, and reasoning.

For comparison: Thinking Machines' best open-source model is 975B parameters — it pales in comparison. And Kimi K3 is significantly better on the intelligence scale too.

The 1-minute demo video (which K3 itself edited) shows incredible 3D asset creation, simulated worlds, real-time reflections, dynamic daylight cycles — basically building a video game that looks like Red Dead Redemption.

▶ 2:09 — US vs China: open-source economics.

▶ 2:34 — K3 demo video (edited by K3 itself).

3 Pricing and Intelligence Density

▶ 3:06

Open source is known for efficiency and cost. Kimi K3 comes in at $3/million input tokens and $15/million output (with cache mix) — about half the price of GPT 5.6 Soul.

But you have to factor in intelligence density: what's the intelligence per token? On the DeepSWE benchmark (Matthew's current favorite): Kimi K3 Max sits right under GPT 5.6 Soul but at effectively the same price (~$4.70/task). K3 takes twice the tokens for the same task — so the half-price advantage disappears.

Bottom line: Token hungry and slow.

▶ 3:56 — DeepSWE benchmark: intelligence density analysis.

4 Industry Reactions

▶ 5:06

David Sacks (US AI Czar): "Concerning. First time a Chinese model has taken #1 on front-end code arena."

▶ 6:40 Guillermo Rauch (CEO, Vercel): "Kimi K3 is the best performing model on Next.js.org evals, ahead of Fable, reaching comparable success rate in less time. First time an open model is ahead of all proprietary ones for comprehensive web engineering."

92% success rate on agent performance results with agents.md. Also excels at writing: jumped from place 21 to #1 on an internal writing benchmark, 5x cheaper than the model it displaced.

▶ 7:15 — K3 excels at writing too.

5 Can We Trust These Benchmarks?

▶ 7:44

Important caveats: many benchmarks are completely saturated.

▶ 7:56 Anthropic accused Moonshot of distillation-attacking their models — essentially stealing data from Anthropic to train K3. Whether they stole enough to be meaningful is unknown.

But it's open source: you can go look at exactly how they built the model, replicate it. They revealed all algorithmic unlocks. We won't know the real quality until production testing.

6 The Real Gap: US Labs Are Still Ahead

▶ 8:32

Matthew argues the perception that open source has reached the frontier is misleading. Chinese/open-source labs release immediately when done baking. But Fable 5.1/5.2 has probably already been baked — Anthropic is just testing it.

Anthropic is likely 8–10 months ahead: they had Mythos back in January, 5 months before public release. GPT 6 is probably being tested internally at OpenAI right now. US closed-source labs are still far in the lead.

7 Why Chinese Open Source Benefits Everyone

▶ 9:57

It helps everybody — including US open source, OpenAI, and Anthropic. Algorithmic discoveries given away for free. Increased competition pressures US labs.

When open-source models are good: every part of the AI stack wins (except maybe closed-source labs). The chain reaction:

  • Models get better + cheaper
  • Jevons Paradox → more tokens used
  • Better apps built
  • Inference providers make more money
  • Nvidia sells more chips
▶ 10:58 The risk: if US enterprise builds on Chinese open-source models optimized for Chinese chips, it creates dependency on Chinese chips.

8 Live Demo: Rubik's Cube Test

▶ 11:14

Matthew kicked off a Rubik's Cube simulator build with K3. After 30 minutes it finished (criticism: K3 is quite slow).

Results: looks really good — reflections on cube sides, scramble works correctly, solve works perfectly. Another model that "absolutely crushes" the Rubik's Cube test.

But confirms K3 is token-hungry and slow.

▶ 12:04 — Final thoughts.

🎯 Key Takeaways

  1. Kimi K3 tops Arena AI's front-end benchmark at 76% vs Fable 5's 63%
  2. 2.8T parameter model — biggest open-source model ever, million-token context
  3. Half the price of GPT 5.6 but takes twice the tokens — effective cost is similar
  4. Excels at front-end dev, web engineering (92% on Vercel's Next.js evals), and writing
  5. Anthropic accused Moonshot of distillation attacks on their models
  6. US closed-source labs (Anthropic, OpenAI) are likely 8–10 months ahead of open source
  7. Chinese open source benefits the entire AI stack through competition and free algorithmic discoveries
  8. Risk: US enterprise dependency on Chinese open-source models optimized for Chinese chips
  9. K3 is token-hungry and slow — but produces high-quality output
  10. You don't need absolute frontier for the vast majority of enterprise work

⏱ Timestamps

▶ 0:00 Kimi K3 drops — "the next DeepSeek moment?" ▶ 0:44 Arena AI front-end benchmark: K3 at #1 ▶ 1:26 K3 specs: 2.8T params, 1M context ▶ 2:09 US vs China: open-source economics ▶ 2:34 K3 demo video (edited by K3 itself) ▶ 3:06 Pricing: $3/$15 per million tokens ▶ 3:56 DeepSWE benchmark: intelligence density ▶ 5:06 David Sacks reaction (US AI Czar) ▶ 6:40 Guillermo Rauch: best on Next.js evals ▶ 7:15 K3 excels at writing too ▶ 7:44 Can we trust these benchmarks? ▶ 7:56 Anthropic's distillation accusation ▶ 8:32 The real gap: US labs 8-10 months ahead ▶ 9:57 Why Chinese open source helps everyone ▶ 10:58 The risk: Chinese chip dependency ▶ 11:14 Rubik's Cube demo: results ▶ 12:04 Final thoughts