Pi Agent code mode video thumbnail
Pi Agent · Deep Dive

Pi's 180: Code Mode and Native MCP Land in the Minimalist Agent

Source: kintu 11:25 · Published Oct 3, 2026 Subject: Pi 1.0 (pi.dev)

0:00For much of the past year, Pi — the minimalist self-modifying coding agent from Mario Zechner and Armin Ronacher — was one of MCP's loudest critics. Its site even declared "no MCP," with the founders arguing the protocol was token-wasting bloat and that agents were "often better off just using bash CLI tools and code." This week that declaration is gone: Pi 1.0 shipped with native MCP support and a new feature called code mode, an embedded JavaScript sandbox designed to make MCP dramatically more efficient. This deep dive explains what changed, what the video's benchmark actually measured, and why the community is split.

The 180: from "no MCP" to native MCP

The timing matters. Pi reached its 1.0 milestone on October 1, 2026, and the MCP inclusion is the headline change — notable because Pi creator Mario Zechner had once dismissed MCP as unnecessary. Pi is the product of two Austrian long-timers: Zechner (libGDX) and Ronacher (Flask), building under the name Earendil. Their philosophy is written into the project: "if I don't need it, it won't be built." The MCP reversal is therefore a real story, not a small feature note — the Register's headline was "Pi coding agent pulls a 180 and adds MCP support."

Why MCP felt like bloat

0:59The video opens with the concrete failure mode. Ask an agent to "check 50 customer support tickets" the traditional way and it does the hard way: query ticket one, read a "giant pile of raw data" into the context window, query ticket two, repeat. The metaphor: "like hiring a master chef, but instead of letting them cook, you force them to inspect every single potato in a truck one by one." All that raw intermediate data clogs the context window.

2:35MCP is the "universal USB cable" that connects an agent to GitHub, Linear, Slack, or a database — one protocol instead of N integrations. But Pi's critique was never the protocol itself: it was how agents use it. "You can connect an MCP server and suddenly the model gets handed a giant technical manual explaining every tool, parameter, and description before you even ask it to do anything." Connect enough tools and "you burn thousands of tokens explaining buttons the model might never press."

Code mode: the QuickJS sandbox

1:37Code mode's fix is structural. Instead of the AI acting as middleman for every tool call, it writes a small JavaScript script that runs inside Pi's embedded QuickJS sandbox. The script can loop through data, call tools in parallel, do the math, remove duplicates, and filter out junk — and only the distilled result (a clean table, five matching tickets, three critical issues) returns to the model. "The bulky intermediate data stays inside the sandbox instead of flooding the model's context," which "massively reduces token usage, cost, and context load."

Pi's stance, as the video puts it: "Fine, we'll support MCP, but we're not dumping all of it into the context." MCP tools are exposed as callable JavaScript functions inside code mode, so the sandbox handles the call and the model sees a normal programming interface — not "a giant pile of chat tools."

Progressive disclosure

3:17If the model no longer has every tool sitting in front of it, how does it know what exists? Pi's answer is tool search + progressive disclosure. Instead of loading the whole catalog, the model gets a lightweight overview — "you've got access to these categories of tools." Ask to "update my DNS record" and the AI doesn't read 2,000 unrelated tools; it searches DNS specifically, finds the relevant tools, and inspects only the one it needs.

Inside code mode that's exposed as functions like search_tools and describe_tool. The video's framing: "instead of giving the model the entire cookbook, Pi gives it the table of contents and lets it pull out the exact recipe when needed."

System-one models in the loop

4:32A second trick: code mode can call a small "system one" model — a fast classifier like Jev, or an open-weights alternative — directly from inside the JavaScript loop. Pull 200 GitHub comments through MCP, and instead of sending all 200 back to the frontier model, the script passes each one through the classifier first ("is this user frustrated? is this urgent? what category?"). The small model handles hundreds of cheap semantic decisions; the frontier model only sees the final filtered or aggregated result.

This is the layering Pi is pushing: MCP for fetching, code mode for processing, a system-one model for fast semantic decisions, and the frontier model for high-level reasoning — "layers to protect the model from bloat and busy work it really doesn't have to do."

The benchmark: plain MCP vs code mode

5:37The video runs two like-for-like tests using GPT-5.6 (thinking set to high). Test one: compare 10 SUVs via the NHTSA MCP server (recalls, complaints, crash ratings, investigations).

  • 6:21Plain MCP: "hit a wall almost immediately" — raw payloads dumped straight into context, six of ten never completed. 95,000 cache tokens, 26¢, and only 4 of 10 ranked.
  • 7:05Code mode: the model wrote a script to fetch, filter, dedupe, and calculate inside the sandbox. Ranked all 10 with zero truncations. Cache input 95,000 → 21,000 (~78% less), cost 26¢ → 9¢.

7:25Test two: analyze 100 public VS Code issues from September. Plain MCP "stopped at 50 fully processed threads" because "large source responses and repeated output truncations leave insufficient context to reliably complete all 100" — over a million total input tokens if you count the cache.

8:13With code mode plus a local 12B classifier model running through llama.cpp, it completed all 100. Uncached input dropped from ~92,000 to ~30,000 tokens (67% less), and cost fell from 35¢ to 14¢. The video's conclusion: "in a production pipeline, this kind of stuff really, really matters."

The community pushback

9:14Alongside the update, Earendil published the "no MCP" reversal, and the response split into three camps:

  • The who-told-you-so crowd: the anti-MCP wave "ignored the enterprise side entirely" — remote credentials, security boundaries, telemetry, managing tools across teams.
  • The middle ground: MCP might be flawed, "but so are USB-C and HDMI" — one massive ecosystem beats "seven different perfect ways of doing the same thing."
  • The minimalist camp: "So Pi is also accruing cruft now" — worry that Pi is becoming the bloated harness it originally tried to avoid.

10:04Armin Ronacher jumped into the replies to push back: "Nothing is loaded by default that was not loaded before," and code mode "isn't just bash with extra steps" — bash runs shell tools, while code mode lets the model orchestrate tools directly inside the harness. The video's closing verdict is the VHS/Betamax line: "sometimes you just have to accept that marketing wins over engineering purity." MCP has flaws, but it won the ecosystem and the distribution battle.

The wider trend

10:42Pi isn't alone. The video notes Claude Code and Codex have both moved toward lazy tool loading — MCP tools searched for and pulled into context only when needed — and OpenCode "runs MCP through its own code mode by default," very close to Pi's approach. The through-line: the industry has converged on MCP as the wire protocol, and the differentiator is no longer "do you support MCP" but how much of it you force into the context window. Code mode — in some form — is becoming the default answer.

Claims checked

Verified

Pi 1.0 + MCP. Pi (pi.dev) reached 1.0 on Oct 1, 2026 with native MCP support (stdio + streamable HTTP). The Register independently reported the "180" — Mario Zechner had previously dismissed MCP.

Verified

Creators. Mario Zechner (libGDX) and Armin Ronacher (Flask), building Pi under the name Earendil. The auto-caption renders Earendil as "Arendelle" — a known transcription garble.

Verified

Code mode = QuickJS sandbox. Confirmed by Pi's own docs and third-party guides: "a QuickJS sandbox that lets the model call tools via JavaScript," plus a model router.

Corrected

Garbled names. The auto-caption mangles several proper nouns: "Cloud Code"→Claude Code, "CodeDex"→Codex, "Open Code"→OpenCode, "Jeff"→Jev (TypeSafe's classifier). The local 12B classifier's name is unclear in the caption and is described generically above rather than guessed.

Note

Benchmark is the creator's own. The 78%/67% token reductions and cost figures (26¢→9¢, 35¢→14¢) come from the video author's single-run tests, not an independent study. Directionally consistent with the mechanism, but treat as an anecdote, not a measurement.

Sources & further reading

☰ View all