Deep Dive · Coding Agents

Fixing the PR Bottleneck: Three Brakes for Agent-Generated Code

✍️ Matt Pocock ⏱️ 23 min 📅 AI Engineer Paris 2026
Video thumbnail

1.The PR Bottleneck, Amplified by Agents0:01

Matt Pocock opens with a blunt claim: the pull request has always been the bottleneck — "huge numbers of PRs just laying around that no one's bothered to review" — and agents have made it far worse. "Agents make it trivial to open a PR. They make it only slightly easier to open one worth reviewing." The central promise of AI — scale yourself up to do more work — has produced the software factory, the buzzword of the moment, where the initiation of work passes from humans to agents.

His definition of a software factory matters: not a human typing prompts, but deterministic code triggering agents. A classifier (he nods to Jev) turns an issue into a fix or a reproduction; PlanetScale hooks report slow queries and trigger a workflow. All of it uninitiated by a human. The problem is what happens next — acceleration without brakes produces a "slop cannon": a ton of crappy PRs you can't even look at.

The thesis: "Code is the environment your agent operates in." Bad code in the codebase begets more bad code. So the fix isn't writing PRs faster — it's fixing what happens around the PR. Agent experience (AX) deserves the same rigor developer experience (DX) got.

2.Brakes: The Three-Layer Cake1:50

His framework is a three-layer cake. At the bottom, automated checks — the deterministic linting, tests, type checking, and code-quality metrics we've had "since the '50s," which work the same every time. On top, automated review — agents that look at the code and catch what the tests didn't, and examine overall structure. Finally, human review — people looking at the PR.

LayerWhat it doesCost
Automated checksLint, tests, type checking, metrics — deterministicCPU cycles (cheap)
Automated reviewAgent catches what tests miss; structureTokens
Human reviewFinal judgment on the PRHuman effort (expensive)

The goal: make human review faster by leaning on the first two phases. "Stop the slop" — raise the quality of shipped code and you'll need fewer human interventions. Automated checks are the cheap layer: they don't cost tokens or human effort, "just CPU cycles," so you can pile on "loads and loads" of them — and most teams aren't creative enough with their use. But checks have a fatal flaw: they can lie. Green CI does not mean ready to merge, which is why the upper two layers are "lie detectors" for the checks.

3.Checks Can Lie — and Agents Are Great at Writing Bad Tests5:14

Pocock shows three ways automated checks lie, all "real code from agents." First, tautological tests — a test that reasserts the implementation. His example: a constant X post character limit = 280, tested with expect(limit).toBe(280). Opus 5 "got addicted to these." They're bad because they're structure-sensitive: you can't change or rename the constant without the test failing.

Second, structure-sensitive tests that don't even execute code — one test checked that "videos" appeared after "content plan" in the UI by reading the source file into memory and searching for the strings in order. Change how the source looks and the test breaks. Third, tests that literally cannot fail — over-mocking (stubbing an audio-context API) until the test bypasses the very error modes it was meant to catch.

Important framing: "the AI isn't trying to write bad tests." It takes your instructions and writes tests tied to structure instead of behavior. So the question isn't "how do I stop agents from cheating" — it's "how do I make automated checks harder to cheat."

4.Fix It With Deep Modules8:25

The first fix is design: you can design your way out of bad checks. The key idea is deep modules, from John Ousterhout's A Philosophy of Software Design — modules that hide complex behavior behind a simple interface. Contrast module A (large implementation, tiny interface) with module B (large interface, lots of functions that each do very little). A deep module produces fewer structure-sensitive tests because the implementation is hidden; if tests only exercise the interface, they're forced to be behavioral. Your job is to make agents use that little interface instead of reaching into implementation details.

Pocock ships a codebase-design skill that runs on "the weirdest vibe-coded codebase" and produces an HTML document of before/after opportunities — reducing duplication, creating a deep, testable module. Alongside it, a shared vocabulary for modules: locality (how well-located code is, how a small change ripples out), leverage (how much value a caller gets from a simple function), and seams.

5.Don't Overload the Implementer: Review as a Sub-Agent11:13

Here's the mental model he wants you to keep: implementation is overloaded. The implementer agent's single context window already has to explore, change, and debug. Impose coding standards on top of that and it "performs worse." So don't put coding standards in your implementer agent — and definitely not in global scope (agents.md), where they "drown out" the implementer, which may or may not even read them.

His fix is the code-review skill: a separate sub-agent with its own context window and budget. It receives a diff, reads a coding-standards.md file you write and customize, and checks compliance. While implementation is overloaded, review is underloaded — it doesn't implement or debug, so you can pile coding standards into it and it'll do a much better job. He frames the whole thing as red-green-refactor: one context window makes it work, another makes it good.

The rule in one line: coding standards live in coding-standards.md, read only by the review agent — never in the implementer's prompt, never in global agents.md.

6.Build Your Own Reviewer — and Make It Commit, Not Comment14:22

On the "just use a third-party service" instinct (Cursor bug bot, CodeRabbit): he's tried building a generic review skill that finds all bugs and does security review, and it's "really, really hard" — make it too general and you get false positives irrelevant to your use case; too specific and only TypeScript (or only Rust) people can use it. So: don't outsource automated review — build your own coding standards over time and share them across the team. Those docs "sitting around that no one reads" are exactly what goes into coding-standards.md.

And the crucial design decision: a review agent that comments on the PR is creating more work for the human, who must read every verbose comment and decide what to do. The reviewer should commit — it should actually make the fixes it finds. Then the human reviews a clean artifact, not a wall of nits. It can comment when it genuinely has questions, but "the default should be commits." Stop trying to one-shot good code out of the implementer.

7.The Human-Friendly PR: Doors, Radius, and Diagrams16:28

With checks and automated review done, the last lever is making the human pass as painless as possible — the PR skill (in progress, shipping in v1.3 of his skills). Three principles. First, some reviews matter more than others: borrow AWS's terminology and ask whether a PR is a one-way door or a two-way door. Most PRs are two-way — merge, revert later; the glorious difference between software and civil engineering. But a "simple change" that blasts an email to 60,000 people, or an expensive migration with data loss, is a one-way door: review the hell out of it. Tied to this is blast radius — what can go wrong, and how bad — distilled into a one-line "merge danger" summary at the bottom of every PR ("two-way door, blast radius localized").

Second, make the why fast to grasp: use pseudo-code and diagrams instead of prose, crediting Dex Hadley's "show me" skill — Mermaid diagrams, UML, and simple visual summaries ("new command added to the CLI, two new flags"). Third, the meta principle: review the system that produces the code, not just the code. "The process that produces your code is just as important as the code itself" — and you never want to write the same review comment twice.

8.Retro: Making Human Review Compound20:17

The closing skill, Retro, closes the loop. Take a session, or a PR plus its session, or all the PRs and reviews from the last week, and run a retrospective. It suggests automated checks and coding-standards.md updates so the next PR is better — the compounding effect where doing human review raises the quality of the next human review and saves you work next time.

It goes further than standards: navigation pointers (did the agent find its information? add a pointer in agents.md), tool economy (are there more token-efficient tools — "amazing how many things this catches"), and bloat (bloated steering files and skills that produce bad results). The whole talk is one argument, restated at the end: layer up automated checks, layer up automated review, and make human review as painless — and as optional — as it can be. "You really don't need to review every single two-way door. Every single one-way door, you do."

Key Takeaways

  1. The PR was always the bottleneck; agents make it trivial to open PRs but only slightly easier to open ones worth reviewing — so a software factory without brakes becomes a "slop cannon."
  2. The fix is a three-layer cake: cheap deterministic checks, automated review, then human review — with each upper layer acting as a "lie detector" for the one below.
  3. Checks lie in three ways: tautological tests (reassert the implementation), structure-sensitive tests (read source instead of executing), and over-mocked tests that cannot fail.
  4. Deep modules (Ousterhout) reduce structure-sensitive tests — hide implementation behind a small interface and force agents to test behavior.
  5. Implementation is overloaded, review is underloaded. Keep coding standards out of the implementer; put them in coding-standards.md and let a review sub-agent enforce them.
  6. Don't outsource automated review — generic reviewers give false positives; build your own standards over time.
  7. A review agent should commit fixes, not comment — comments create more work for the human; a clean artifact doesn't.
  8. Categorize PRs as one-way vs two-way doors with a "merge danger" blast-radius line, use diagrams over prose, and review the system that produces the code.
  9. Retro turns every human review into better automated checks and standards for the next PR — the compounding loop.

Timestamp Index

☰ View all