Matt Pocock opens with a blunt claim: the pull request has always been the bottleneck — "huge numbers of PRs just laying around that no one's bothered to review" — and agents have made it far worse. "Agents make it trivial to open a PR. They make it only slightly easier to open one worth reviewing." The central promise of AI — scale yourself up to do more work — has produced the software factory, the buzzword of the moment, where the initiation of work passes from humans to agents.
His definition of a software factory matters: not a human typing prompts, but deterministic code triggering agents. A classifier (he nods to Jev) turns an issue into a fix or a reproduction; PlanetScale hooks report slow queries and trigger a workflow. All of it uninitiated by a human. The problem is what happens next — acceleration without brakes produces a "slop cannon": a ton of crappy PRs you can't even look at.
His framework is a three-layer cake. At the bottom, automated checks — the deterministic linting, tests, type checking, and code-quality metrics we've had "since the '50s," which work the same every time. On top, automated review — agents that look at the code and catch what the tests didn't, and examine overall structure. Finally, human review — people looking at the PR.
| Layer | What it does | Cost |
|---|---|---|
| Automated checks | Lint, tests, type checking, metrics — deterministic | CPU cycles (cheap) |
| Automated review | Agent catches what tests miss; structure | Tokens |
| Human review | Final judgment on the PR | Human effort (expensive) |
The goal: make human review faster by leaning on the first two phases. "Stop the slop" — raise the quality of shipped code and you'll need fewer human interventions. Automated checks are the cheap layer: they don't cost tokens or human effort, "just CPU cycles," so you can pile on "loads and loads" of them — and most teams aren't creative enough with their use. But checks have a fatal flaw: they can lie. Green CI does not mean ready to merge, which is why the upper two layers are "lie detectors" for the checks.
Pocock shows three ways automated checks lie, all "real code from agents." First, tautological tests — a test that reasserts the implementation. His example: a constant X post character limit = 280, tested with expect(limit).toBe(280). Opus 5 "got addicted to these." They're bad because they're structure-sensitive: you can't change or rename the constant without the test failing.
Second, structure-sensitive tests that don't even execute code — one test checked that "videos" appeared after "content plan" in the UI by reading the source file into memory and searching for the strings in order. Change how the source looks and the test breaks. Third, tests that literally cannot fail — over-mocking (stubbing an audio-context API) until the test bypasses the very error modes it was meant to catch.
The first fix is design: you can design your way out of bad checks. The key idea is deep modules, from John Ousterhout's A Philosophy of Software Design — modules that hide complex behavior behind a simple interface. Contrast module A (large implementation, tiny interface) with module B (large interface, lots of functions that each do very little). A deep module produces fewer structure-sensitive tests because the implementation is hidden; if tests only exercise the interface, they're forced to be behavioral. Your job is to make agents use that little interface instead of reaching into implementation details.
Pocock ships a codebase-design skill that runs on "the weirdest vibe-coded codebase" and produces an HTML document of before/after opportunities — reducing duplication, creating a deep, testable module. Alongside it, a shared vocabulary for modules: locality (how well-located code is, how a small change ripples out), leverage (how much value a caller gets from a simple function), and seams.
Here's the mental model he wants you to keep: implementation is overloaded. The implementer agent's single context window already has to explore, change, and debug. Impose coding standards on top of that and it "performs worse." So don't put coding standards in your implementer agent — and definitely not in global scope (agents.md), where they "drown out" the implementer, which may or may not even read them.
His fix is the code-review skill: a separate sub-agent with its own context window and budget. It receives a diff, reads a coding-standards.md file you write and customize, and checks compliance. While implementation is overloaded, review is underloaded — it doesn't implement or debug, so you can pile coding standards into it and it'll do a much better job. He frames the whole thing as red-green-refactor: one context window makes it work, another makes it good.
coding-standards.md, read only by the review agent — never in the implementer's prompt, never in global agents.md.On the "just use a third-party service" instinct (Cursor bug bot, CodeRabbit): he's tried building a generic review skill that finds all bugs and does security review, and it's "really, really hard" — make it too general and you get false positives irrelevant to your use case; too specific and only TypeScript (or only Rust) people can use it. So: don't outsource automated review — build your own coding standards over time and share them across the team. Those docs "sitting around that no one reads" are exactly what goes into coding-standards.md.
And the crucial design decision: a review agent that comments on the PR is creating more work for the human, who must read every verbose comment and decide what to do. The reviewer should commit — it should actually make the fixes it finds. Then the human reviews a clean artifact, not a wall of nits. It can comment when it genuinely has questions, but "the default should be commits." Stop trying to one-shot good code out of the implementer.
With checks and automated review done, the last lever is making the human pass as painless as possible — the PR skill (in progress, shipping in v1.3 of his skills). Three principles. First, some reviews matter more than others: borrow AWS's terminology and ask whether a PR is a one-way door or a two-way door. Most PRs are two-way — merge, revert later; the glorious difference between software and civil engineering. But a "simple change" that blasts an email to 60,000 people, or an expensive migration with data loss, is a one-way door: review the hell out of it. Tied to this is blast radius — what can go wrong, and how bad — distilled into a one-line "merge danger" summary at the bottom of every PR ("two-way door, blast radius localized").
Second, make the why fast to grasp: use pseudo-code and diagrams instead of prose, crediting Dex Hadley's "show me" skill — Mermaid diagrams, UML, and simple visual summaries ("new command added to the CLI, two new flags"). Third, the meta principle: review the system that produces the code, not just the code. "The process that produces your code is just as important as the code itself" — and you never want to write the same review comment twice.
The closing skill, Retro, closes the loop. Take a session, or a PR plus its session, or all the PRs and reviews from the last week, and run a retrospective. It suggests automated checks and coding-standards.md updates so the next PR is better — the compounding effect where doing human review raises the quality of the next human review and saves you work next time.
It goes further than standards: navigation pointers (did the agent find its information? add a pointer in agents.md), tool economy (are there more token-efficient tools — "amazing how many things this catches"), and bloat (bloated steering files and skills that produce bad results). The whole talk is one argument, restated at the end: layer up automated checks, layer up automated review, and make human review as painless — and as optional — as it can be. "You really don't need to review every single two-way door. Every single one-way door, you do."
coding-standards.md and let a review sub-agent enforce them.