The Pragmatic Engineer Podcast

AI Skills with Matt Pocock: Grill Me, Wayfinder, and Software Fundamentals

Matt Pocock 1:36:36 Gergely Orosz
AI Skills with Matt Pocock

1. The 275k-Star Skill

▶ 0:00

Matt Pocock created the grill me skill — the one that interviews you relentlessly about what you are building before you write a line of code. His skills repo, mattpocock/skills, is now one of the most-starred on GitHub (275,890 stars at the time of writing, MIT, active) — and he says it is the second most-starred skills repo in the world, somewhere around the 20th-to-25th most-starred repo of all time.

Before the skills, Matt was best known as the creator of Total TypeScript, the course that generated over $2.5 million in sales. His path there is unusual — six years as a voice coach before touching code — and it turns out that background explains a lot about why his skills work the way they do.

2. From Voice Coach to Developer

▶ 5:48

For six years he was a voice coach — a singing teacher who taught accents, taught public speaking, ran Shakespeare workshops at drama schools, and coached consultants at big companies on how to deliver speeches. He did a master's in it and genuinely thought that was his career. What pushed him out was geography: doing it well meant living in London, and he hated London.

So he taught himself to code in order to get a remote job, building a web-audio analyzer to see the resonant frequencies of a voice. The communication skills turned out to be an "unfair combination": he could go into interviews and sound like a reasonable person, explain technical ideas clearly, while his technical knowledge — which he was passionate about — grew quickly. He rose through the ranks fast, "different from the other software developers."

3. Open Source, Vercel, Total TypeScript

▶ 10:14

His open-source break came through XState: building type-safe tooling around the state-machine library got him noticed by its creator David Khourshid, and he became a core team member at Stately — his first job paid in "American money," which changed how he thought about compensation and flexibility. From there he was recruited to Vercel under Jared Palmer, where he took an unusual three-days-a-week contract.

The contract was a deliberate hedge: he was already floating Total TypeScript, and Vercel was a "weird backup" while he tested whether the course would work. It did — a pre-release sale earned 30-to-40x his Vercel salary almost immediately, and the course reached seven figures within months, later publicly passing $2.5 million in total revenue. His model, worked out with Egghead's Joel Hooks, was honest: a product people pay for because they want to learn, an extended refund policy, and a refusal of most other forms of money.

4. AI's Impact: Knowledge vs Wisdom

▶ 23:21

Then AI arrived. Matt's framing of what it changed is the thread that runs through the rest of the episode: he teaches two layers — the what (syntax, knowledge) and the why (wisdom, judgment). Knowledge has become very cheap; you can just look it up or have an agent teach it to you. But wisdom has gotten no easier to learn — you still run into the same issues you would have without it.

That is why Total TypeScript revenue has fallen: the tactical material it taught is now commoditized. His own inflection point came around December of last year — the "Opus 4.5 winter break" where everyone came back AI-pilled and realized you can now genuinely delegate to agents. Drawing on John Ousterhout's distinction, his conclusion is that AI has largely eaten tactical programming, and the strategic layer is what is left for humans — which is exactly where the skills live.

5. Skills as Distribution

▶ 30:32

His years trying to put agents into applications turned out to be the ideal training: building an app that contains an agent forces you to think about data flow — where state lives, how it is passed in, what shape it takes, how it is compacted. When he moved to harnesses, it all felt familiar. He started stringing agents together into "loops" (which he prefers to think of as processes, or finite state machines — a direct echo of XState) and saw dramatically better results than the default setup.

The distribution mechanism he landed on was skills: a folder of markdown files the agent can invoke itself (model-invoked) or that you invoke with a slash command (user-invoked). He put his process up as a set of skills, did almost no marketing, and came back to find it had more stars than anything he had ever made — "word of mouth." A talk at AI Engineer London, titled Software Fundamentals Still Matter, pushed it to 1.2 million views, and the repo kept climbing.

6. The Grill-Me Skill

▶ 33:12

The most popular skill is a small one: get the agent to interview you relentlessly about the topic. Matt credits the idea to Tarik at Claude Code, and its emergent behavior is the surprise — the models start thinking outside the box and throwing ideas at you. Gergely's own grill-me story: designing a simple API endpoint, the skill asked 35 questions, digging into whether he wanted a bearer token or JSON, and how to enforce rate limits.

Two things make it valuable. First, it forces you to make the decisions, while reminding you of things you forgot — it sent him off to research authentication methods, making him a better professional. Second, it closes what Gergely calls the communication gap: the agent cannot read your mind or your hierarchy of values, so grilling is how the agent gets to know you — what is in scope, what matters, what to optimize for. Matt's own line: "As long as we're learning, I think we're fine. As long as we stop learning and outsource the learning to this thing, trouble will be brewing."

7. The Smart Zone vs the Dumb Zone

▶ 40:46

The deeper idea comes from Dex Horthy, a big influence: every token is shouting for attention, and the more voices you put in the room, the harder it is to hear the important ones. A model's performance degrades slowly as the context window fills, but there is a region where it is sharp — the smart zone, roughly the first 150,000 tokens of a frontier model's window, regardless of the window's total size.

The practical question becomes: how do you take work larger than 150k tokens and portion it across multiple context windows and sessions? That is what the Ralph loops were one attempt at, and what the skills are built to solve systematically.

8. Ralph Loops, Specs, and the Night Shift

▶ 42:15

A Ralph loop gives the agent a goal, tells it to make the smallest possible change that gets closer, then clears the context and starts fresh — with the codebase itself holding the only persistent state. Matt's skills formalize this into two documents: a spec (the destination — what it means to be done) and tickets (one per session, broken out of the spec). The flow is: grill → turn the grilling into a spec → run an implement loop over the tickets until the work is done.

He designed this for the day shift / night shift pattern: plan during the day, let agents work while you sleep, wake up to clean code. The goal is to stop context-switching between terminals and instead spend decent 15-minute chunks planning, then review the agent's output in bulk. The realization that made it possible: as of December, agents are good enough to delegate to, so you can run them AFK.

9. Wayfinder: The Map Metaphor

▶ 45:02

Sometimes a single grilling session is not enough — "build me a Stripe clone" cannot be planned in 150k tokens. Wayfinder solves infinite-length planning with a map: a center point holding everything decided so far, with milestones on it and a fog of war hiding the unexplored territory. Every grilling session opens more points on the map, and the process walks a directed graph of tickets until it reaches the destination.

Tickets come in types — not just implementation, but prototyping, research, even provisioning infrastructure. Matt has run maps with 50 to 100 tickets, and found the metaphor transposes beyond engineering: he has used wayfinder for course planning and to design the garden office in his garden. Engineering, he notes, is just a discipline; once you see it as discussing and doing things in the world, the tools map onto any domain.

10. Why Agents Excel at Software Engineering

▶ 47:52

Gergely brings up Hillel Wayne's observation that software's "material" behaves differently from other engineering — deterministic, unlike the load-bearing thresholds of mechanical or civil engineering — until LLMs introduce that variance into code. Matt's answer is more direct: agents are good at software because every input and output is text — code, docs, instructions in, more code, test results, type-checking and lint output out.

Anything non-text is where agents struggle. Those one-shot "perfect UI" demos fall apart the moment there is an interaction problem — a hover animation that looks wrong cannot be described to the agent, and vision is not good enough yet. His rule: the more you can make your work agent-friendly — turn your services into text the agent can consume — the better results you get. That is exactly what he is doing now, plugging all his services into agents.

11. Leading Words: Mining Old Books

▶ 50:54

When spec-driven development kept producing garbage code, Matt opened a book still wrapped in plastic: The Pragmatic Programmer. Its chapter on software entropy explained what he was seeing — agents produce entropy at a higher rate than ever. Almost every line felt written for today: don't outrun your headlights, work within your feedback loops, programming by coincidence, tracer bullets.

Then he discovered the trick that became central to his whole method: these phrases are in the model's training data. Start saying tracer bullet or vertical slice in your prompts, and the model repeats them back and changes its behavior — a leading word, a "leitmotif" that leads the agent with a phrase repeated a couple of times in the skill. Tracer bullets (get feedback fast on one thin path that works) directly fixes the worst agent habit: building the whole database, then the whole app layer, then the whole React library, and only plugging them together at the end. John Ousterhout's Philosophy of Software Design added deep modules, another leading word. He now mines old books systematically for these terms.

12. Domain Language and Memento-Driven Development

▶ 57:10

Agents are awfully verbose, so Matt went looking for a common language — and found it in Eric Evans' Domain-Driven Design and the concept of ubiquitous language, again deep in the model's priors. This became the (badly named, he admits) grill-with-docs skill, which builds a domain language as you go. In his own app, a tangle of "ghost lessons inside ghost sections inside ghost courses becoming real" collapses into one term the agent can navigate by: the materialization cascade.

The payoff is night and day: you describe changes in far fewer words, and the agent can grep the codebase for functions using that terminology. The deeper principle is what he calls memento-driven development — an agent is like the protagonist of Memento, waking up every morning with no memory. A human can work around a bad codebase by slamming their head against it until they build memory; an agent starts fresh every session. So you must optimize the codebase for that amnesiac newcomer — and it turns out software fundamentals have been telling us to do exactly that the whole time.

13. Learning the Fundamentals

▶ 1:01:10

Why is strategic programming hard to learn? Because the feedback loop is brutally long — a strategic mistake takes nine months to come back and get you, so people who quit after six months never see the consequences. Matt's metaphor: a giant mixing desk of sliders (microservices vs monolith, and so on), where you are mastering a track but cannot hear what is wrong until months later.

AI changes this in one specific way: it lets you move faster, so your strategic mistakes come back quicker. The core idea is that your code is the environment the agent operates in — you should always be improving that environment. The strategic layer has become more valuable than ever, because you can now get so much leverage out of it. And on the classic question of how to convince stakeholders, his answer is blunt: you could have asked the same question ten years ago about tech debt, and the answer starts with observability over what your agents are actually doing.

14. Observability and the Experimental Mindset

▶ 1:06:00

The first step to justifying fundamentals investment is observability over every agent — its success and failure rates. You could never instrument human developers this way (it would be invasive), but for agents it is fine; you are paying for the service, so you need to know how well you are using it. Some repos will have better success rates than others, and you extract and propagate those lessons.

The second step is a common set of skills — a shared workflow everyone can contribute back to and A/B test, one team on one approach, another team on another, then compare. Everyone working with agents needs this experimental mindset: how do we get more juice out of the tokens we are spending? Matt also runs automated improvement — every morning an improve-codebase-architecture skill proposes a change, which he can turn into tickets and ship. He floats the idea of spending 20% of your time on the factory that builds your software, not just the software.

15. Local vs Cloud Agents

▶ 1:09:17

Matt's recent provocative tweet: "I'm moving away from my local dev setup. Makes zero sense to me now." The driver is collaboration — a shared grilling session needs more than your terminal; you need a space where you can tag someone in and say "do this," which means Slack, Discord, Teams, or Linear. He already chats with his agent through Discord from the train, fixing bugs for his course.

The remote box is always on, so he can schedule a morning standup where the agent plans his day against his Discord chats — and a remote setup makes provisioning resources on demand easier than a thousand git work trees and Docker containers locally. Gergely confirms the trend from Ramp, Stripe, and Uber: 70-to-80% of devs are voluntarily going to cloud agents, with the main exception being front-end work, where you still want the fast local feedback loop.

16. Planning vs Course-Correcting

▶ 1:12:36

The devil's-advocate question: if agents are so fast, why not skip upfront planning and just course-correct? Matt's rule is a simple three-way split. Small enough to align afterward — a five-line bug fix or a button moved three pixels — don't grill, just shift right. Fits in a single session — use grill me. Spans multiple sessions, where wrong code will pollute the context window and be expensive to rework — use wayfinder. The key criterion: if the agent gets it wrong, is the wrong code going to influence everything that follows?

He already runs this in production: a feedback button in his video editor files a GitHub issue, an implement agent picks it up, a review agent checks it, and he does his alignment at the end. On the "isn't this waterfall?" criticism, he notes agents make prototyping the cheapest it has ever been — churn out three or four versions and pick your favorite — and Grady Booch's point that the problem with waterfall was never the plan-build-ship cycle, it was planning taking a year and implementation taking three.

17. TDD, Tech Debt, and Gardening

▶ 1:18:13

On TDD, Matt is genuinely mixed. TDD optimizes for small working memory — the failing test reminds a distracted human where they were — but agents have a much larger working memory, so TDD aims at the wrong problem. What agents do need is feedback loops, and building the failure first is hard for an agent to cheat. His practical take: even when not doing strict TDD, ask the agent to "give me TDD evidence" — prove the change does what it claims. The failure mode is agents writing tautological tests that just assert the implementation.

On tech debt, quoting Jared Friedman: "technical debt used to be something you had to live with in a large codebase — no longer," to which Matt replied, "yes, now you can live with it even in a tiny codebase." Agents are great at producing rubbish because they cannot think strategically, so he pairs an implement agent with an automated review agent that enforces coding standards. And on the gardener metaphor, his reply was that the only thing a team needs is gardeners — because we are now our agents' platform team, building the environment for them to succeed and diagnosing entropy before it becomes a problem.

18. Teaching, Juniors, and the Three Books

▶ 1:24:21

On teaching, the human part survives because it is curation: information is a graph, and a good teacher finds the linear path through it — "Dijkstra's algorithm through the graph" — so you learn in the most sensible order. That is strategic work, and AI is not good at it. For juniors, his advice is to use the agents as much as possible and stay introspective — a "navel-gazing programmer" constantly examining their own process. The people thriving now are the same ones who thrived ten years ago: curious and adaptable.

His book list closes the loop on the whole episode. The Pragmatic Programmer, Philosophy of Software Design by John Ousterhout, and the first three chapters of Domain-Driven Design by Eric Evans — three books, 20-plus years old, that turn out to be the best instruction manual for working with AI agents. There is real irony, Gergely notes, in the fact that the best practices for this new technology were documented decades ago — and only became more important once the agents arrived.

Key Takeaways

  1. Knowledge is cheap, wisdom is not. AI ate tactical programming; the strategic layer is what is left for humans.
  2. Skills are folders of markdown — the distribution mechanism for a repeatable workflow, model-invoked or slash-invoked.
  3. Grill me closes the communication gap — the agent interviews you so it learns your values, scope, and priorities before writing code.
  4. Stay in the smart zone — roughly the first 150k tokens — by splitting work into a spec plus one-ticket-per-session.
  5. Day shift plans, night shift builds — grill into a spec, run the implement loop AFK, wake up to code.
  6. Wayfinder maps infinite planning with a fog-of-war graph of milestone tickets, spanning 50-100 tickets.
  7. Leading words steer the model — "tracer bullet," "vertical slice," "deep module," "ubiquitous language" are in the priors, so naming the concept changes behavior.
  8. Memento-driven development — agents start fresh every session, so optimize your codebase for an amnesiac newcomer.
  9. Your code is the agent's environment — observability over agent success rates, a shared skill set, and an experimental mindset.
  10. We are our agents' platform team — the gardener role is now the essential skill.
  11. Read the old books: Pragmatic Programmer, Philosophy of Software Design, and DDD's first three chapters.

Timestamp Index

  • 0:00 — Intro
  • 5:48 — How Matt got into tech
  • 10:14 — Open source, Vercel, Total TypeScript
  • 23:21 — AI's impact on education
  • 30:32 — Building reusable skills
  • 40:46 — The smart zone vs the dumb zone
  • 45:02 — The wayfinder skill
  • 47:52 — Why agents excel at software engineering
  • 50:54 — Leading words
  • 1:01:10 — Learning the fundamentals
  • 1:09:17 — Local vs cloud agents
  • 1:12:36 — Planning vs course-correcting
  • 1:18:13 — TDD and agents
  • 1:24:21 — Teaching and junior engineers
  • 1:31:07 — Gardeners and great engineers
  • 1:34:01 — Book recommendation
☰ View all