HangarDX Podcast ยท AI-Era Software

Uncle Bob: Why He Stopped Reading Code and Building Harnesses

Robert C. Martin (Uncle Bob) 35:03 Ankit Jain ยท Aviator
Uncle Bob on the AI era

1. The Thesis: Humans Are Too Slow for Code

โ–ถ 0:00

Uncle Bob opens with the working hypothesis that frames the whole conversation:

"Human beings are too slow when dealing with code. If we're going to move at the speed of the agents, we're all going to have to back off a little bit and maintain high-level supervision โ€” but at the level of code, that's probably out of our domain now."

Robert C. Martin โ€” co-author of the Agile Manifesto, author of Clean Code, programming since 1964 in assembly, COBOL, Fortran, C, C++, Java, and "a billion other languages" โ€” is now deep in the AI world. The author of the most influential book about reading code has stopped reading it.

2. How Uncle Bob Works Today

โ–ถ 1:48

His day-to-day no longer involves an IDE. He barely ever sees code โ€” "sometimes I do, but very seldom." His primary tool is a terminal window running the Grok TUI, though he has also used Codex and Claude. He tells it what he wants, how it should get there, which tools to use, and which constraints and thresholds to apply โ€” then watches it very carefully.

The supervision happens at the level of architecture: he tries to extract as much intelligence about what the agent has done as possible, without forcing himself down to the code. The reasoning is the thesis from the top โ€” at the speed agents move, a human reading line-by-line is the bottleneck.

3. What Changed His Mind

โ–ถ 3:12

Two years ago he publicly said AI is not the future of coding. Even a year ago, agents did things "rather stupidly" and he was convinced they would never go anywhere. Then, around January, the inflection point: he could ask agents to do something and they did it better โ€” not great, but better. "Something's going on here."

Virtually every month after that it got better. He describes extracting himself from a "rat hole" โ€” a two-to-three-month effort building constraints to keep the models in place โ€” only to realize the models have now completely outgrown those constraints. The models do "extremely good jobs. Not perfect, but very good."

4. Three Constraints: Tests, CRAP, Mutation

โ–ถ 4:52

He still uses the same constraints as three months ago, just with less austerity. The first is unit tests, which he calls a "double-entry bookkeeping" approach: the agent has to say everything twice, and that keeps it in line. Without tests, agents "go off into wonderland." He has watched an agent break a test and then back off, realizing it needs a different direction โ€” visible even as the code streams by in the window.

The second is CRAP, a metric from around 2007 that combines test coverage with cyclomatic complexity. The number skyrockets for a complicated function that is not tested, so keeping CRAP low keeps code both simple and tested. He used to keep it at 4, relaxed to 6, and now thinks 12 is probably fine โ€” a deliberate loosening as the agents improve.

The third is mutation testing. A coverage tool only tells you what was executed, not what was tested. Mutation testing modifies the source and expects every modification to fail a test; any change that survives is a "surviving mutant," proof of a gap in the tests. He has the agents kill the surviving mutants, within limits, which makes the coverage numbers genuinely meaningful.

5. Aligning Tests With Intent

โ–ถ 8:06

The host's worry: agents are trained to reach a goal, so they can write tests that pass without matching the actual intent. Uncle Bob has two answers. The first is that he is the final arbiter โ€” a very tight loop where he forces changes, executes the code, plays with the system, and tells it to fix whatever it misunderstood. That works for applications where you can move through the steps quickly on screen.

The second answer is Gherkin. But a subtlety emerges here: agents can't really see the screen, so they ignore a lot of human-factors detail that he then has to go back and fix by hand. That gap โ€” between what passes a test and what a human actually wanted โ€” is exactly why the second constraint has to be authored by humans.

6. Gherkin: Specs a Human Must Author

โ–ถ 9:38

For elaborate systems with a lot of behavioral specification, he likes Gherkin โ€” the behavior-driven-development language โ€” with agents building the interpretation tools that parse and execute it. That gives a second set of tests authored, or at least reviewed, by humans. But he admits mixed luck, and a lesson: the farther away a human is from the Gherkin, the weirder the Gherkin gets.

When he had agents writing and testing the Gherkin themselves, it eventually became "an almost nonsensical description of the system." The prose reads plausibly โ€” you nod along โ€” but in the end it does not mean anything. He backed away almost entirely. His current position, which he half-jokes will change next week: humans need to author the Gherkin, period, in real human language, not agent language that only appears human.

7. Relaxing Constraints as Agents Improve

โ–ถ 12:52

The host pushes on why the thresholds are loosening: these are objective metrics, so relaxing them is essentially relaxing readability requirements. Uncle Bob's reasoning is about cognitive load. His old rule of "four" (very small functions) existed because humans have a short-term memory limit of five to seven items. Agents do not have that constraint โ€” they can hold several dozen facts at once and manipulate many simultaneously.

So the metric relaxes not because the code is worse, but because the reason for the metric has changed. The way he now measures the right threshold is empirical: watch whether the agent gets confused โ€” it starts breaking tasks, breaking other things, doing exactly what a task-saturated human would do.

8. Code Written for Agents, Not Humans

โ–ถ 13:54

The host names what has just been declared: code that is now only meant to be read by agents, not humans. Uncle Bob agrees โ€” some complexity is acceptable because agents can understand it even when humans cannot. It is a clean break with the premise of Clean Code, which was written for human readers.

This is the pivot point of the conversation. If code is written for agents, then the entire review process โ€” built around humans reading each other's code โ€” has to be re-examined, which is where the interview turns next.

9. The Future of Teams and Code Review

โ–ถ 15:10

His "wild conjecture" about teams: a single human can now do the work of a team. Instead of a tight-knit group of five or six developers constantly communicating, he predicts isolated individuals producing much more code, communicating only at a narrow interface boundary. The amount of collaboration decreases even as the amount of functionality increases. He is careful to flag it as pure conjecture that could go the absolute opposite direction.

The host disagrees, and the exchange is worth holding onto: his code-review talks are about building mental models and knowledge sharing, and he argues that architects and engineers โ€” the Jeff Deans of the world โ€” are made by learning practices, understanding trade-offs, and making decisions in collaboration. Even when you are not writing the code, that is still the most exciting part of software engineering. The two positions are not fully reconciled; they are the tension that will define the next few years.

10. The Dynamic UML Review Surface

โ–ถ 17:55

Uncle Bob then shares his screen โ€” a first for the podcast โ€” to show the tool he actually uses. It breaks a project down into a UML-like diagram, but one that is dynamic and coupled to an agent. He can click, expand, collapse, and play "what if" games. Problems surface as red: CRAP too high, mutation rate too low, red arrows for misaligned dependencies. He can drill from the high-level structure down to the code level when he wants โ€” though usually he does not need to.

This is the old dream of UML, inverted. UML was a static thing you did before writing code; this is a dynamic thing you do after the code is produced, to inspect what the agents have done "and undo the horror that they have created." He can point at an arrangement he dislikes, ask the attached agent to move things into different layers, and get back a proposed structure โ€” which the agent can then turn into real source code if he approves.

11. Disciplines Every Team Needs

โ–ถ 23:05

Asked what teams need to kill or reimagine code review, his first answer is a set of disciplines: rules the team agrees to follow until they meet and decide to break one. Those rules are concrete โ€” unit tests, driving coverage as high as possible, keeping CRAP below a threshold, mutation-testing within limits โ€” plus a verification step that the work was actually done correctly. For a big enterprise system, Gherkin might be part of it, and so might a human review in a GUI cycle.

The thing to avoid, in any setup, is the "vibe coder" โ€” someone who prompts his way to something that appears to work. You want to make sure it actually does work, which is precisely what the disciplines enforce.

12. Why Harness Engineering Is Overrated

โ–ถ 25:25

For several months he built an "austere" harness โ€” one agent for specification, one for coding, one for review, one for architecture, arranged in an assembly line handing off to each other. He spent most of his time keeping the hand-offs from leaking information, because "if they leak anything into each other, they all become the same agent after a while." He got it working and was genuinely proud โ€” little cards moving across swim lanes on screen.

Then he realized it was "inefficient as hell." The giveaway: he never used the harness to create the harness. The whole time, he was driving a single agent, and that agent kept getting better. The breakthrough came about ten days earlier โ€” a 45-minute discussion with Grok, proposing ideas and getting counter-proposals back, the kind of conversation he could only have with an extremely seasoned senior engineer, which would throw darts at his proposals and be right every time. His conclusion: "If I am willing to invest that much trust in a single agent, why am I putting it into a harness that makes it act like a menial slave?" He had the wrong model in his head โ€” a point echoed, he notes, by Steve Yegge's Gastown also going away. Are software factories doomed? He no longer knows what a software factory is.

13. The AI Exponential Curve and Its Limits

โ–ถ 28:30

He does not know where the single-agent approach stops, because he does not know how long the current curve lasts. He lived on Moore's law from the 1960s on โ€” exponentially shrinking electronics, bounded eventually by atoms and the speed of light. The AI curve is different: it is driven by growing power and growing real estate, which has a "very severe physical limitation." That, he half-jokes, is why Elon Musk is going to orbit.

The counterpoint, offered by the host, is that we already have a proof of efficiency: the instance in our skulls โ€” two pounds, twenty watts. Maybe once models hit a certain equilibrium we start optimizing again, and quantum computing might eventually change the equation. Uncle Bob is agnostic but clearly thinks the physical limits are real and near.

14. Responsibility and Ownership

โ–ถ 29:58

On responsibility, his answer is sharp: agents are not responsible for anything โ€” they are entirely irresponsible. Only humans can be responsible, now and probably for the next several years. Developers have to take their skill and experience, apply it to what the agents produced, and sign off โ€” with all the repercussions that implies. On a team, that means doing the UML work, the Gherkin work, the CRAP and coverage and mutation work, and making sure the system is deployable and rational โ€” without devolving all the way down to reviewing the code.

His reasoning for skipping code review is two-fold: it is not very useful anymore, and it slows everything down "like crazy." The new division of labor is clean: leave the syntax to the agents, hold everything above it near and dear, and remain responsible.

15. Rewriting Clean Code for the AI Era

โ–ถ 32:26

He wrote the second edition of Clean Code just before he started the AI work โ€” right at the moment he thought the AI stuff was not going anywhere. What went into the book is largely the same as before, with more examples, because all the same principles still apply even when agents write the code: you still want clean code, good partitioning, good dependencies, good names, properly managed comments.

What might change are the thresholds and the boundaries โ€” the part he is actively exploring, like the loosening of CRAP. But he will not change any fundamental principle. The closing exchange lands the humility of it: humans were never perfect at this either, so if agents become as good as humans were, that is enough โ€” and on the average, he suspects they might already be a little better. If the exponential curve continues, at some point the agents beat the chess master.

Key Takeaways

  1. Humans are too slow for code. Uncle Bob's bet: back off to high-level supervision and inspect architecture, not lines.
  2. Three constraints survive the AI transition: unit tests as double-entry bookkeeping, CRAP (coverage + complexity), and mutation testing to prove coverage is real.
  3. Humans must author specs. The farther a human is from the Gherkin, the weirder it gets โ€” agent-written specs read plausibly but mean nothing.
  4. Constraints relax because the reason changed. Small functions existed for human short-term memory; agents have no such limit, so the threshold loosens and you watch for confusion instead.
  5. Code is now written for agents, not humans โ€” which breaks the premise that review means humans reading code.
  6. Teams may collapse to isolated individuals producing more code through narrower interfaces โ€” flagged as pure conjecture.
  7. The review surface is a dynamic UML diagram coupled to an agent, built after the code, to inspect structure and dependency alignment.
  8. Harness engineering is overrated โ€” he abandoned a multi-agent assembly line because he never used it to build itself, and a single trusted agent beat it.
  9. Only humans are responsible. Sign off above the code level; leave syntax to the agents.
  10. Clean Code principles still apply โ€” only the thresholds and boundaries change.

Timestamp Index

  • 0:00 โ€” Introduction and the thesis
  • 1:48 โ€” How Uncle Bob codes today
  • 3:12 โ€” What changed his mind about AI coding
  • 4:52 โ€” Unit tests, CRAP, and mutation testing
  • 8:06 โ€” Aligning tests with intent
  • 9:38 โ€” Gherkin and human-authored specs
  • 12:52 โ€” Relaxing constraints as agents improve
  • 13:54 โ€” Code written for agents, not humans
  • 15:10 โ€” The future of code review and teams
  • 17:55 โ€” The dynamic UML review tool
  • 23:05 โ€” Disciplines every team needs
  • 25:25 โ€” Why harness engineering is overrated
  • 28:30 โ€” The exponential curve and its limits
  • 29:58 โ€” Responsibility and ownership
  • 32:26 โ€” Rewriting Clean Code for the AI era
โ˜ฐ View all