Loop Engineering Β· Claude Code

Loop Engineering: Getting Claude Code to Build on Its Own

AI LABS 12:51 Karpathy's loop Auto loop
Loop engineering with Claude Code

1. The Loop That Builds On Its Own

β–Ά 0:00

People have been getting AI agents to build real things on their own for hours at a time, and the reason it works is the structured workflow behind each loop. Loop engineering is that workflow, built into a repeatable form: an agent makes a change, a separate check scores the result, and the agent keeps the change only if the score improves.

One of the best-known versions is Karpathy's loop. AI LABS rebuilt it on their own apps, watched it hold up β€” and found one issue the loop cannot see on its own, then fixed it by adding a second loop that improves the first. What follows is what the Karpathy loop is, how to set it up, and the change that fixes its blind spot.

2. Karpathy's AutoResearch Method

β–Ά 0:51

Andrej Karpathy's AutoResearch (97,158 stars on GitHub, Python) put an AI agent to work training a model in a loop. The agent was allowed to change exactly one file β€” the one that trains the model. Each round was one experiment: change that file, train for a few minutes, then read a score from a second file. If the score improved, the change stayed; if it was the same or worse, the change was undone.

The critical constraint: the agent could never touch the scoring file. If it could, it would simply make the score easier to beat instead of making the model better. A third file, program.md, was a plain-English instruction file Karpathy wrote to tell the agent how to run each round. Left running for two days, the agent ran 700 experiments and found 20 changes that made the model train faster. The CEO of Shopify ran the same loop on one of his models overnight: 37 experiments, and the model performed 19% better by morning.

3. When to Use a Loop

β–Ά 2:32

Loops are not free β€” set one up on the wrong task and you burn tokens for nothing. A loop is worth building only when the task meets four conditions:

  • You repeat it often. A loop takes time to build and only pays back through repetition. For a one-off job, a single good prompt is enough.
  • Your usage limit can handle the token cost. The agent re-reads the project and tries a new fix every round, burning tokens even on failing rounds β€” a long loop on a cheap plan hits its limit before the work finishes.
  • The work is checkable without a human. There must be a clear score β€” in an app, small pieces of code that exercise a feature and confirm it works.
  • The agent can run what it built and see what breaks, so it knows what to fix next round.

Two guardrails from AI LABS' own practice: they only loop over tasks with a measurable score, and they never let a full app get built in one loop and declared done β€” one feature at a time, or a simple version the checks can verify.

4. The Setup: Skills and Agents

β–Ά 4:02

Karpathy's loop, translated into a reusable Claude Code workflow, is a set of skills and agents β€” each one you can ask Claude to build for you:

  • Project context skill β€” the app's memory bank: what it does, which pages it has, its conventions, and what to avoid. It grows as the app grows. A skill beats a plain file here because only the short description stays in the context window; the full details load only when needed.
  • Build skill β€” drives the loop, carrying the instructions for running it on each feature.
  • Write checks skill β€” writes the checks before any feature is built, so the agent verifies against something other than its own judgment.
  • Approve checks program β€” after you review the checks (and add or remove any), this moves them into a locked folder, and a Claude Code settings rule blocks the agent from editing it.
  • Feature builder agent β€” a fresh agent per feature, so each feature's context stays separate.
  • Results file β€” records every round, kept for later use.

The build skill commits the approved checks (a second layer of protection against tampering) and writes all the loop's fixed rules and "how to work" into program.md, in the same format Karpathy used.

5. Running the Loop: Pickup Ordering

β–Ά 7:37

On a half-built restaurant website, they prompted Claude to build a feature letting guests order food for pickup. Claude added the feature to the tracking file, wrote the online-ordering rules, and listed its checks β€” ten of them. They asked for the guest's email to be verified too, making eleven. After approval, the feature builder agent built the feature in about six minutes.

The report showed all 11 checks passed on the first round. But the loop caught something the score did not: the checks only tested the ordering rules, not the order form guests actually use. So the loop built the order form and connected it to the rules before marking the feature done. That is the loop's hidden value β€” the checks gate quality, but the loop still has to wire the feature into the app.

6. The Habit Problem

β–Ά 8:38

Every loop has a blind spot: the agent has habits. If a particular approach fails, the agent corrects course within that run β€” but it does not remember the lesson for the next one. Because every feature starts with a fresh feature builder agent and the same instruction files, a mistake made in one round simply carries over to the next.

The fix is a second loop that reads how the first loop ran and changes how the first loop works. In Karpathy's method a human writes program.md; in AI LABS' version, a skill called auto loop rewrites the "how to work" part of program.md automatically.

7. The Auto Loop: A Loop on the Loop

β–Ά 9:56

The auto loop reads the results file β€” which stores every round, whether it was kept or undone, what the agent tried, and which checks failed β€” and looks for problems that keep recurring: the same mistake, or the same gap showing up in two features. It writes down each habit, then rewrites the "how to work" section so the next feature starts from the method that actually worked. It can never edit the checks, or it could make them easier to pass instead of fixing the habits.

Tested on a project-management app, it caught two real problems. A shared database passed all ten of its checks while the app never saved anything to it β€” so the auto loop added the habit "connect each feature to the app in the same round its checks pass." A mentions feature passed its checks while an older part of the app still read the @-sign the old way β€” so it added "find every place the app already does a job and make each one follow the new rules." Both habits stuck, and the next features came with their screens already connected. In a later shared-projects feature the builder took three rounds β€” round one broke mentions (undone), round two split three people's share into 33/33/33, totaling 99 (check failed), and round three fixed it to pass all ten. The whole method is model-agnostic: it runs the same on Claude Code regardless of model, and the ideas carry over to ChatGPT, Codex, or any other agent.

Key Takeaways

  1. Loop engineering = agent changes, a separate check scores, keep only if better. The scoring check must be untouchable, or the agent games it.
  2. Karpathy's AutoResearch ran 700 experiments in two days for 20 improvements by changing one file and reading a score from another it could not edit.
  3. Four preconditions: a repeatable task, a usage limit that fits the token cost, a human-free score, and an agent that can run what it builds.
  4. Write checks before the feature, lock them in a folder the agent cannot edit, and commit them β€” three layers against self-grading.
  5. One feature per fresh agent, with a results file recording every round.
  6. The blind spot is habits: a fresh agent with the same instructions repeats the same mistakes β€” so run a second loop that rewrites program.md's "how to work" from the results.
  7. The auto loop caught what scores missed: a database that passed 10 checks but saved nothing, and a mentions feature that clashed with legacy code.

Timestamp Index

  • 0:00 β€” Intro
  • 0:51 β€” Karpathy's AutoResearch method
  • 2:32 β€” When to use loops
  • 4:02 β€” The setup: skills and agents
  • 7:37 β€” Running the loop: pickup ordering
  • 8:38 β€” The habit problem
  • 9:56 β€” Auto loop in action
☰ View all