Decision #699AcceptedTrack · AI in Product4 min read

Verification loops ship code. PMs own the loop that ships impact.

Boris Cherny stopped prompting Claude Code this summer and started writing loops. The work-grade-rework cycle is now called a verification loop. PMs should own the grader and the post-ship optimize loop.

Context

  1. Boris Cherny, who built Anthropic's Claude Code, said this summer: "I write loops and the loops do the work."

  2. Both Claude Code and OpenAI's Codex now ship a /goal command that runs an agent until a specific condition is met.

  3. A verification loop has five parts: context, goal, agent, output, and grader — with the grader as the only PM-owned component.

  4. Five common grader checks are deterministic tests, LLM-as-a-judge, browser-based visual verification, simulated environments, and offline experiments.

  5. Amplitude built Wave to answer "did the change move the metric enough?" once code ships, closing the loop the build-side verification cycle can't close.

"I write loops and the loops do the work. My job is to write loops." That's Boris Cherny, the engineer behind Anthropic's Claude Code, explaining this summer how he stopped prompting and started orchestrating. Addy Osmani framed the consequence plainly: software has dreamed of becoming a "repeatable and instrumentable production process" for half a century. "A software factory is many harnessed loops running at once," he said."

When the human stops prompting, something else has to decide what work gets done and when it ends. Both Claude Code and OpenAI's Codex now ship a /goal command that runs an agent until a specific condition is met. That work-grade-rework-regrade cycle has a name: a verification loop.

What is a verification loop?

An agent uses context to meet a goal, alternating between building and evaluating. Each failure feeds into the next attempt until the output hits a success criterion. In a build loop, an agent writes code, a grader scores it, and the verdict decides whether the change ships or iterates. The structure is adapted from Andrew Ng's three-loop framework for agentic development.

A basic verification loop has five parts:

  • Context: codebase, error traces, product data you feed in
  • Goal: the success criteria you set
  • Agent: the AI doing the work
  • Output: typically a pull request
  • Grader: a separate evaluator

Four of the five are mechanical. The grader is where decisions live, and where PMs should be writing.

What does the grader actually check?

Take an AI support agent. Its grader scores every reply against a rubric: did it resolve the issue, did it stay factually accurate, did it sound like the brand, did it escalate when it should have. Engineers can wire the LLM judge. Only a PM can decide that a refund-policy violation is an automatic fail, or that a correct answer with a cold tone isn't good enough.

Five grader checks are common today:

  • Deterministic tests: CI pass/fail, schema validation, risk classifiers. Fast, objective, cheap. They answer "did this break something," not "is this any good."
  • LLM-as-a-judge: scores readability, style, fit to goal. Catches obvious failures but adds noise. Products pass the test and miss the goal.
  • Browser-based visual verification: an agent opens a browser and confirms rendering. Codex and Cursor both ship this. Catches the case where code is correct but invisible to the user.
  • Simulated environments and synthetic users: model user behavior pre-ship. Powerful, expensive, bounded by the quality of the world model.
  • Offline experiments: run configurations against past data, score against a rubric, pick a winner. Useful backward-looking; a weak predictor of live behavior.

When does the build loop stop working?

Verification loops confirm that code runs. They can't confirm that a change moves a metric. Picture a checkout conversion agent: tests pass, the LLM judge approves the copy, the browser confirms the button renders, the simulation returns a guess. All four checks fire green. None of them can tell you whether more users finished checkout.

That's the limit every PM eventually hits. The build loop is competent on "does it work." It is silent on "did it move the metric."

What does the optimize loop need?

The same loop concept applies after shipping, with every part swapped:

  • Context: live product data
  • Goal: a metric, not a spec
  • Output: the shipped user experience
  • Grader: how users actually react, measured through a controlled experiment or holdout

Product analytics shows what happened after shipping. An experiment shows what caused it. In the optimize loop, the only honest grader is real user behavior captured through controlled experiments or holdouts — not the in-session model that handles the build loop.

Today, a general-purpose coding agent can't verify live product data against user-experience goals in-session. Even a skilled PM couldn't wire that up this quarter. It requires a long-horizon agent that carries product rubrics, which is what Amplitude built Wave to be. Verification loops answer "does this change work"; Wave answers "did the change move the metric enough."

What should PMs do on Monday?

Two moves change how you show up next sprint:

  • Design the build-loop grader. A vague spec makes the loop converge confidently on the wrong thing, faster than any human can review. PMs write the rubric; engineers wire the judge.
  • Own the post-ship optimize loop. If users don't react positively, the build work doesn't matter. Behavioral metrics have to be defined before a change can be called successful.

Engineers are getting very good at automating the build loop. If PMs are still running the optimize loop by hand, the whole process slows down or breaks. As the cost of building approaches zero, shipping working code stops being a differentiator. Improving the user experience and moving business metrics becomes one — and that's the loop PMs will own going forward.

via x.com (Original)

More from Daniel Okafor

Daniel Okafor

Show full bio

Staff writer covering media and advertising at Roadmap File.

18 articles