Decision #414AcceptedTrack · AI in Product4 min read

AI evals become the PM's loop: 20–50 failures is all you need to start

Amplitude's Darshil Gandhi argues 20–50 real failures are enough to start an AI eval program — and explains why traces, not dashboards, are now the PM's source of truth for agent quality.

AI evals for product managers: A beginner’s guide to getting started
AI evals for product managers: A beginner’s guide to getting startedAI-generated

Context

  1. 20–50 real failures are the working minimum to produce a useful first AI eval suite, per Amplitude's Darshil Gandhi.

  2. 34% of inventory-lookup traces ended in a user follow-up, which became two new evals in the guide's example.

  3. Online LLM-as-a-judge evals typically sample 5–15% of production traffic to control token cost.

  4. An online eval caught a helpfulness pass-rate drop from 81% to 64% within six hours of a Sunday-night deploy.

  5. Amplitude Agent Analytics treats agent interactions as events in the same stream as retention, conversion, and adoption.

20 to 50 real failures is the working minimum for a useful first eval suite, according to Amplitude's Darshil Gandhi.

Gandhi, Director of Product Marketing at Amplitude and a former solutions engineering principal, argues that product analytics built for clicks and form submits cannot see inside an AI agent. Users now type their intent directly into a chat box, and the agent replies in nondeterministic ways. "Your dashboards can't see inside your agent. Evals can," Gandhi writes in a beginner's guide published this week.

The pitch is blunt: "If you treat evals as a chore, you'll ship features you can't measure, debug, or defend. If you treat eval design as part of your job, you'll own the loop between agent quality and business outcomes." For product managers shipping agents in 2026, fluency in trace analysis and eval design is the new craft line.

Why does the trace replace the dashboard?

A trace is the complete record of one agent interaction: user input, every tool call, retrieved context, the model's responses, latency, cost, and any feedback signal. Each step is a span, and one user session can contain many traces. Gandhi calls the trace "the new source of truth for what users experience" for PM teams working on agents.

Two failure patterns from the guide show why:

  • Duplicate billing rows. A support agent returns the wrong answer to "Why was I charged twice?" The trace shows the billing tool returned a duplicate row the agent didn't catch — a retrieval bug session recording would never surface.
  • Inventory without restock dates. Over a 7-day window, 34% of traces where the inventory lookup tool fired ended in a user follow-up. The tool returned stock counts but not restock dates, so the agent gave a technically correct but incomplete answer.

In both cases, the trace analysis produces two outputs: a fix that ships, and an eval that prevents regression.

What kind of eval do you actually need?

Gandhi splits evals along two axes: scoring method (code-based vs. LLM-as-a-judge) and where they run (offline in development vs. online against live traffic).

  • Code-based evals score with deterministic logic — regex, JSON schema, expected tool calls. Fast, cheap, run in milliseconds. Best for verifiable properties: valid JSON, correct tool call, required legal disclaimer.
  • LLM-as-a-judge (LLMaaJ) evals use a second model against a natural language rubric. Each grade costs tokens and takes seconds. Best for subjective quality: helpfulness, tone, groundedness.
  • Offline evals run in development against a fixed dataset and act as a CI gate. They catch regressions before merge but only test what you anticipated.
  • Online evals score live production traffic continuously and surface new failure modes. Most teams sample 5–15% of traffic for the LLM judge to control cost.

The failure mode to watch: a judge that passes everything. "A judge that gives every answer a passing score is worse than no judge at all because it creates false confidence," Gandhi writes. Mature teams sample judge decisions, compare them to a human reviewer, and adjust the rubric until the two align. Anthropic's engineering guide on evals is cited as a reference.

Do you actually catch regressions in CI?

Yes, and the guide includes a working example. After shipping a new system prompt, an offline eval caught the agent failing 4 of 12 edge cases on ambiguous date ranges — cases added to the dataset three months earlier after a wave of complaints. The eval blocked the PR.

Online evals catch what offline sets can't. In one case, a helpfulness pass rate dropped from 81% to 64% on Monday morning across account-migration queries. A backend change deployed Sunday night had altered the context the agent received. The online eval surfaced the regression in six hours, and the failure pattern was fed back into the offline dataset.

What does this change for the PM job?

Three shifts Gandhi calls out:

  • Eval ownership is shared. PMs define what counts as success and contribute cases from real user behavior. Engineers build the infrastructure and wire it into CI. Evals become a shared artifact, edited like product specs.
  • A/B tests without evals optimize for shallow proxies. Response length, latency, and thumbs-up rates rise without proving quality improved. Pairing experiments with evals lets teams measure both quality and business outcomes per variant.
  • Eval scores have to join engagement data. A high pass rate doesn't tell you whether successful interactions drive retention, whether failures concentrate in high-value segments, or whether your most expensive queries are your lowest-converting ones. Amplitude Agent Analytics, launched earlier this year, treats agent interactions as events in the same product stream as retention, conversion, and adoption.

The next eighteen months will tell whether eval design settles into a standard PM skill the way SQL did, or stays a specialist practice siloed in ML teams — and the answer likely depends on whether vendors keep making traces as cheap to instrument as page views.

via anthropic.com (Original)

More from Marcus Bennett

Marcus Bennett

Show full bio

Market editor covering marketplaces and e-commerce at Roadmap File.

21 articles