Decision #797AcceptedTrack · AI in Product4 min read

AI Evals Move From Engineering to Product Teams: A Practical Framework

ProductTalk's Teresa Torres publishes a hands-on guide for product teams on building AI evals, with a three-step framework, four eval categories, and production numbers from her own products.

Context

  1. AI evals have been the 'it' skill for product teams for over a year, per ProductTalk's Teresa Torres

  2. Torres's framework has three steps: error analysis, picking an eval type, and running experiments

  3. Four common eval types exist: Golden Dataset, Code Assertion, LLM-as-a-Judge, and Customer Feedback

  4. Torres's Interview Coach baseline scored 15 of 105 suggested questions as leading, 3 as general, and 9 as already-answered

  5. The next article in the series will detail how 16 experiment variants were needed to fix a single customer complaint

AI evals have been the "it" skill for product teams for over a year, yet most product teams still only have a vague idea of what evals are or how to build them. In a hands-on guide published this week, ProductTalk founder Teresa Torres argues that AI evaluations—the methods teams use to measure whether an AI product or workflow is performing well—are the missing feedback loop for any team shipping LLM-powered features.

The guide matters because most writing on this topic targets engineers or stays too abstract for practicing PMs. Torres, author of Continuous Discovery Habits, makes the case directly: product teams that use AI to write PRDs, synthesize customer feedback, or analyze behavioral data can measure that work with evals.

Why do traditional tests fail with LLMs?

Unit tests assume the same input always produces the same output. LLMs do not work that way. Torres describes the challenge as twofold: probabilistic output and semantic tasks.

Probabilistic means the model can return 5 for "2+3" one moment and 4 the next. Semantic means many tasks—writing a joke, summarizing an interview—have more than one right answer. Sometimes better or worse. The job of an eval is to count how often the model gets it right, not to assert it always does.

Torres is blunt about one trap: don't outsource the definition of "correct." Big eval tool providers ship built-in checks for conciseness or helpfulness, but correctness is context dependent. A student learning astrophysics needs verbose responses; a professor does not.

"Don't let a vendor define correctness for your product," Torres writes. "This is the product team's job."

What are the three steps?

Torres's framework breaks into three moves:

  • Step 1: Error analysis. Manually review a range of inputs and log the types of mistakes the model makes. For a personal workflow, one or two deep passes work. For production, she recommends hundreds of cases.
  • Step 2: Pick an eval type for each error. Four options exist, each with different cost and judgment needs.
  • Step 3: Experiment. Collect a baseline, run a variant, compare scores.

She ran Step 1 on 17 Lovable interview transcripts she wanted to turn into blog vignettes. After one in-depth review, she identified the recurring errors: hallucinated job titles and fabricated quotes that merged two real statements into a new meaning. That review produced two eval needs: a fact-checker and a hallucination guard.

What are the four types of evals?

Torres breaks evals into four categories, with clear tradeoffs:

  • Golden Dataset. A set of inputs paired with ideal outputs. Best for small inputs, small outputs, one correct answer—classification, factual recall, routing. Fails when inputs or outputs are large, or when many answers qualify.
  • Code Assertion. Deterministic checks against the LLM output. Fast and cheap. Torres used one to check every quote in her ChatGPT-generated story appeared verbatim in the transcript—string match or no match.
  • LLM-as-a-Judge. A second model scores the first. Best when no string exists for evaluation and judgment is required. Expensive and slow. Works only when the judge gets a simpler task than the original model, returns a binary verdict, and gets aligned with human judgment.
  • Customer Feedback. Direct user ratings or inferred behavior. The ultimate judge for customer-facing products, but often tough to make actionable beyond thumbs up or down.

Torres mixes them. In her Interview Coach, a code assertion flags red flag words like "typically" and "usually" that indicate a general question. An LLM-as-a-Judge evaluates whether a suggested question is leading or already answered—judgments no string search can catch.

What does the math look like in production?

Torres's Interview Coach produces concrete numbers. After running evals on a batch of suggestions:

  • Leading Question: 15 out of 105 suggested questions
  • General Question: 3 out of 105 suggested questions
  • Already-Answered: 9 out of 105 suggested questions

For her Lovable blog work, the LLM-as-a-Judge fact-checker returned 15 out of 17 facts grounded in the transcript. The code assertion verified 3 out of 4 quotes verbatim.

What infrastructure does the loop need?

Torres built a Python test harness that runs every input through the LLM service, transforms outputs into the format each eval expects, runs each eval, and scores the full run. Each eval lives as a module with a run function the harness calls without understanding internals. For personal workflows, she uses Claude Code sub-agents to drive the same loop without automation.

Where does this leave product practice?

The shift from engineering-owned quality to product-owned correctness is the thread Torres pulls hardest. When error analysis draws on customer interviews and outcome metrics, evals stop being a technical chore and become another discovery habit, sitting beside assumption testing and story mapping. The next article in the series promises a deeper dive into LLM-as-a-Judge alignment and a story about how 16 experiment variants were needed to fix a single customer complaint. That kind of iteration candor is what teams will need as AI products move from prototype to production.

via vistaly.com (Original)

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering marketplaces and e-commerce at Roadmap File.

22 articles