Decision #325AcceptedTrack · Discovery & Research4 min read
4 Evals and 16 Experiment Variants to Fix One Customer Complaint
One Vistaly customer complaint about a flat tree branch took 3 weeks, 4 new evals and 16 experiment variants — a case study in calibrating LLM judges and knowing when prompts stop working.

Context
Fixing one customer complaint about unstructured branches took 3 weeks, 4 new evals and 16 experiment variants at Vistaly
The LLM-as-a-Judge eval initially scored 100% recall but only 43.75% specificity; after 7 iterations specificity reached just 60%
Final variant 16 delivered 78% fewer parent-restaters, 29% fewer poorly framed key moments, and a 65% increase in mid-level parents
A single beta customer's comment — "I wish I could click on this card and be like, 'Clean this up'" — took three weeks, four new AI evals and 16 experiment variants to properly fix. That ratio is the real story of building AI-generated opportunity solution trees at Vistaly, where Teresa Torres (author of Continuous Discovery Habits) is partnering on AI-generated interview snapshots and trees.
The customer was reviewing her AI-generated opportunity solution tree and spotted a branch full of flat opportunities — no hierarchy, no sub-groupings. The tempting fix was a "clean up" button. Torres wanted to fix the agent itself: why didn't it structure the branch?
Start with evals, not fixes
Before touching prompts, she measured how often the error occurred. She built two evals. The first, tree-shape, is a deterministic code-assertion eval that counts nodes, children per parent, mid-level parents and depth — letting her apply rules of thumb like "three to seven key moments" and "trees should get deeper, not infinitely wider, as interviews accumulate." The second, missed-groupings, is an LLM-as-a-Judge eval: a second LLM call reviews each parent-child cluster and flags missing sub-groupings.
Here's the failure mode most teams hit: you can't trust a judge until you calibrate it against your own labels. Torres built a calibration set of manually labeled tree nodes and scored the judge on specificity (accuracy when there's no error) and recall (accuracy when there is one). Recall was perfect at 100%. Specificity was 43.75%. Seven iterations of simplified instructions, added examples and a smarter model got specificity only to 60%. The judge kept inventing missed groupings that didn't exist.
The culprit: upstream errors. A child that simply restated its parent confused the judge into proposing redundant sub-groups — three nodes saying the same thing in different words.
Upstream errors first — but with a catch
Hamel Husain's advice was to fix upstream errors first. Torres resisted — customers weren't complaining about poorly framed customer key moments (CKMs), only about flat branches. But the judge couldn't do its job with those errors present, so she built two more evals: one for poorly framed CKMs, one for parent-restaters (children that add no new information). Her gating strategy: only run the missed-groupings judge on nodes clean of the other two errors.
Then another surprise. Her newly calibrated judges, run against real production trees, reported >90% error rates on both categories — saturated. The calibration sets hadn't matched production data. She rebuilt them, eventually reaching 100% recall / 93% specificity on parent-restaters and 90% / 100% on poorly-framed-CKMs.
The whack-a-mole problem
Prompt changes exposed a structural tension: the more constraints she added on framing parents well, the more reluctant Sonnet 4.6 became to add parents at all. Upgrading to Sonnet 5 (variant seven) didn't save her — it showed zero parent-restater errors because it barely added nodes, creating one mid-level parent where 4.6 created six or eleven.
By variant thirteen, after a full rewrite of her sub-grouping rules, parent-restaters had dropped roughly 70% and poorly framed CKMs about 54% — but missed groupings barely moved.
From eval to guardrail
The breakthrough came from orchestration, not prompts. Her agent already generated trees in steps mirroring how she teaches the method (key moments, then grouping, then structuring), with an audit loop that sends errors back for self-correction. She taught the deterministic auditor to detect parents with too many children using tree-shape, then shipped prompts optimized for well-framed parents — accepting that the agent would miss groupings on first pass, because the auditor would catch and route them back. Two more variants fixed the new overcorrection (too many single-child parents).
Final results for variant 16: 78% reduction in parent-restaters, 29% reduction in poorly framed CKMs, and a 65% increase in mid-level parents.
The takeaways for product teams building AI features: calibrate judges against production-representative data, measure every intertwined error category before fixing any of them, and expect prompt engineering alone to plateau — reliable AI products need context engineering, orchestration and guardrails working together. As AI features move from demo to production, this unglamorous eval-and-iterate grind — not model upgrades — is where product craft now lives.
via vistaly.com (Original)
More from Rebecca Stone
Show full bio
Correspondent covering marketplaces and e-commerce at Roadmap File.
2 articles