Decision #733AcceptedTrack · Discovery & Research4 min read
Vistaly's 2.5-Month Rewrite Replaced Three Years of Product Code
Vistaly rebuilt its product in 2.5 months after three years of V1. The agentic rewrite forced four evals, 16 experiment variants, and an orchestration repair loop.

Context
Vistaly rebuilt its product in 2.5 months after spending 3 years on V1.
Four new evals and 16 experiment variants tested a seesaw between missing subgroupings and badly framed parents.
Vistaly shut off V1 signups during the rewrite to protect the V2 launch.
A 40-point system generates synthetic interview transcripts for eval data.
European data residency pushed inference onto AWS Bedrock with EU hosting for SOC 2 and GDPR compliance.
Vistaly rebuilt its entire product in two and a half months—after three years building the first version. The new platform, V2, takes uploaded customer interviews and uses an agentic workflow to draft an opportunity solution tree, the canonical artifact Teresa Torres popularized for continuous discovery. The team shut off V1 signups to protect the rewrite.
"We rebuilt nearly all of V1's functionality in two and a half months," Torres, host of Just Now Possible and a partner building Vistaly's AI synthesis services, said on the show. The compressed timeline was possible because the team already knew what the product needed to do. The harder question was whether the same workflow, routed through AI agents, would still hold up.
Why the chat interface failed
Vistaly's first AI attempt was a chat-based synthesis tool that walked users through interview insights one at a time. Users found it too slow, even when each step was "still faster than before." They abandoned it and stopped layering AI features onto V1, choosing a ground-up rewrite instead.
V2 works in three stages. Users upload three interviews. The system generates an interview snapshot for each. It then synthesizes a first-draft opportunity solution tree from those snapshots. Users edit from there.
The house-of-cards problem
The bottleneck isn't latency. Layered analysis compounds errors. If the interview snapshot is wrong, every layer above it inherits the failure. Co-founder CP Dehli named it the "house of cards" problem—one badly framed opportunity corrupts everything above it.
That insight reset the quality bar. The competitive threat isn't faster good synthesis; it's shallow synthesis that never really happened. "Good enough" stops meaning acceptable latency and starts meaning the snapshot layer is correct, or the rest is decoration.
When prompts run out
A single user complaint forced the eval work. "I'll have to come back and clean up this branch," one customer told Vistaly. The team needed to find why.
They built four new evals and ran 16 experiment variants against a seesaw between two opposing failure modes: missing subgroupings and badly framed parents. Improving one regressed the other. An LLM-as-a-judge couldn't be calibrated to catch the right failure cleanly. The fix came from orchestration, not prompting—moving the check into an agentic repair loop that detects the failure mode and reruns the relevant step. That eval graduated to a production guardrail. Prompt iteration alone had run out.
Before paying for any LLM judge, Vistaly runs a cheap code assertion: if a tree node has too many children, the synthesis has likely failed upstream. The deterministic check pre-filters obvious failures.
Change sets as moves in a game
Logging what the agent changed turned out to be harder than generating the change. The same input/output tree pair admits multiple valid change sets; only the semantically meaningful one makes sense to a user. Diffing at the end doesn't work.
The team reframed change sets as moves in a game with rules: merge, move, reframe. The agent now logs semantic moves as it goes. Stable IDs and data provenance—"basic computer science" Torres said she had swept under the rug in the prototype—became load-bearing infrastructure.
Infrastructure and inputs
Beta traffic surfaced bad inputs: sales demos, stakeholder meetings, and LLM-generated fake transcripts uploaded as "interviews." To test against realistic inputs, the team built a 40-point system for generating synthetic transcripts.
For deployment, European data residency pushed Vistaly onto AWS Bedrock with EU-hosted inference. SOC 2 and GDPR followed. Torres also flagged that model upgrades aren't drop-in: prompts are model- and version-specific, and Bedrock constrains tokens independently of the model itself.
External editing runs through the Model Context Protocol (MCP). One customer drove a full Jira epic backlog into the opportunity space via MCP.
What this signals for product teams
The hardest problem isn't output quality—it's helping users comprehend what changed. Vistaly's rebuild argues that agentic product tools need three things: repair loops in orchestration rather than better prompts, layered evals with deterministic pre-filters above LLM judges, and change sets designed as semantic moves, not end-state diffs. As more product teams rewire continuous-discovery and planning workflows around AI agents, those three patterns will likely separate prototypes that demo from products that ship.
via open.spotify.com (Original)
More from Nathan Brooks
Show full bio
Senior reporter covering consumer brands and retail at Roadmap File.
16 articles