Decision #943AcceptedTrack · AI in Product5 min read
Amplitude cut its MCP catalog from 96 tools — and retention climbed
Amplitude cut its MCP server tool count, merged cohort and chart flows, and watched 67% of external users return after one week. The rebuild traded endpoint-mirroring for usage-shaped tools.
Context
Amplitude's MCP server reached 96 tools by June before the redesign
A Sonnet user with a 200K context window burned 57% of it on connection alone
Anthropic's Opus 4.5 picks the correct tool only 88% of the time, per the company's own evals
Cohort tools fell from nine to one; chart tools fell from 16 to four after the rebuild
Post-rebuild retention: 67% returned after one week, 59% after four, 50% after eight, over a 12-week window
A customer running Sonnet with a 200K context window burned 57% of it just connecting to Amplitude's MCP server, before a single tool call.
By June, the server exposed 96 tools, most of them endpoint-shaped and most irrelevant to any given user. The team's response was a rebuild around usage, not around the API. AI engineer Chanaka Perera led the work and laid out the rationale in a public post.
Why did 96 tools become the wrong shape?
Each API endpoint had become a candidate tool. Customer requests prompted exposure. Niche needs from one account became context weight for everyone. Most users paid the token bill for tools they never touched.
Perera's instrumentation pointed to three structural issues:
- Most MCP clients load every tool definition upfront by default; tool search is opt-in and client-controlled.
- On Amazon Bedrock, tool search works only through InvokeModel, not the Converse API most teams use.
- Enterprise gateways cache the tool list once and rarely refresh it.
"Every tool costs tokens on every connection" became the working assumption.
Tool count also raises error rates. Anthropic's own evaluations show Opus 4.5 picking the correct tool only 88% of the time. Older models perform worse. Anthropic flags wrong tool selection as one of the most common failures, especially when names overlap.
What three rules shaped the rebuild?
Don't mirror the API. Endpoint-shaped tools forced agents to chain calls and pass IDs between them, a sequence they handle poorly. Perera's team collapsed groups into one tool per object. Cohort work went from nine tools to one, with eleven actions routed to existing handlers.
The merge surfaced a missing capability. There was no list operation because the API had no listing endpoint, so agents were calling the cohort lookup with invented IDs. The team added a new route rather than renaming an existing one.
Each action keeps its own permission level and audit event. Risk gating matters: setting it at the riskiest action breaks read-only access; setting it at the lowest exposes writes. Amplitude resolves access per action before the handler runs.
Merge across product boundaries when usage demands it. Search previously had one tool per product — charts, dashboards, notebooks, cohorts. Agents were calling two or three in sequence for ambiguous phrases like "find the retention analysis," because the prompt didn't specify which entity type.
The new search tool accepts a list of entity types and queries them in one call.
Chain internally, not at the agent layer. REST-shaped flows assume a caller can order calls correctly. Creating a chart in Amplitude requires three ordered steps: create a chart edit, save it, add to a dashboard. The dashboard rejects unsaved edits.
The old tool catalog mirrored those steps: query data, render chart, save edit. Of 18,703 users who queried, 20% rendered a chart and fewer than 5% saved one.
The team read the "Brief explanation of why you are calling this tool" rationale field that 75% of agents populate voluntarily. Stopped calls read like this: "save chart edit for dashboard analysis." The agent attempted to save but never finished the chain.
The replacement takes a definition and returns data in one call. Saving is now a button on the rendered chart. Chart tools fell from 16 to four.
Where does context still leak?
Tool descriptions load on every connection, so long instructions cost every user. Amplitude moved 37 of them into skills — markdown files loaded only when a task requires them. They live in amplitude/mcp-marketplace under MIT license, pulled in as a git submodule.
Skills serve two ways: as MCP resources for the roughly one in ten clients that read them, and through a regular tool that fetches by name for everyone else. The team logs which surface delivered each read.
For rarely used operations, the team hides a CLI behind one tool so a new endpoint works the day it ships. Permissions derive from the HTTP method, and unknown commands get write treatment by default.
How do MCP Apps fit in?
MCP Apps let a human approve destructive or multi-user actions before they run. A confirmed: true parameter doesn't work; the model will set it. Amplitude uses the same flow for entity sharing and space administration.
Apps also let the team add functionality without expanding context. A tool marked app-only is callable from the UI but invisible to the model, so it costs nothing on connection. Not every host supports MCP Apps yet, so a fallback remains necessary.
Which two metrics mattered during rollout?
Each change shipped behind a feature flag, one per product area. Internal orgs went first; customers followed in stages. CSMs briefed accounts with automations built on the old tools.
Two numbers watched during ramps:
- Error rate on tool calls
- Recovery rate: did the agent make a successful call within five minutes of a failed one?
Recovery was prioritized. A failing tool with a clear message is acceptable; a failing tool that leaves the agent guessing is not. Several merged tools added schema validation the originals lacked, so error rates climbed first. Recovery stayed high, meaning agents read the messages and corrected themselves.
What did retention look like after?
After full rollout, Amplitude analyzed the server in its own product. Over 12 weeks of external usage: 67% of users returned after one week, 59% after four weeks, and 50% after eight. Heavy users grew fastest.
For product managers building MCP servers, tool catalog size is a product decision, not a coverage decision — and the shape of that catalog will increasingly determine retention as agent workflows become the default surface for analytics tools.
via platform.claude.com (Original)
More from Nathan Brooks
Show full bio
Senior reporter covering consumer brands and retail at Roadmap File.
21 articles