Decision #236AcceptedTrack · AI in Product5 min read
Amplitude Cuts Agent Costs by Two-Thirds With Kimi, Matching Sonnet
Amplitude swapped Sonnet 4.6 for Kimi K2.7 in Global Agent: after fixing five failure modes, Kimi scored 73.7 vs Sonnet's 72.7 at roughly one-third the cost.

Context
Kimi K2.7 scored 73.7 vs Sonnet 4.6's 72.7 on Amplitude's 186-case eval after remediation, at roughly one-third the cost.
Microsoft's July 24 open letter on open weights grew from 25 to over 200 signatories in its first week.
Stanford HAI reports the Chatbot Arena open-vs-closed preference gap fell from 8% to 1.7% in a year.
About 80% of failed eval cases traced to the head agent; Global Chat is 44% of Global Agent spend.
Kimi K2.7 ranks #1 on τ²-bench Telecom at 93%, trained for 200–300-step tool chains.
Amplitude cut the cost of its in-product Global Agent by roughly two-thirds by swapping Anthropic's Sonnet 4.6 for the open-weight Kimi K2.7 — and after targeted remediation, Kimi scored 73.7 on Amplitude's internal 186-case evaluation suite, edging out Sonnet's 72.7 on identical test cases. Principal Product Manager Jacob Newman published the full methodology, arguing the win came from tuning, not from the model being better out of the box: "Kimi K2.7 wasn't superior out of the box. It was a Sonnet-class open-weight model that needed Amplitude-specific tuning and correcting to account for specific failures."
The backdrop is an industry-wide realignment on open weights. On July 24, Microsoft published an open letter, "Open Weights and American AI Leadership," arguing open-weight models are critical infrastructure for competition, security, and AI diffusion into everyday business. The original 25 signatories — Nvidia, Meta, Dell, IBM, Palantir, Mistral, Hugging Face — grew past 200 in the first week, adding OpenAI, Google, Amazon, Uber, Databricks, GitHub, and Fireworks AI. Amplitude signed on. CEO Spenser Skates described the company in his joining note as "heavy users of open-weight AI," noting that Kimi and GLM are already core to operations, internally and in customer-facing features.
Why open weights, and why now?
The Stanford HAI AI Index puts the human-preference gap between open and closed models on Chatbot Arena at 1.7%, down from 8% a year earlier. Kimi K3, the 2.8-trillion-parameter model Moonshot released in late July, reportedly performs competitively with Anthropic's Fable 5. But open weights are not an automatic upgrade. The New Stack frames the Kimi K3 tradeoff bluntly: "same results, one-third the cost, four times slower." Beyond price, open weights offer autonomy — download, install on your own infrastructure, fine-tune against a specific product surface, and keep data from leaving on every third-party call.
How Amplitude ran the evaluation
The existing stack ran Sonnet 4.6 across all Global Agent subagents: Global Chat, Chart Agent, Data Assistant, Session Replay SQL, Preference Extractor, and smaller ones. Because Sonnet's weights are closed, Amplitude could prompt around failure patterns but never fine-tune against them. The test goal: hold the existing customer experience at meaningfully lower cost.
Infrastructure choice came first. Fireworks (managed catalog, per-token billing, nothing to operate) beat Modal (raw GPU containers, own your quantization and batching). Inside Fireworks, dedicated GPU instances matched serverless on Amplitude's quality suite, delivered the most consistent latency, and carried a fixed hourly cost that scales with volume rather than token count.
Model selection mapped three candidates against Claude's lineup:
- gpt-oss-120B — near Haiku: strong tool-calling on τ-bench, 11–25x cheaper, but weak instruction-following suits narrow subagent work, not orchestration.
- Kimi K2.6/K2.7 — near Sonnet: trained for 200–300-step agentic tool chains, ranked #1 on independent τ²-bench Telecom at 93%, at roughly double the thinking-token burn.
- GLM 5.2 — near Opus, not quite: within about four points on SWE-bench and Terminal-Bench at roughly 6x lower cost, but loses blind preference reviews and struggles on long-horizon tasks.
Kimi K2.7 won the head-agent slot not for paper scores but for fit.
What failed, and what fixed it
Kimi initially scored 64 on the eval suite — a real gap behind Sonnet. Production telemetry showed subagents rarely trigger but dominate wall-clock time when they do: Data Science at 299 seconds, Chart Deep Research at 217. Routing chart work through a single-agent compiler flow matched production correctness at roughly 15% lower mean latency. Caching and parallelizing permission checks cut seconds of blocking overhead to sub-second and eliminated false chart-permission denials; typed, validated chart contracts stopped invalid definitions before persistence.
About 80% of failed cases traced to the head agent. Five recurring patterns, nearly absent under Sonnet, drove the fixes:
- Capability hallucination — inventing UI paths or plan limits from training data
- Degenerate tool loops — retrying the same wrong tool until the turn died
- Hand-done math and dates — epoch conversions off by a year, weekdays shifted by one
- Chart semantic near-misses — syntactically valid, semantically wrong, reported as success
- Output leakage — raw reasoning in the final answer, or empty messages
None of the remediations were better-worded prompts. A loop breaker cuts tool access after repeated bad calls and forces an honest answer. A dedicated compute tool and date table bypass the model's own arithmetic. Grounding rules require a docs or tool check before any capability claim. A sanitizer strips leaked reasoning tags; a context clamp ties history compression to the actual context window — the root cause behind nearly every "prompt too long" failure.
What does the swap actually save?
The same workload on Kimi K2.7 costs roughly a third of Sonnet's price — and the real gap may be wider, since Anthropic bills cache writes and Fireworks does not. Global Chat alone accounts for roughly 44% of Global Agent spend, which is why the swap concentrated there first. Caveats remain: Kimi's launch benchmarks are vendor-reported, independent coverage is thin, and some third-party trackers still give Sonnet the edge on tool-routing reliability. Amplitude is not moving every customer at once; it is testing across segments to see whether one model serves everyone or whether cohorts split between Kimi and Sonnet, letting live traffic settle what the eval suite can't.
Newman's sequencing is the transferable part: baseline cost per session and failure patterns on production traffic, test the candidate against your own evals, fix the named gaps, then let a live slice validate. As token prices climb and open models close the benchmark gap, deployment decisions like this — grounded in instrumentation, not vendor marketing — are becoming the default practice for AI product teams.
via microsoft.com (Original)