Decision #918AcceptedTrack · Product Analytics4 min read
Jev decision model: 70–500ms per call, but accuracy gaps remain
Jev catches every meaningful chart in Amplitude's synthetic benchmark at $0.06 per 1,000 charts — but confident errors and ~12-point overconfidence mean teams need outcome-based measurement, not just a cascade.
Context
Jev, TypeSafe AI's decision model, runs 70–500ms per call; Amplitude measured it at roughly $0.06 per 1,000 charts screened.
An AY Automate study of 791 decisions found a 0.80-gated cascade matched frontier accuracy at ~25% of cost and half the latency, but the decision model matched ground truth 6–8 points less often than it matched the frontier model.
A 77-case BANKING77 pilot showed 87.4% mean confidence vs 75.3% actual accuracy — ~12 points of overconfidence; jevbench.xyz accuracy ranges 62.6%–95.4%.
In Amplitude's synthetic benchmark (64 charts, 190 time series, 24 injected changes), Jev 1.13 caught every meaningful chart at under half the cost and a quarter of the latency of DeepSeek V4.1 Flash, with lower precision.
Analyzing 27,000 agent sessions, Amplitude found 98% gave judges no quality signal, while users with clean first sessions who saved output retained at 3x the rate.
Jev, TypeSafe AI's flagship "decision model," runs a call in 70 to 500 milliseconds and costs so little that Amplitude put a price on broad screening: roughly $0.06 per 1,000 charts. But in independent benchmarks and Amplitude's own tests, decision models share the same accuracy pitfalls as their slower, more expensive LLM cousins — and teams that deploy them without a measurement layer are, in the authors' words, "on a faster and cheaper road to potential slop."
The analysis comes from Vinay Goel and Ram Soma, staff AI engineers at Amplitude, published September 21. Their argument targets the whole class of decision models, not just Jev.
What is a decision model, and where does it fit?
Like an LLM, a decision model takes text and code as input. Unlike an LLM, it returns typed decisions or probabilities instead of prose. Ask an LLM where your funnel drops off and it may answer conversationally; a decision model with a configured answer space simply returns "checkout." The constrained output makes it fast and token-cheap.
Decision models do two jobs: they perform work (routing tickets, gating tool calls in production) or grade work (scoring sessions as a cheap judge). The first kind runs all day inside agents — and is the harder one to measure.
What do the benchmarks actually show?
Mixed results. An independent AY Automate study of 791 labeled decisions found that a cascade gated at 0.80 matched frontier-model accuracy at roughly a quarter of the cost and half the latency. But the decision model agreed with the frontier model 6 to 8 points more often than it matched ground-truth labels — the study measured agreement between models, not correctness.
The failure mode is precise: confident errors clustered on semantically similar intents, such as "Direct debit payment not recognised" versus "card payment not recognised." Five confident answers were wrong that way; the frontier model missed the same five. The cascade, built to catch exactly this, did not.
Anthony Maio, AI Engineer at Pieces, put it cleanly: the type system constrains the shape of the output, not the judgment. Clean types, high confidence, no exception — and the ticket still goes to the wrong queue.
Calibration doesn't save it. jevbench.xyz records 91.7% on a 60-case tool-call set but a range of 62.6% to 95.4% across everything it has run, and declines to publish one headline number. A 77-case BANKING77 pilot found 87.4% mean confidence against 75.3% actual accuracy — about 12 points of overconfidence. And as Maio notes, individually calibrated judgments do not compose into a calibrated workflow once you add thresholds, weights, and branches.
How did Jev perform in Amplitude's analytics test?
Amplitude built a synthetic benchmark of 64 charts and 190 time series covering segmentation, funnels, retention, and customer journeys. They injected meaningful changes into 24 charts; about 15% of series contained spikes, drops, level shifts, segment divergence, volatility changes, or instrumentation breaks. The remaining 40 charts were realistic-looking controls. Jev 1.13 was compared against Luna, DeepSeek V4.1 Flash, GLM-5.3 Flash, and Sol.
Jev caught every meaningful chart, at under half the cost and under a quarter of the latency of DeepSeek V4.1 Flash. The tradeoff was precision: Jev escalated more benign charts than the best general-purpose models. In a manage-by-exception workflow, that may be acceptable — unless every flag triggers an expensive root-cause analysis, in which case false positives can erase the savings.
The authors caution this is synthetic data; production evaluation should use a blinded sample of real charts labeled by analysts.
What should teams do?
Goel and Soma recommend decision models as the first layer of a cascade:
- Run Jev frequently across the full inventory
- Filter to high-probability candidates
- Route them to an expensive model, an automated deep-dive, or an analyst
- Tune the threshold against review capacity and false-positive tolerance
A second stage catches uncertain decisions, but confident errors are not uncertain — they cluster where the schema is ambiguous, and the expensive model inherits that ambiguity. In some industry tests, both LLMs and decision models missed against labeled data.
Their stronger claim: every proposed accuracy fix is more judgment (bigger model, reviewer, hand-labeled set), which is why public evidence spans dozens to hundreds of cases — what anyone can hand-label in a sitting. The right instrument is product data. Earlier this year, Amplitude ran Agent Analytics over 27,000 Amplitude Global Agent sessions; in 98%, users gave no signal a judge could use. But users whose first session tripped no failure flag and saved the agent's output retained at 3x the rate of those who did — a correlation no transcript reviewer could have surfaced, because the decisive save event happened after the session ended.
A ticket routed to billing at 0.94 certainty that reopens on Thursday is a graded decision, observed rather than inferred. As agents take over more branch points, teams that grade decisions by downstream behavior — not by another model's opinion — will be the ones whose speed advantage doesn't compound into silent errors.
via docs.typesafe.ai (Original)
More from Rebecca Stone
Show full bio
Correspondent covering marketplaces and e-commerce at Roadmap File.
22 articles