Decision #593AcceptedTrack · AI in Product3 min read
Cheaper Tokens Won't Fix Enterprise AI, Salesforce VP Tells Fortune
Salesforce's VP argues in Fortune that cheaper inference cannot plug the structural leaks between AI pilots and production, urging product teams to invest in evaluation and integration before negotiating model cost.

Context
A Salesforce vice president told Fortune that cheaper tokens cannot fix enterprise AI's deployment gap.
The executive described enterprise AI rollout as a 'leaky pipeline' of structural failures that price cuts cannot plug.
The argument reframes AI cost from per-token inference to full-pipeline plumbing including integration and evaluation.
Integration and evaluation are the most common leak points, and cheaper tokens offer no remedy for either.
Token price is rarely the binding constraint in mid-to-large enterprise rollouts.
Salesforce's vice president pushes back on the inference-cost narrative in a Fortune interview, arguing that reducing token prices will not close the gap between AI pilots and production systems. The executive frames enterprise AI deployment as a "leaky pipeline," in which structural failures overwhelm any savings from cheaper inference.
What's actually leaking?
The deployment chain rarely fails at a single visible step. It leaks across several: data readiness, system integration, evaluation rigor, change management, and post-launch monitoring. Each stage carries its own failure modes, and a failure at one compresses the value of the others.
A model that performs well in a controlled lab often falters when it meets production data, real customer permissions, or existing workflow latency budgets. Cheaper tokens reach the inference line item. They do nothing for the surrounding plumbing.
Why does the cost narrative persist?
Token pricing is easy to compare and easy to lower. It dominates vendor roadmaps, RFPs, and procurement discussions. The shift toward smaller models, distillation, and aggressive discounts has produced real per-call savings. Budget conversations reward teams that can show the inference line shrinking.
The trade-off: those savings get credited to the model layer even when most deployment cost sits in integration, evaluation, and human review. The result is a roadmap that underinvests in the components that determine whether an AI feature ships successfully.
What should product teams track instead?
Treat the full pipeline as the unit of cost, not the token. Roadmaps that bind AI deployments to production outcomes typically invest in:
- Data preparation and labeling with quality bars the model can rely on
- Connectors and retrieval pipelines that survive real permissioning and schema drift
- Evaluation infrastructure including offline evals, shadow deployments, and outcome metrics
- Human-in-the-loop designs for high-stakes calls and edge cases
- Post-launch monitoring with feedback loops that retrain evaluation datasets, not just models
Each item has different failure modes, and each carries its own cost line. Token price is rarely the binding constraint in mid-to-large enterprise rollouts.
Where does the leak usually show up?
The loudest leaks appear at integration. A retrieval-augmented system without proper indexing returns confident wrong answers. An agent without tool scoping exceeds its mandate. A copilot without telemetry cannot show whether it helps. In each case, cheaper tokens offer no remedy.
Evaluation is the second frequent leak. Teams that cannot score their own outputs cannot tell whether a model swap improved anything. They end up debating vibes in standups while production behavior drifts.
What are the tradeoffs?
Cheaper tokens do expand the feasible surface area. They make high-volume, exploratory, and consumer-scale assistants viable in ways 2023 pricing did not. The Salesforce framing accepts this. It argues that enterprises optimize the wrong lever when they treat per-token price as the primary constraint.
A defensible product strategy sequences the work: instrument the pipeline first, then negotiate model cost. Reversing the order produces assistants that score well on a benchmark and underperform on the metrics the business tracks.
What's the forward view?
Expect the binding constraint in enterprise AI to shift from model capability to operational discipline. Teams that build evaluation, integration, and monitoring muscle now will own the production workloads of the next cycle. The "leaky pipeline" framing works best as a forcing function: spend where the water actually leaves the bucket, not where the vendor invoice is easiest to read.
via Google News - AI Product Management (Source)