Decision #217AcceptedTrack · AI in Product5 min read

Orchestration Beats Model Choice: 33–61% AI Cost Cuts Without Switching Models

Research ran 22 enterprise tasks on six models: changing only the orchestration harness cut costs 33–61%, tokens 38% and latency 44% — beating model switching.

Context

  1. Changing only the orchestration harness cut costs 33–61% across six models on the same 22 enterprise tasks.

  2. Tokens per task fell 38% and median latency fell 44%, with completion quality broadly unchanged.

  3. The orchestration layer moved cost more than switching between the cheapest and most expensive model tested.

  4. Anthropic's MCP work confirms loading tools on demand and filtering results reduces latency and cost.

  5. IBM's expanded shadow AI definition covers ungoverned agents, MCP servers and autonomous workflows — a cost problem as much as a governance one.

Changing only the orchestration harness around a model cut enterprise AI costs by 33–61% across every model tested — a bigger saving than switching between the cheapest and most expensive model in the same experiment. That is the headline result from recent research that ran the same 22 enterprise tasks on six foundation models and varied nothing except the layer surrounding them. Tokens per task fell 38%, median latency fell 44%, and aggregate completion quality stayed broadly level.

The finding should reorder how product and platform teams attack AI efficiency. Most cost conversations start with the model price table: who cut input-token charges, whether a smaller model can handle the task, whether routine work can be routed away from the frontier model. Those questions are sensible but downstream of the more important one: why are we sending the model so many tokens in the first place?

Why falling token prices can hide rising costs

Token prices have dropped fast, which creates the impression that usage will naturally get cheaper. But agentic systems consume more tokens per task as they mature: longer reasoning traces, larger tool catalogues, more retrieved context, more turns, and full conversation histories replayed at every step. The unit price can fall while the total bill rises.

The useful unit of measurement is therefore not cost per token. It is cost per successfully completed task. A cheap model in a wasteful loop can cost more than a frontier model inside a disciplined workflow — and deliver a worse result that requires more human checking.

Where do all those input tokens come from?

A model controls the length of its own answer. It does not decide most of what appears on the input side. The application assembles the system instructions, conversation history, tool schemas, retrieved documents, intermediate results and the current request — four of those six categories are constructed by software before the model sees anything.

That makes orchestration a P&L concern rather than plumbing. The layer around the model decides whether every turn receives the entire transcript, every available tool and every related document, or only what the current step needs.

The naive agent loop is the worst offender: it replays full history each turn, so total input grows much faster than useful work. A larger context window makes this possible. It does not make it sensible. Better orchestration separates stable from volatile context — cache system instructions and tool definitions, compact older conversation into structured state, store large tool outputs outside the prompt and refer to them when needed, and suspend a workflow waiting on a person rather than polling the model. The goal is not crude prompt shortening; it is preserving the information that changes the decision and refusing to pay to resend everything else.

Can a workflow choose the cheapest reliable path?

Cost control is not only about context. Different task steps need different capabilities. A bounded classification suits a fast, inexpensive model. Synthesis across contradictory evidence may justify a frontier model. Identifier validation, arithmetic, permissions and completion checks should run deterministically rather than be delegated to any model.

ProdPad Conductor, the orchestration product referenced in the source analysis, makes these choices step by step. It can stop a workflow once the completion rule is satisfied, repair a failed structured response without restarting the task, and stop an agent repeatedly exploring the same dead end. That beats assigning one model to a whole process and hoping its average cost is acceptable.

Shadow AI is shadow spend

Employee-built agents compound the problem. Each carries its own long prompt, duplicates product context, defines overlapping tools and repeats work another agent already did. Individually none looks expensive; collectively they become shadow IT with a token bill — duplicated context, duplicated tools, duplicated mistakes, and no shared way to verify completion.

Anthropic has made the same point from another direction in its work on MCP and code execution: loading every tool definition and passing every intermediate result through the context window inflates both latency and cost. Loading tools on demand and filtering results before they reach the model is orchestration hygiene, not a cheaper-model trick. IBM's expanded definition of shadow AI — teams spinning up agents, MCP servers and autonomous workflows faster than the organisation can govern them — frames the same phenomenon as a discoverability gap. The answer is not banning experiments but giving successful ones a route into shared orchestration where context, routing, traces, completion and cost are managed once.

What should teams actually measure?

Teams that report only answer quality or raw model usage will token-max, because more context and more reasoning look safer when the compute cost sits on someone else's line item. A better set of measures connects spend to outcome:

  • successful task completions per million tokens
  • quality or acceptance rate per pound or dollar
  • cost and elapsed time per completed workflow
  • retry, repair and human-intervention rates
  • cost broken down by workflow step and model

These measures reveal whether a pricier model genuinely reduces total workflow cost or just moves spend to a more visible line.

Product management is especially exposed because its context is broad and connected — feedback, ideas, objectives, roadmaps, decisions — and not all of it is relevant to every step. The model price still matters. The bigger lever is architectural: cache the stable, compact the old, retrieve the relevant, route the bounded, stop the failing. The cheapest token is not the discounted token; it is the token the workflow never needed to send. As agentic estates grow, expect cost reviews to move from procurement's price table to the orchestration layer where the tokens are actually spent.

via x.com (Original)

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering marketplaces and e-commerce at Roadmap File.

22 articles