Decision #182AcceptedTrack · AI in Product4 min read

Orchestration, Not Agents, Cut Enterprise AI Costs by Up to 61%

Research shows swapping the orchestration layer, not the model, cut task costs 33–61%. ProdPad's Conductor shows why enterprise AI reliability lives around the model.

Context

  1. Research on 22 enterprise tasks across six foundation models found changing only the orchestration layer cut cost per task 33–61%, tokens per task 38%, and median latency 44%.

  2. ProdPad built Conductor, an orchestration layer powering CoPilot PM, handling state, routing, permissions, validation, retries and human hand-offs.

  3. OpenAI describes the harness as a control plane owning the agent loop, tool routing, approvals, tracing, recovery and run state.

  4. IBM has introduced the terms 'agentic estate' and 'agentic control plane' for governing agent sprawl across organizations.

  5. Anthropic's workflow-vs-agent distinction underpins the design advice: bounded workflow steps contain model variability better than one open-ended agent loop.

Changing only the orchestration layer around a model — not the model itself — cut cost per task by 33–61% across six foundation models, reduced tokens per task by 38% and median latency by 44%, while task quality stayed broadly at parity. That research result is the core argument of a new ProdPad essay: enterprise AI fails or succeeds in the layer wrapped around the model, not in the model.

ProdPad, which built its own orchestration layer called Conductor to power CoPilot PM, frames the current agent boom as a repeat of the chatbot era. "'Just add an agent' is becoming the new 'just add a chatbot'", the company writes. "It sounds like an answer because it names a technology. It usually avoids the harder question: what, exactly, is responsible for getting the work done?"

What does the orchestration layer actually do?

OpenAI now describes the harness around a model as a control plane: it owns the agent loop, tool routing, hand-offs, approvals, tracing, recovery and run state. ProdPad prefers the broader term orchestration layer, because an enterprise system must also coordinate deterministic software, business workflows and people — not just the model loop.

Inside ProdPad, Conductor's job list is concrete: state management, routing, permissions, validation, retries and hand-offs back to a person. The company's verdict after shipping it: the hard part was never getting a model to do something impressive once. It was making the result dependable enough to become part of somebody's work.

Why do demos mislead?

Most AI demos are, in ProdPad's words, "happy paths wearing a lab coat" — clean context, narrow task, an expert nearby. Production brings missing context, inconsistent data, unclear requests, permissions, tool failures and undocumented exceptions. Enterprise controls — access, approval, traceability, ownership, exception handling — are not details bolted on later. They have to live in the architecture: what each step can see, what it can do, how results are checked, when the decision returns to a human.

Agent flexibility compounds this. The same request may trigger different context retrieval, different tool sequences, redundant reasoning steps, plausible-but-adjacent answers, or work that continues when it should have stopped. Reviewing the prose is a poor way to judge whether the workflow succeeded.

How do you contain model variability?

Anthropic's distinction anchors the design advice: workflows follow predefined code paths; agents dynamically decide how to pursue a task. Enterprise systems need both. The mistake is assuming autonomy removes the need for workflow design — in practice, autonomy makes orchestration more important.

ProdPad's approach: break work into bounded steps, each with a clear purpose, limited context, constrained tools, an expected output and an explicit definition of completion. Some steps are read-only; a smaller number change product data; consequential decisions return to the user. Planning, data gathering, interpretation and execution are separated, because collapsing them into one open-ended loop makes failures impossible to localize.

The target is dependable outcomes, not deterministic wording. Wording can vary where it helps the model reason or communicate. Variability is unacceptable around authority, evidence and completion.

Why is orchestration also cost control?

Most input tokens are not chosen by the model. The application decides which system instructions, history, tool definitions and retrieved material get sent every turn. A naive agent loop replays all of it repeatedly; a disciplined orchestration layer caches stable context, compacts older history, offloads bulky results and stops failed workflows before another retry cycle burns budget. Model choice cannot compensate for an architecture that sends the model far more than it needs.

What did reliability require in practice?

Much of Conductor's engineering was unglamorous:

  • Clean empty or default arguments before tool calls
  • Validate structured output rather than trusting well-formed JSON
  • Check identifiers and permitted values
  • Retry or repair output that fails validation
  • Stop when the system lacks enough information to proceed safely

Models fail in ordinary software-shaped ways — omitted fields, wrong identifiers, invalid options, well-formed structures inappropriate for the workflow state. Asking the model to be more careful helps only up to a point.

Testing follows the same logic: scenarios, not scripts. Can users express the same intent differently, skip ahead, change their mind? Can an empty result or a failed structural validation be handled without collapsing the flow?

Where is enterprise AI governance heading?

IBM has started using "agentic estate" for the sprawl of agents, models, tools and workflows accumulating across organizations — an estate that exists before anyone agrees how to operate it. IBM pairs this with a proposed agentic control plane for visibility, credentials and tracing across frameworks. ProdPad positions that control-plane problem as adjacent to orchestration, not a replacement: the control plane gives enterprise-wide authority; the orchestration layer decides how a specific piece of work moves between models, tools, software and people.

As vendors embed agents into every product and prototypes escape into daily work, the teams that treat orchestration as the actual product — bounding authority by design, testing workflow states, not sentences — will be the ones whose AI survives contact with production.

via ibm.com (Original)

More from Priya Raman

Priya Raman

Show full bio

News editor covering business strategy at Roadmap File.

27 articles