Decision #257AcceptedTrack · AI in Product3 min read
Periodic AI Governance Fails Once Agents Start Acting
IBM's 'continuous AI assurance' reframing exposes a gap most governance frameworks miss: policy states intent, but only runtime traces prove what an agent actually did in production.

Context
IBM used the phrase 'continuous AI assurance' at Think 2026 to describe its shift from periodic review.
IBM separately used the phrase 'orchestration-led governance' to describe where policy must be enforced.
Agentic systems can attach new tools, prompts or MCP servers after the original risk review, causing the approved diagram to diverge from the running system.
ProdPad's Conductor keeps workflow state outside the model, limits tools by step and validates actions.
Sharjeel Ahmad argued, in the context of banking agents, that production is where the demo meets the control framework.
At IBM's Think 2026 conference, the company used the phrase "continuous AI assurance" to describe its shift from periodic review to ongoing visibility, enforceable controls and accountable ownership. That language captures a hard truth for product managers shipping agents: a signed risk review on Monday tells you nothing about what the agent did on Tuesday.
Sharjeel Ahmad, writing about production banking agents, argues that production is where the demo meets the control framework. The framework has to produce evidence while the work happens, not merely describe what should have happened. That is the difference between reviewing a policy and operating a live system.
Why does periodic governance break for agents?
Traditional governance assumes assets can be inventoried, assessed and approved at sensible intervals. That assumption works when a model runs a narrow prediction and ships through a managed release train. Agentic systems act across tools, contexts, permissions and collaborators, and their risk depends on the combination.
A model can change. A prompt can be edited. A new tool or MCP server can connect. An agent can begin collaborating with another agent that did not exist when the original risk review completed. The approved diagram stays the same while the running system quietly becomes something else.
The ProdPad post frames this drift as the way the "agentic estate" grows: every connected tool or collaborator is a new asset the original review never saw.
What does continuous assurance actually require?
Governance tells you the rule. Assurance shows that the rule held. The two are not the same deliverable.
A governance policy may state that a product recommendation requires human approval. Assurance should show that the workflow actually stopped, presented the relevant evidence, recorded the decision and blocked the action through any other path.
A policy may restrict access to sensitive customer feedback. Assurance should show which context the model retrieved, which model received it and which tools were available at that step. None of it reconstructs reliably from the final answer alone.
Traces, evaluations and runtime state are the raw material. Without them, assurance is a claim, not a measurement.
Where do the controls have to live?
IBM has also used the phrase "orchestration-led governance," which gets closer to the architectural point. Rules become real when the orchestration layer applies them while assembling context, exposing tools, validating outputs and authorising actions.
A governance dashboard beside the workflow reports a violation after the event. An orchestration layer can prevent the invalid action or route it to a person before it occurs.
Risk owners still define policy, compliance teams still need oversight, and an enterprise control plane still needs visibility across the wider agentic estate. The orchestration layer is where policy meets execution, and the two cannot be separated without losing enforcement.
How do evals fit into assurance?
Enterprise evals should ask whether the workflow respected the organisation's definition of good work. Did it use authoritative evidence? Did it stay within its permissions? Did it complete the intended task? Did it stop when confidence or authority ran out?
Traces turn a failed outcome into something a human can examine. Evals turn that failure into a test. The next release can then demonstrate that the workflow improved, rather than producing a more convincing answer on a hand-picked example.
ProdPad's Conductor keeps workflow state outside the model, limits tools by step, validates actions and reconciles what actually changed in the product. Human involvement is one of the resources the workflow deliberately invokes, not an emergency brake bolted on at the end.
As the ProdPad post put it: "Governance records the promise. Assurance needs evidence from the running system." Teams that treat that distinction as documentation rather than architecture will ship the next agentic failure mode before they recognise it.
via linkedin.com (Original)