As AI agents take action, CIOs have to govern the outcome

As Oracle and Progress push AI agents deeper into business processes, CIOs need governance that measures outcomes and defines when humans intervene.

AI agents can follow every access rule, pass every technical check and still produce a bad business outcome. An agent might increase throughput while worsening customer satisfaction or mark a workflow complete even though the expected change never appears in the destination system.

The AI can be operating as designed while the business absorbs the damage. That problem is becoming more immediate as vendors push agents deeper into live business processes.

Oracle is pushing more autonomous agentic AI into ERP, while Progress is adding workflow auditing, AI operations logging and permission-bound assistants to the Domo AI and data platform business it acquired this month. For CIOs, the post-deployment question is no longer just what an agent is allowed to do. It is whether its actions are producing acceptable outcomes -- and how quickly the organization can detect, own and correct them when they are not.

Those failures may not trigger a conventional technical alert. Post-deployment AI governance therefore has to connect real-world outcomes to human accountability and corrective action.

"The governance question isn't 'Is this AI safe in theory?'" said Tom Romanoff, global policy director at the Association for Computing Machinery. "It's 'Do we have the organizational capability to catch problems and act on them before they damage the business?'"

That requires governing AI as an active business operation rather than technology approved once and left to run.

Define what good looks like

Organizations first need to define an acceptable outcome in precise, measurable business terms. Technical measures such as accuracy, drift, latency and task completion can explain system performance, but they do not determine whether the AI is helping or harming the business.

A stronger measure pairs productivity with quality, cost and harm indicators. Depending on the workload, those indicators might include customer complaints, corrections, reversals, reopened cases, manual rework, unusual spending, failed handoffs, exceptions and employee overrides.

Julie Irish, CIO at Alteryx, recommends defining what good looks like and then monitoring the system against it. Organizations need enough logging and alerting for a human to identify a problem and intervene quickly, she said.

The level of supervision should reflect the consequences of a mistake. Irish said agents can earn greater autonomy as their work becomes more trusted, but a human should retain final approval for complex activities involving legal matters, billing or customer impact.

That does not mean applying human review indiscriminately. It means deciding which outcomes can be checked later and which actions must be checked before execution. Drafts and internal summaries are generally reversible. Moving money, deleting data, contacting customers or affecting someone's legal rights may not be. For higher-impact actions, organizations may also need runtime controls that can intervene before execution.

Give the outcome a human owner

Dashboards cannot be held accountable, and committees may not be able to respond quickly enough to every live incident. Each consequential AI workflow needs a named business owner who understands what the system is supposed to accomplish and has the authority to change or stop it.

That person is not necessarily the technical owner. Technology teams may detect that a model, integration or usage pattern has changed. The business owner must decide whether the resulting behavior remains acceptable.

"One person must be able to stop it," said Jyotsna Jha, vice president of product strategy, development and services at Innodata. "A quarterly committee cannot intervene on a Tuesday afternoon."

A practical rule is to pin responsibility to the person who manages the function and would own the result if people performed the same work. A customer service leader should remain responsible for automated customer decisions. A finance leader should remain responsible for automated financial actions. Using AI should not create a gap in accountability.

It should also be clear who can restrict the system without waiting for several levels of approval. If no one can name the person with that authority, the governance structure is incomplete.

Verify the outcome independently

Organizations should not rely only on an AI agent's account of what it did. The system's report is part of the evidence, not proof of the outcome.

An AI's report of its own work is a lead, not a fact.
Wesley Cablefounder, PipelineOS

Wesley Cable, founder of PipelineOS, learned this while using agents that write, publish and interact with live client systems. In one case, a package was approved for publication, but the separate publishing job hit its budget limit, stopped and never retried. The expected page never went live.

Cable's company now checks the destination itself by loading the public page and searching for the approved text.

"An AI's report of its own work is a lead, not a fact," he said.

Independent verification should occur in the business system where the result appears: the public website, customer record, payment ledger, inventory system or communications platform. The same principle applies to incident investigation. Reviewers need records of the actions actually taken, approvals obtained and downstream outcomes, not simply the model's explanation of its behavior.

Establish intervention triggers before an incident

The term "human oversight" has little meaning unless an organization defines what will cause a human to act. Intervention thresholds should be established while stakeholders are still objective rather than during a crisis, when revenue, deadlines and internal politics may discourage action.

Triggers might include a sudden increase in complaints, repeated corrections, an unusual override pattern, actions outside the intended purpose, mounting financial exposure or deteriorating outcomes for a protected group.

The response should already be attached to the trigger. That could mean requiring additional approval, reducing transaction volume, removing access to one tool, returning part of the workflow to people or suspending the system.

A complete shutdown should not be the only option. Companies can hesitate to turn off an AI system when doing so would interrupt the underlying business process. Jha recommends preparing smaller moves that reduce scope, dial back autonomy or send more decisions to humans. Those interventions should also be tested before they are needed.

Watch where humans and AI disagree

A useful governance record that companies might overlook is the disagreement log.

Patrick Bryant, co-founder and CEO of CODE/+/TRUST, recommends examining where humans override AI decisions and where AI systems escalate decisions to humans. Those events show where the system's judgment and assigned mandate diverge.

A high override rate may mean the AI has too much autonomy, its instructions are inadequate or business conditions have changed. An unusually low rate may not be reassuring either. It could indicate that reviewers are rubber-stamping output rather than exercising meaningful judgment.

Frontline employees provide another early warning system. Workers can notice questionable recommendations before a problem appears in aggregate reports. They need to know what normal behavior looks like, how to flag an anomaly and who is responsible for responding. Otherwise, governance exists on paper but not in practice.

Preserve the chain of responsibility

As agents begin invoking tools and delegating work to other agents, organizations must be able to reconstruct the resulting chain of action and authority.

Jason Sabin, CTO of DigiCert, said every agent should have a verifiable identity, a designated human or business owner and defined boundaries around the systems, data and actions it can access. DigiCert calls its branded approach an "AI Agent Passport."

The governance question is not only what an agent did, but on whose authority it acted. Organizations should be able to connect the original task with the information used, actions taken, approvals received, other agents involved and final business outcome.

Mehdi Houdaigui, principal and Cyber AI leader at Deloitte Consulting, said that without that traceability, it becomes difficult to determine whether a failure arose from the AI, a control problem, misuse or a flaw in the surrounding business process.

Post-deployment governance cannot guarantee that AI will never make a bad decision. Its purpose is to recognize an unacceptable outcome, identify who owns it, intervene before damage spreads and apply what was learned to the next decision.

That is what turns oversight from a policy document into an operating capability, Houdaigui said.

Related video: Targeting AI: Building the AI-Powered Enterprise examines how enterprises can govern agentic AI as it moves deeper into workflows and decision-making. 

Pam Baker is a freelance journalist and the author of books including ChatGPT For Dummies and Generative AI For Dummies. Baker is also an instructor on AI topics for LinkedIn Learning and a member of the National Press Club, the Society of Professional Journalists and the Internet Press Guild.