Self-improving AI is widening what counts as a software change

AI systems are beginning to propose changes to agents' prompts, tools and models. CIOs must decide what counts as a change and how far the evidence lets it travel.

"Self-improving AI" sounds like the kind of phrase designed to make a CIO sit up. If software can look at what happened in production, decide it could have done better and then change how it behaves the next time, the obvious question is whether the enterprise still controls the change -- or whether the software has begun to control itself.

At least for now, it is less dramatic than it sounds.

Autoheal, an AI automation startup building what it calls a self-improving software factory, offers a useful example. Its Evaluator agent scores other agents' work. A second agent, called the Healer, proposes changes to their prompts, tools, skills or model selections based on those results.

But those proposed improvements do not simply become the new production behavior. Autoheal says they are tested against historical benchmarks, version-controlled in Git and require an engineer to approve them.

That looks less like software rewriting itself on the fly than AI inserting itself into a familiar process: something identifies a problem, proposes a change, tests it and puts it in front of somebody who decides whether it should go into production.

This example raises a less dramatic but more useful question for CIOs: If the change is still versioned, tested and approved, what exactly has changed about IT change management?

The answer so far does not look like "everything." Familiar controls still matter. What AI changes is the number of things that can alter production behavior -- and therefore the number of things an enterprise might have to treat as a meaningful change.

The change goes beyond the code

A conventional application change can often be identified because somebody changed the code, configuration or infrastructure. An AI agent makes that fuzzier.

For example, its prompt, model or routing policy might change. A tool might be added or removed. Its skills, permissions or workflow might change. Any of those could alter what the system does the next time it runs, even if little or no conventional application code has changed.

The AI platform vendor Harness shows both sides of the problem. Harness is still using recognizable software delivery controls -- testing, human gates, progressive rollout, canaries, risk scoring and rollback. But it is applying them to nondeterministic agents whose behavior can change when a prompt, model setting or other component changes.

So, the controls are familiar. The harder part is deciding what they now must control.

Not every prompt edit, model switch or new tool deserves the same scrutiny. And technical size is probably the wrong way to decide which ones matter most. What matters more is what the change can affect in the business.

A small prompt adjustment that changes the wording of an internal summary is one thing. A similarly small change that alters when a customer can receive a refund is something else. In HR, a model or workflow change could affect candidate screening. In ERP, it could affect an approval path or financial transaction.

IT still has to know what technically changed because that determines what can be versioned, tested, deployed and rolled back. But the potential business consequence tells you how much the change matters.

Julie Irish, CIO at Alteryx, recently told TechTarget that the level of supervision for an AI agent should reflect the consequences of a mistake. Easily reversible work, such as drafts and internal summaries, falls into a different category from actions such as moving money, deleting data, contacting customers or affecting someone's legal rights.

That was about what an agent is allowed to do. The same logic applies when something changes what the agent will do next time.

Evidence should buy scope, not certainty

Once a change could materially affect the business, the next question is harder: What would make the evidence good enough to approve it?

Historical testing helps. So does comparing the proposed version with the existing one, exposing it gradually to actual production conditions, getting business-owner review and knowing that the change can be rolled back.

However, none of that gives you certainty. Even if a proposed version performs better against every historical case the enterprise tests, those cases describe conditions the organization already knows.

The enterprise is approving the change for what comes next.

Applications change. Business processes change. Regulations, customers, data, integrations and surrounding systems change. Something can look universal now and still turn out not to be universal later.

There is always a line beyond the horizon where foreseeable becomes unknowable. The problem is that nobody gets to plant a flag showing exactly where that line begins.

So "Is the evidence good enough?" may not be the best final question. A more useful one is: How much change does this evidence justify?

For a low-consequence change, the answer might be quite a lot. If a prompt adjustment improves an internal summary and the result is easy to check, relatively modest evidence could support broad use.

The calculation changes when the modification affects ERP approvals, customer refunds, employment decisions or financial transactions. The greater the potential consequence and the harder the result is to reverse, the harder it becomes to justify an unrestricted change.

For some sweeping changes, the bar might simply be too high to clear.

That does not mean rejecting the improvement. It can mean breaking it into more manageable bits: one workflow, one group of users, transactions below a certain threshold or a defined period before reevaluation.

In that sense, evidence should buy scope, not certainty.

A human approval gate helps, but it does not settle the question. It also risks AI agents overwhelming human reviewers.

Evidence should buy scope, not certainty.

Taking the person out does not make the judgment disappear. It just moves the judgment somewhere else. If an automated policy says a change has passed enough tests to deploy, somebody still had to decide which tests count, what passing looks like and how much risk that threshold is supposed to permit.

The same is true over time. A change that was reasonable when approved could operate six months later around a different model, business rule or downstream system. Nothing about the agent itself must change for the assumptions behind an old approval to stop being true.

Approval, therefore, does not necessarily have to mean forever.

Reconstruction looks backward

Once prompts, models, tools, permissions and workflows can all change, an enterprise also needs to know what was running when something happened.

Recent updates to Domo's data and analytics platform provide one example of familiar controls being applied around increasingly agentic systems. The platform added pipeline versioning and more visibility into processes initiated by agents and assistants, including a Model Context Protocol (MCP) trigger that identifies when an agent or assistant starts a process.

Organizations also need an audit trail connecting the original task with the information used, actions taken, approvals received, other agents involved and final business outcome.

Those are related but slightly different problems. One is preserving versions and system activity. The other is reconstructing how an outcome happened. Together, they point toward the need to know enough about a system's state and actions to investigate what went wrong and, where possible, get back to a known state.

But reconstruction has one big advantage over approval -- it looks backward.

The state you are trying to reconstruct actually existed. The prompt was there. The model ran. The transaction happened.

It is a little like the easier half of a time-travel story: you are trying to get back to a place that actually existed.

Approval looks the other way. You are deciding whether to let a change operate in a future state that has not happened yet.

That is why reconstruction matters but comes after the harder questions: What counts as a consequential change? What can it affect? What evidence supports it? And how far should that evidence allow the change to travel?

Self-improving AI still needs IT change management

So far, self-improving AI does not look like a reason for CIOs to throw out conventional change management. Versioning, testing, staged deployment, human review and rollback still matter.

What's changing is the definition of the change itself.

Even when the evidence says new behavior is better, CIOs still must decide what "better" actually entitles the system to do.

Maybe it justifies one workflow but not the whole enterprise, one kind of transaction but not another, or a defined period of operation but not permanent approval.

The goal is not to stop AI systems from getting better. It is to avoid confusing evidence that something worked under known conditions with proof that it should be allowed to go everywhere, affect everything or last forever.

For CIOs, that may be the more useful way to think about self-improving AI: not as software escaping change control, but as software forcing enterprises to get much more precise about what counts as a change, what evidence justifies it and how far that change should be allowed to travel.

James Alan Miller is a veteran technology editor and writer and Lead Editor for CIO News at Informa TechTarget. He directs coverage of enterprise technology strategy, AI, software, data, infrastructure and the decisions shaping how CIOs manage increasingly complex IT environments.