When is AI work really done? CIOs need a clearer finish line
AI can finish a task while the business result remains unfinished. CIOs must define what "done" means to measure AI's true productivity, cost and value.
When the AI says the work is done, the work is done. Well, not so fast.
Maybe the task is done. Maybe the workflow is done. Maybe the business result is done. Those are three very different definitions of completed work. None is completely right or wrong.
They simply describe "done" at different levels.
When an agent drafts a customer response, its assigned task can be considered finished. The response can then be checked, sent and recorded correctly, completing the workflow. And the customer's problem can actually be resolved without the case being reopened, completing the business result.
All three -- task, workflow and business outcome -- are legitimate finish lines.
The trouble starts when one gets mistaken for another. The AI can be done while the business is still waiting. That matters because enterprises are using those finish lines to judge AI's productivity, cost, safety and business value.
The distinction kept showing up in TechTarget reporting last week, even though the stories themselves were about different things -- productivity, cost, governance and pricing.
Together, they raise a pretty basic question for CIOs:
That's real progress. But the company isn't treating the capacity itself as the final result.
Its next question is what happens to that extra capacity. Does it improve client retention? Generate more business? Create more value for customers?
So when was the work finished?
When the AI completed its portion? When the employee got the time back? Or when the company converted that time into something the business actually values?
Those can all be valid answers. They are not the same answer.
The AI can be done while the business is still waiting.
Counting the work requires defining it first
Cost measurement runs into the same problem.
Token consumption and model calls can tell CIOs something about what an AI system used. They don't necessarily account for retries, failed attempts, tool calls or human intervention along the way.
Chris Bennett, vice president of the Global AI Practice at Unisys, put it simply: Establish the "finish line" first, then measure cost, quality and employee involvement against it.
For example, consider an invoice.
Is the work done when an agent extracts the information? Or is it done when it is verified and entered correctly? When somebody approves it? When the invoice is actually paid?
A smaller task might be easier to count, but that doesn't necessarily make it the right finish line.
Done depends on where you stand
A task can be complete while the workflow continues. A workflow can be complete while the business result remains unresolved. None of those finish lines is wrong; each is complete at its own level and incomplete from the level above. That makes "done" less a single technical state than a matter of perspective -- what work is being measured, who owns the result and how far downstream success has to travel before the enterprise calls the job finished.
A completed task can still produce an unfinished result
Governance creates another version of the same problem.
Accuracy, latency, drift and task completion can tell an enterprise how the technology performed. They don't necessarily tell it whether the AI helped the business.
That doesn't make task completion meaningless. If an agent was supposed to draft a response and drafted it correctly, it did its job. But its job might be one small cog in a much bigger wheel.
The customer still needs the right answer. The claim still needs to be settled correctly. The transaction still has to post.
The agent can finish its job without finishing the enterprise's job.
Then somebody puts a price on "done"
Outcomes-based AI pricing sounds appealing because the customer presumably pays for what the technology accomplishes rather than merely for access or consumption.
Except buyer and seller first have to agree on what was accomplished.
Even contact center vendors can disagree over something as basic as what constitutes a resolution. That makes standardized comparisons difficult before anybody even gets to the invoice.
Now "done" isn't merely a technical or operational definition.
It's a commercial one.
In this TechTarget video, Sutherland's Jim Dwyer explains why proving that AI works is not the same as proving business value. The harder question is defining success and measuring what happens downstream.
Who gets to decide when the work is done?
This may be the part CIOs need to pay the most attention to.
Is the work done when the AI says it completed its task? When IT sees a workflow finish successfully? When the manager who owns the process says it's done?
Or is it done when the business leader accountable for the result says the outcome actually happened?
There probably isn't one universal answer. Nor should there be.
A task, workflow and business result can each be complete on their own terms. The mistake is assuming that completion at one level automatically proves completion at the next.
Enterprises are increasingly trying to determine whether AI is cheaper, more productive, safer, more valuable and even worth paying for based on outcomes.
Before CIOs can answer any of those questions, somebody has to decide what "done" means.
And that somebody probably shouldn't be the AI.
James Alan Miller is a veteran technology editor and writer and Lead Editor for CIO News at Informa TechTarget. He directs coverage of enterprise technology strategy, AI, software, data, infrastructure and the decisions shaping how CIOs manage increasingly complex IT environments.