Faster drafts, more code and higher adoption can look like AI productivity. David Linthicum argues CIOs should measure end-to-end outcomes, quality and hidden rework instead.
A large financial services company that I worked with a few years ago rolled out generative AI tools across its knowledge workers. The early metrics looked terrific. Employees were drafting emails faster. Analysts were producing summaries in minutes instead of hours. Developers were generating more code. Customer service teams were closing more tickets per shift. The dashboard was green, the vendor case study was almost writing itself and the executive team was ready to declare victory.
Then someone (meaning me) asked the question that should have been asked on day one: Are we making better decisions?
The room got quiet.
The company had more output, but not necessarily better outcomes. More documents were being produced, but the quality varied wildly. Some of it was AI slop. More tickets were being closed, but escalations were rising. More code was being generated, but senior engineers were spending more time reviewing subtle defects. Analysts were creating more summaries, but business leaders still did not trust them enough to act without rechecking the source material.
That is the problem with how enterprises are measuring AI productivity. They are confusing activity with value. I have spent 35 years watching companies do this with every major technology wave, from client-server to web, cloud, big data and now AI. We take the easiest thing to count, put it on a dashboard and pretend we are measuring business impact.
We are not. Stay with me here.
The productivity trap
Many AI productivity metrics focus on speed and volume. Time saved. Documents produced. Tickets resolved. Lines of code generated. Meetings summarized. Emails drafted. These measures aren't useless, but they are dangerously incomplete. Moreover, this is how many AI providers are selling AI today.
If an AI tool helps a claims processor handle 40% more cases per day, that sounds like productivity. But if error rates rise, appeals increase, customer satisfaction falls or managers must add a second review layer, the enterprise did not gain productivity. It moved work from one part of the system to another and congratulated itself for the transfer.
This is the same mistake enterprises made with cloud cost metrics. They measured server reduction and ignored architecture complexity. They measured migration progress and ignored application modernization. Now they are measuring AI usage and ignoring decision quality, and that's everything.
Here is my blunt take: many AI productivity dashboards are executive comfort food. They make leaders feel progress is happening because the numbers are easy to collect and easy to present. But if a metric doesn't connect to quality, risk, revenue, cost or customer outcomes, it is not a business metric. It is telemetry.
If a metric doesn't connect to quality, risk, revenue, cost, or customer outcomes, it is not a business metric. It is telemetry.
More output is not the same as better work
The first illusion is that more output means better work. AI is very good at increasing output. That is the easy part. It can quickly produce a first draft, a summary, a recommendation, a query, a test case, a slide or a service response. Enterprises see that speed and assume value.
Sometimes it has. Often it has not.
A mediocre report produced in 10 minutes is still mediocre. A flawed software module generated quickly still has to be debugged, secured, tested and maintained. A customer response generated instantly can still be wrong, tone-deaf or legally risky. The enterprise isn't paid for producing more artifacts. It gets paid for better outcomes.
This matters because AI often shifts work into hidden review loops. The person using the tool may save time, while the expert downstream spends more time validating, correcting and explaining. The productivity gain shows up in one department. The cost appears somewhere else. If CIOs do not measure the full workflow, they will approve systems that make one group look efficient while making the enterprise less effective.
I have seen this pattern repeatedly. An AI pilot can show a substantial productivity lift because it measures task completion time. Production shows much less value because nobody measured rework, exception handling, audit exposure or decision latency. The happy path looks great. The enterprise lives in the exceptions.
Ford offers a useful example. The automaker acknowledged that it had overestimated what AI and automated quality systems could deliver without experienced human judgment. As part of a broader quality overhaul, Ford brought in more than 300 experienced engineers -- including former employees and people from suppliers -- to strengthen its quality process, mentor younger workers and help improve the AI systems themselves.
The lesson was not that AI had no value. It was that automation could not replace the expertise the company needed to make the technology work.
The real test is the business outcome
The metric CIOs should care about is not simply whether AI made workers faster. It is whether that speed improved an end-to-end business outcome without creating unacceptable risk.
That changes the measurement conversation. For an AI sales tool, do not just measure emails generated or leads contacted. Measure conversion quality, deal velocity, customer fit and downstream churn. For an AI coding assistant, do not just measure code volume. Measure defect density, security findings, maintainability and cycle time from request to stable production release. For an AI support assistant, do not just measure tickets closed. Measure first contact resolution, escalation rate, customer satisfaction and repeat contacts.
This is not complicated, but it is inconvenient. It forces leaders to admit that AI value is contextual. It also forces them to measure across organizational boundaries, which is where enterprises usually fail. The CIO may own the platform, but the business owns the outcome. If both sides are not looking at the same scorecard, AI measurement becomes another political exercise.
Vendors are not going to fix this for you
The vendor community has a strong incentive to sell simple productivity stories. Time saved is easy to understand. Usage is easy to show. Adoption curves look good in quarterly business reviews. Nobody wants to walk into the boardroom and explain that the real value of AI depends on redesigning workflows, changing incentives and measuring outcome quality over several quarters.
But that is the truth.
I am not saying vendors are dishonest. I am saying their metrics are often optimized for selling software, not running enterprises. If a vendor tells you its tool saves employees seven hours per week, ask what happens to those hours. Are they converted into revenue? Better service? Faster decisions? Lower risk? Or are they absorbed into more meetings, more drafts, more internal chatter and more rework?
Time saved is only valuable if the enterprise captures it. Most do not. They just create more work, faster.
The most overhyped metric in enterprise AI is adoption. High adoption tells you people are using the tool. It does not tell you whether the tool is improving the business. Plenty of bad systems get used because employees are told to use them, because they are embedded in existing software or because people enjoy experimenting with them.
The second overhyped metric is prompt volume. Counting prompts is like counting keystrokes. It may tell you something about usage, but almost nothing about value.
The third is generic time savings. I distrust any AI business case built mainly on assumed hours saved. That number is usually soft, self-reported and disconnected from financial outcomes.
What matters is outcome-based measurement. Did the decision improve? Did the process shorten end-to-end? Did risk go down? Did quality improve? Did customers get a better result? Did the business actually capture the savings? If not, the productivity story is mostly theater.
David Linthicum is a globally recognized thought leader, innovator and influencer in AI, cloud computing and cybersecurity. He has more than 30 years of experience in enterprise technology and, until early 2024, served as managing director and chief cloud strategy officer at Deloitte Consulting LLP.