Getty Images/iStockphoto
Why more data will not deliver AI data readiness
Volume-first data strategy served the BI era but falls short for agentic AI, where the constraint on readiness is context and governance rather than how much data exists.
For the better part of two decades, enterprise data strategy ran on a single assumption: more data meant better outcomes. AI has broken that assumption, and the organizations discovering it are learning that the bottleneck isn't volume, but context.
The accumulation strategy was rational for the workloads it served. Enterprises consolidated data lakes, scaled warehouses and ingested everything available, because BI consumed structured data in predictable patterns. Having more data in one place made reporting faster and dashboards richer.
That logic has started to break down. As enterprises move from BI-era workloads to agentic AI, the question is no longer how much data they have, but whether that data carries the context, definitions and governance required for an AI agent to use it correctly. The gap between what enterprises have built and what AI requires is widening, and closing it demands a fundamentally different kind of investment.
The workload that broke the old model
Traditional BI workloads were structured, predictable and scoped. An analyst queried a warehouse, built a dashboard and refined a report. The data was organized for human consumption, and the questions it answered were framed by humans.
AI workloads do not operate in the same way. Retrieval-based systems, including the RAG architectures underpinning most enterprise agents, assemble their context at query time, pulling from multiple sources, interpreting definitions on the fly and executing tasks without a human framing each query in advance. What the system retrieves in the moment determines the quality of what it produces. And unlike a human analyst who can recognize when a metric looks wrong, an agent will use what it is given.
The difference has made context the highest-returning investment in the data stack. Felix van de Maele, CEO of Collibra, pointed to evidence from Anthropic's own self-service analytics work. Without business context, Anthropic's agents achieved roughly 20-25% accuracy on analytical tasks, van de Maele said. With governed business context layered in, accuracy climbed to 95%.
The numbers suggest that improving the context fed to a model now produces a larger performance gain than improving the model itself or expanding the data it can access. For organizations that spent years investing in volume, the implication is uncomfortable. The constraint they optimized for is not what matters most for their current needs.
How the gap became structural
An important question to consider is why context and governance were neglected in the first place. Enterprises were not careless; the way organizations fund and deliver technology projects made reuse and shared semantic foundations structurally difficult to build.
Terry Dorsey, senior data architect at Denodo, spent much of her career in industry roles before moving to the vendor side. She described an environment in which every data initiative was scoped and funded as a standalone project, with resources tied to specific deliverables and no mechanism for cross-project reuse.
"A lot of it has to do with how organizations structure: how they do funding, who gets funding and how do you use that funding," Dorsey said. Each new initiative rebuilt its own data foundations from scratch, even when an adjacent team had already done similar work. Over time, the pattern produced overlapping silos, duplicated logic and mounting technical debt that individual project budgets were never structured to address.
The technology changed across eras, but the organizational pattern did not. Dorsey noted that at a recent industry conference, practitioners were already discussing agent sprawl as the latest iteration of a familiar cycle. Report sprawl became data sprawl, and data sprawl is becoming agent sprawl. The same project-scoped thinking that created redundant dashboards and disconnected data pipelines is now producing redundant agents with no shared governance layer.
"Fundamentally, all we've done is move data from one place to another," Dorsey said. "You can get at it faster, but we haven't worked on how we actually deliver it to people. And we keep the same processes we've had for the past 20-plus years."
The cost of the gap is visible
For years, the absence of shared semantic foundations was a latent inefficiency. AI has made it an active cost.
Van de Maele described a "hallucination tax" that enterprises are paying when agents operate without governed context. When AI output cannot be trusted, organizations default to human-in-the-loop verification, double-checking every result. That verification bottleneck erodes the productivity gains that justified the AI investment in the first place.
The cost extends beyond labor. Agents consuming ungoverned data burn tokens processing irrelevant or duplicative information, driving up compute expenses without improving outcomes. Van de Maele said token efficiency has become a key metric alongside task completion because throwing more unstructured data at an agent makes the operation more expensive and degrades accuracy.
The prototype-to-production bottleneck tells the same story from another angle. Van de Maele said getting agents from prototype to production "is turning out to be a lot harder than people expected," and attributed the difficulty to governance gaps rather than technical limitations. The models work, but organizations are still building the context that makes them work reliably in production.
Dorsey offered a practitioner's version of the same observation. Enterprises often believe that hiring skilled AI researchers would be sufficient to operationalize AI. Researchers could build models in weeks, but the surrounding work turned out to be the far larger component. Governance had to be established for what data an agent could expose. Business processes needed to change, and the organizations had to be brought along. The development was the smaller part of a much bigger effort that the organization had not anticipated.
What the response looks like
No single playbook exists for closing the gap between what enterprises have built and what AI needs. Organizations are approaching the problem from different starting points, and the strategies reflect those differences.
Some enterprises are beginning with discovery and classification, according to van de Maele. They inventory their data, tag and classify what they have and build a semantic structure before deploying AI against it. The emphasis is on understanding the estate before asking agents to operate within it.
Others are working backward from a specific AI use case. Rather than cataloging everything, they identify a high-value workflow, build a curated knowledge base to support it and expand outward from there. This approach is narrower and faster to deliver results, though it risks recreating the project-scoped pattern Dorsey described if the knowledge base remains siloed within a single initiative.
What both strategies share is a recognition that the next investment must go somewhere different. The volume era asked how much data an enterprise could bring together. What matters now is whether the data already in hand carries the context, definitions and governance to be trusted by the systems consuming it. That shift, not the failure of the old strategy, is what the next investment has to answer.
Scott Thompson is the Site Editor for TechTarget's Data Technologies group, covering data management and business analytics topics for senior enterprise data leaders. He has edited data and analytics content for TechTarget since 2021.