Getty Images/iStockphoto

How to prepare enterprise data for advanced AI

AI projects need reliable data, but it's not necessary to prepare all assets at once. Focus the improvements on the structured and unstructured data each use case requires.

Executives have praised data as the "new oil" for years, and its value has taken on new significance in the AI era. But relatively few organizations have their data ready for modern AI use.

In a May 2026 research report from professional services firm Accenture that surveyed executives at 2,000 companies, just 7% of respondents said they had achieved the degree of data readiness necessary to scale advanced AI, which includes generative, agentic and physical AI.

In another report from software maker Cloudera and Harvard Business Review Analytic Services that surveyed more than 230 members of the Harvard Business Review audience, only 7% believed their data was ready for use with AI.

Both reports also indicated that most organizations struggle with data readiness, but in different ways. In Accenture's report, 72% of respondents lacked both reliable data of the right quality and standardized governance procedures. In Cloudera's study, 73% of respondents said that processing and preparing data for AI was a challenge.

Together, the findings show organizations continue to have problems with data quality, preparation and the governance practices that can affect the reliability of outputs from their AI systems. That's because many organizations are moving into the era of GenAI, retrieval-augmented generation (RAG) and agentic AI with data foundations built for earlier machine learning and analytics capabilities. For example, they may have the structured data quality and precise labeling that work well for predictive analytics. However, they may be missing high-quality unstructured data or a semantic layer that provides business context, ensuring that advanced AI deployments operate accurately and securely at scale.

Without all these pieces in place, AI deployments are hamstrung. More than 80% of organizations have paused, downsized or adjusted AI initiatives because of data-related risks, according to Accenture's research.

"AI-ready data has to be governed, traced, protected and used properly because these are the crucial inputs in decisions," said R "Ray" Wang, principal analyst and founder of Constellation Research.

What makes data ready for use with AI

Modern AI, especially GenAI and agentic AI, usually needs a higher level of data readiness than older enterprise AI projects. That shift helps explain why so many organizations lack the data maturity needed for large-scale GenAI, RAG and agentic use cases despite a decade-plus of investments in their data environments.

The types of data involved also affect readiness. Modern AI relies on both structured and unstructured data. While many organizations have focused on improving the quality, integrity and accessibility of structured data, unstructured data often comes from diverse sources and exists in formats, such as text, audio and images.

While requirements vary by use case, data preparation for AI commonly focuses on the following criteria for structured and unstructured data.

Data quality

AI-ready data should have a high degree of accuracy, completeness, consistency, timeliness and uniqueness, free of unwanted duplicates.

"For organizations, a lot of times today, all the data is still not mastered, still not cleansed," said Pradeep Suryanarayan, chief solutions officer at UST, a digital services company. "Data quality is still a problem in many organizations."

Semantic structure and context

Data also needs rich metadata, taxonomies and consistent business definitions. Depending on the use case, a semantic layer can provide common definitions for data and metrics and show relationships among data elements across multiple data sources.

"That layer is often missing today," Suryanarayan said.

Data assurance

Governance, security and policy controls can limit access to authorized users and approved uses, reducing the risk of hacks, data leaks and unintentional exposure. Organizations should also include processes to monitor adherence to standards and guardrails, and maintain metadata and logs to understand why something did not go as planned.

"With monitoring, if something went wrong, you can tell what happened and when it happened so you can determine how to make sure it doesn't happen again," said Alan Cecil, data analytics manager at professional services firm BPM.

Data movement

Traceability requires practices that track data lineage throughout its entire lifecycle, from creation to archival. For AI applications, organizations may also need to monitor how data is transformed, accessed or utilized by a model or application.

"You have to imagine that your data is like your water supply or energy supply. Every step needs to be accounted for," Wang said.

Prepare the right data, not all data

Accenture's research associated stronger AI-ready data capabilities with greater financial and non-financial value, including improvements in productivity, decision quality, customer experience, risk reduction and margins.

But organizations don't need to have all their data prepared for advanced AI to move forward, Cecil said.

"Don't think of AI readiness as an all-or-nothing situation," he advised. "It's too big to think you're going to get all your data AI ready."

In fact, he said that approach may mean "doing work you don't have to do."

Rather, the expert consensus is to take a phased approach, prioritizing advanced AI use cases by impact and then improving the integrity of the data required to support each use case.

"What [are] the most important processes that you could start to support with AI and then build up the readiness for the relevant data?" Cecil said.

Mary K. Pratt is an award-winning freelance journalist with a focus on covering enterprise IT and cybersecurity management.

Dig Deeper on Data Management