In his new book titled "Designing the AI-Driven Data Foundations: Architecture, Principles, and Practice," veteran analyst Sanjeev Mohan, who headed Gartner's data and analytics research team before founding SanjMo in 2021, lays out the critical nature of a proper data foundation for AI and how organizations can build that bedrock.
AI, he states, is fundamentally just the most modern use case for data.
For decades, data in the enterprise was kept under wraps by specialists and largely used to inform business intelligence tools such as reports and dashboards. In recent years, the paradigm has evolved. Now, data is earmarked to underpin agents and other AI tools, and it is becoming the infrastructure for widespread AI-driven analysis and process automation.
In the coming years, data will be the foundation for technologies such as robotics and quantum computing.
But just as BI without reliable data couldn't be used to inform decisions, AI without a strong data foundation can't be trusted to generate insights and run business processes.
In a recent interview with TechTarget, Mohan discussed the essential nature of a strong data foundation for AI, what makes a good data foundation and how AI can play a role in preparing data for AI as well as benefitting from it.
In addition, he spoke about data governance's role in the evolving AI workflow, the barriers that still prevent some enterprises from building a strong data foundation and how data for AI will continue to evolve.
Editor's note: This Q&A has been edited for clarity and conciseness.
Before we get into the details of what makes a good data foundation for AI, tell us why a strong data foundation is so important for AI.
Sanjeev Mohan: When I watch vendors demo new capabilities these days, they all open a natural-language interface, they ask a question, and then an answer magically appears. But that's a fallacy. It makes it look like AI has overpowering capabilities that wash over the cesspool of data that [organizations] have collected.
Your AI is only as good as your data, which is easier said than done. Data doesn't have to be pristine [for AI], but you at least have to know what you have.
Why don't many organizations know what they have and instead have a cesspool of data?
Mohan: Data estates are exploding at their seams. Most organizations today have over 100 SaaS products. All of them collect data, but they don't share it because there is no mechanism [to connect them], and there are no standards to make it consistent across applications. Instead, there are all these SaaS products with local copies of data -- maybe some gets moved into data warehouses -- leading to too many data siloes.
I'm a very strong believer that if we can fix the data foundation, we can derive the success from AI that has so far eluded us.
Sanjeev MohanFounder and principal analyst, SanjMo
In organizations, sales and marketing should be connected at the hip. Instead, sales has its ecosystem of data, and marketing has its ecosystem of data. As a human consumer in a senior position, if I need to know how marketing is leading to sales, I can connect them. But if AI agents are analyzing data, how in the world will an agent know that 'X' in one system is the same as 'Y' in another system? There are no standards, and the two are coded differently.
I'm a very strong believer that if we can fix the data foundation, we can derive the success from AI that has so far eluded us.
Fixing the data foundation to get value from AI seems obvious, so what are you experiencing when you talk to organizations about doing so?
Mohan: If I talk too much about data, people start shaking their heads and say, 'What are you talking to us about data? Talk to us about AI.' When I tell them that AI is a use case of data, some take offense. To them, AI is a whole, gorgeous, brand-new world. But it's a use of data. Tomorrow it will be robotics, then it will be quantum computing, and then it will be something we haven't even thought about.
All of these things are sitting on top of a substrate, and that substrate is a data foundation.
Before the AI era began a few years ago, what did organizations need to do to build a proper data foundation?
Mohan: That process was determined by technology. When computers came in during the 1960s, 70s and 80s, memory was very expensive, so when networks arose, they were very slow. Data foundations evolved from that era. It was almost always structured data,it was all batch file processing, there was no concept of real time and there were a lot of manual checks with humans in the loop. Automation was something only enabled by the cloud.
How has AI changed what the data foundation looks like -- how is it different than it was just a few years ago?
Mohan: With AI, there is an amazing ability to detect similarities across any data source. Now, I can say, 'Give me information structured sources such as databases and CRM applications and do sentiment analysis on customer support calls for a given customer.' In one query, AI can join audio transcripts, structured data and sentiment analysis. There is data [that informs] AI, and AI [that helps manage] data. AI plays two roles. It fixes the data foundation, and it helps derive new insights from that data.
If you were advising an enterprise, how would you tell them to build a data foundation for AI -- what tools do they need from start to finish to have a proper data and AI stack?
Mohan: I tell them to first pick one business process that they're trying to turn into an agent, then lay out the data floor from ingestion, through integration and transformation, to observability and reporting.
Once they understand their data process, they can look to see where they can insert AI into the process, such as at the ingestion level. And once they understand the process, they can turn their natural language documentation into a skill.
The first skill becomes ingestion so they can instruct their AI to ingest, classify, profile and find anomalies. Others are data integration, data quality, and so on, but start with ingestion and automate it, and have a human in the loop to check if the data was ingested correctly and whether there were duplicates.
In AI, the most important thing is evaluations, and what you're doing in this process is creating relevant context for large language models (LLMs) to act on your behalf.
What does an LLM acting on a user's behalf look like?
Mohan: The final thing is to be able to tell your application something such as, 'Build me a dashboard for SEC filings,' or, 'Do my taxes.' That combines all the other skills. To do my taxes, first the LLM needs to get my 1099s and W2s -- that's ingestion and integration. Then a human needs to cross-check to make sure it was correct. Ingestion is not an outcome. Integration is not an outcome. These are means to an end.
AI should be considered a new paradigm, a new way of doing things. In the past, humans built their own reports but were limited. Now, AI can access more information.
Where do capabilities such as data governance enforcement fit in now?
Mohan: People ask how to go from data governance to AI governance, and I tell them that there is a problem in how they're asking the question. You're not going from one to the other. You're upscaling from data to AI governance. You still need all the lineage, business glossaries and catalogs at the data level. But now, with AI, everything changes.
The idea of data governance was to govern source data and build deterministic applications on top that would always run the same way. AI breaks that thinking because AI is probabilistic. Even with highly governed data, AI can misbehave and provide wrong answers. Now, you have to govern the output.
That's why humans need to be in the loop. You have to govern your input, which is data governance. You have to govern your output, which is AI governance. And with agents, you have to govern all the things that are happening in a multi-agent systems. Data governance is still needed, but you have to upskill agents so that before something bad happens, it can be stopped.
As you interact with organizations, are you seeing that most now have their data in order for AI, or do many still not have proper data foundations?
Mohan: It's getting better.
What I mean by that is that when I started my career, all the data was proprietary. SQL was the open thing, but everyone had their own flavor, so it wasn't really open. Today, there are open standards throughout the data stack, such as object store with Apache Parquet files and Apache Iceberg as a table format. Even vendors such as Snowflake with proprietary formats are embracing Iceberg. That allows users in Snowflake to directly see what's going on in SAP, which is being published to the outside world in Iceberg. Now, you also have standards for semantic layers with the Open Semantic Interchange, for agents with Model Context Protocol.
We've broken free of proprietary technologies, and for the first time, we are seeing standards span the entire stack. … The data foundation is a living, breathing organism where cells die and new ones generate.
What are some of the biggest hurdles preventing organizations from building a data foundation for AI?
Mohan: The hurdles that organizations face are in keeping up with technology in a manner in which their investment is safeguarded, and that is very difficult. For example, 15 years ago, we thought Hadoop was a savior. So, now we ask how to know whether agents will be successful. Right now, they're running haywire. Knowing what to do is a big challenge.
A second challenge is how to do it. Everyone is talking about context, but how many are talking about how to build that context? Two people building context might have very different opinions about what context is. When you haven't even agreed upon a definition of context, how can you build it. The idea is great, but we don't know how to do it yet.
Looking into the future, how will the data foundation evolve, or do you think that AI's current data requirements will remain somewhat static over the next few years?
Mohan: I think the data foundation will evolve tremendously.
In my book, I introduce a concept which should be the next version of the extract, transform and load process (ETL). I call it ECL -- extract, context, link. ETL was great for batch jobs, for moving data. But why are we moving data? If you tell AI where the data is, where the context is for a query, let the AI go where the data is. You don't want to move data unnecessarily because it's expensive, it opens up security holes and there's latency. What changes is that you need access to contextual metadata, and your sources are [unstructured data]. … And you need the context to be linked in a knowledge graph or some other way. This is how AI changes the ETL paradigm into ECL where context is created in real time. But there is work to be done. All this tooling still has to be developed.
A lot of enterprises are still trying to solve that context layer.
Beyond figuring out context, is there anything you want enterprises to be able to do to build a data foundation that still isn't possible?
Mohan: I was very hopeful that by now, four years into this, there would be an ability for organizations to train their own models, but we haven't gotten there. And so far, it doesn't look close.
There are frontier models that are trained on common data, but enterprise data is locked away. That's why we started doing retrieval-augmented generation and other experiments. Eventually, every company might have its own model that is constantly refreshed, and that might change things again.
Eric Avidon is a senior news writer for Informa TechTarget and a journalist with more than three decades of experience. He covers analytics and data management.