Getty Images
Why AI should build on, not copy, your data stack
The model might be the face of AI, but the data architecture performs the heavy lifting. Its strength determines the consistency and reliability of AI-generated results at scale.
Many enterprises treat the "AI stack" as a new frontier alongside the data platform they already run. It's a comforting story, but it's wrong.
Models, vector databases, orchestration layers and AI agents might require new capabilities, but they still depend on enterprise data. They inherit every strength and weakness of the architecture that supplies, defines and governs that data. Your data architecture is now your AI architecture, whether you designed it that way or not.
The model is not the moat
The models are the most visible part of AI and the least durable. They commoditize monthly; the capabilities of today's frontier model can become tomorrow's baseline. What does not commoditize is the architecture that decides what data the model can access, whether that data means the same thing across the business, and whether anyone can trust it.
Two companies can license the same model and get wildly different results. One feeds it governed, well-defined, reusable data; the other feeds it whatever it can find.
The model alone is not the differentiator. The data architecture behind it is.
Deployment is outpacing discipline
A July 2026 survey by Omdia, a division of Informa TechTarget, of 400 organizations found that 39% were running vector databases or vectorization processes at production scale. Meanwhile, 47% had deployed governed semantic layers in some domains, but only 37% had deployed them across the enterprise.
Read those numbers carefully: many organizations have established governed semantic layers in parts of the business, but far fewer have extended them across the enterprise. A vector database retrieves what you embed. It has no opinion about whether "revenue" or "active customer" means the same thing across two systems. Without a semantic layer to enforce shared meaning, you do not get smarter AI. You get inconsistent answers, delivered with confidence, at scale.
The foundation that traditional analytics and AI share
Here is the reframe for data and analytics leaders: you have already been building most of what AI needs.
The semantic layer that gives analytics consistent definitions can also reduce the risk of AI misinterpreting ambiguous terms. Likewise, the reusable data products that feed dashboards can provide the same governed, trustworthy supply for AI models.
The metadata, lineage and governance that make analytics auditable also help make AI-supported decisions defensible. These are not separate analytics capabilities or AI capabilities. They are a shared foundation, and every dollar spent improving them can benefit both.
From data architecture to intelligence architecture
The deeper shift is that data architecture is no longer just plumbing for moving and storing data. It is becoming the enterprise's intelligence architecture.
These capabilities do not curb hallucinations because they are AI technologies. They help by reducing ambiguity in the underlying data.
And, unlike a one-time project, the value of that capacity can compound. Every reusable definition lowers the cost of the next AI use case, every data product shortens the next build and every governed relationship deepens what the organization can learn.
This is why a well-built data architecture does not need to lose value over time, but can grow more beneficial with ongoing maintenance and reuse. The architecture is not just supporting AI; it is helping determine how far your organizational intelligence can scale.
Where AI genuinely needs something different
Some components are genuinely new and are worth the investment.
AI works extensively with unstructured content -- documents, images and transcripts -- that many traditional pipelines were not designed to handle.
AI architecture often needs vector storage and embeddings for semantic retrieval, pipelines to chunk and enrich that content, and infrastructure to serve models in production. Those are justified additions. The mistake is treating them as a reason to build a parallel universe, creating a second data environment with its own copies and controls outside the organization's established governance.
That path duplicates data, fragments ownership and reintroduces problems that governance was built to solve: reconciling multiple versions of the truth and tracing decisions back to a reliable source.
Add the machinery, keep the meaning
The goal is to extend your existing architecture by adding specialized AI components where justified on top of the semantic layers, data products and governance that already supply the enterprise.
The test is simple: if a capability defines, governs or gives meaning to data, it belongs to the shared foundation and should serve reporting, dashboards, traditional analytics and AI. If it only stores, retrieves or serves -- vector indexes, embedding generation and model endpoints -- it can be AI-specific. Meaning is shared. The mechanics can be specialized.
AI that runs on your governed foundation can draw on established trust. AI that runs beside it, on copied data, must earn that trust separately, usually only after something breaks.
So before you fund the next AI platform, ask one question: Are you building a genuinely new architecture or an expensive duplicate of the data architecture you already have?
Bill Schmarzo, "The Dean of Big Data," teaches AI-driven innovation at Iowa State University and advises organizations on data science, AI and data monetization. He is a former executive at Dell Technologies, Hitachi Vantara and Yahoo. He has written books on data-driven innovation and applied AI strategy.