Getty Images

Security, cost benefits drive the rise of bringing AI to data

Natively integrating models where data lives has proven safer, more cost efficient and effective than porting data into LLMs when building agents and other AI tools.

Pipelines that port data into large language models are going the way of the dinosaur when it comes to AI development. Instead, bringing AI models to data in the secure, stable environments where data is created and stored is proving to be more effective, efficient and safe.

AI needs relevant data and business logic -- and lots of it -- to be accurate.

In the initial rush to build generative AI tools after OpenAI's November 2022 launch of ChatGPT, developers constructed complex retrieval-augmented generation (RAG) pipelines to find relevant context, including unstructured data, and deliver it to large language models (LLMs). Those pipelines involved exporting data into specialized vector databases, connecting vector databases to LLMs so the models could discover data through similarity searches, and enabling the LLMs to pull in relevant data to inform their outputs.

Amid the excitement, the security risks of moving data to public models and high cost of using LLMs were secondary to the potential business benefits of GenAI, according to Sanjeev Mohan, founder and principal analyst at SanjMo.

"The best models were only available as hosted APIs, so data had to travel to them," he told TechTarget. "Data platforms lacked … vector search at the time, and urgency won. Teams wanted GenAI demos fast, and copying data into a vector database was the quickest path. Governance came second."

However, a confluence of factors led enterprises to reconsider how they built AI workflows.

RAG pipelines were not effectively delivering the relevant context required to make AI tools accurate enough to put into production, the rise of autonomous agents necessitated not only higher degrees of accuracy but also real-time data, and many data platform providers were building vector database capabilities into their broader data management offerings. Meanwhile, with enterprises getting little return on their AI investments, cost control and security concerns became more prominent.

Consequently, expensive and risky ETL pipelines are being phased out in favor of bringing AI to the data instead.

"The market evolved in new directions," Kevin Petrie, an analyst at BARC U.S., told TechTarget, citing the evolution of data platforms, growing ability of LLMs to navigate unstructured data without vectorized similarity searches and rising security and regulatory concerns. "All these forces make it more appealing for organizations to bring AI to their data."

Removing the pipelines, preserving context

Bringing AI to data means locating capabilities such as model inferencing, vector search and semantic logic within databases and other data management platforms.

Critically, by keeping data where it gets created and stored, all of the governance, lineage and business context that gets applied to data in its native environment through tools such as data catalogs and semantic layers remains intact when it is joined with models to build agents and other AI tools.

"Bringing AI to data means running AI models and agents where a company's most important information is already stored," Craig Wiley, vice president of AI product at Databricks, told TechTarget. "By operating within a governed environment with existing security and access policies and lineage intact, you remove the need for any data migration or pipelines."

By contrast, when data gets copied into models through external pipelines, versions of data need to be created, and governance and business logic get lost along the way, he continued. As a result, developers and engineers need to reconstruct governance as security in each new environment, which is both time-consuming and costly.

In addition, Wiley noted that when data, Model Context Protocol servers and AI models are all in different environments with different governance and discovery controls, agents and users accrue spent tokens and wasted time searching for the right context.

"By centralizing discovery, access control and monitoring in a single platform, they accelerate development and reduce time and token costs," he said.

Once companies started asking hard questions about how to deploy AI on their customer data, at scale, without violating their own compliance policies, they realized the 'send your data out' approach wasn't going to work.
Pavan PothukuchiHead of AI platform, Snowflake

Databricks is now one of the many data management providers that enables users to bring AI to their data by turning its Unity Catalog a control plane for data, machine learning and AI and making models such as OpenAI's GPT-5 natively available in its platform alongside data. Rival Snowflake similarly allows users to bring AI to their data, with Cortex AI serving as a central location and models from Anthropic and Google natively integrated.

Among others, hyperscalers AWS, Google Cloud and Microsoft enable customers to bring AI to their data, as do platform vendors including Cloudera, MongoDB and Teradata.

Snowflake head of AI platform Pavan Pothukuchi noted that a critical advantage of bringing AI to data is that the context AI tools require remains intact, which improves the likelihood of an agent or other AI application avoiding hallucinations and performing as intended.

"With this architecture, AI can work directly with a company’s unique business information without requiring them to build a separate data environment for each AI application," he said.

The rush to build GenAI tools in 2023 resulted in enterprises first exporting data to vector databases and LLMs, which worked well in demos and experiments, Pothukuchi continued. But when it came time to put the applications into production, these projects no longer met most enterprise standards.

"Once companies started asking hard questions about how to deploy AI on their customer data, at scale, without violating their own compliance policies, they realized the 'send your data out' approach wasn't going to work," Pothukuchi said. "Building real, production-grade AI systems on proprietary data requires governance, access control and auditability, which can't easily be replicated in new tools."

Savings and security

While preserving context is one benefit of bringing AI to data, lower development costs and improved security are others.

Moving data into AI models requires building custom storage, governance and ETL pipelines for every AI application. In addition, it racks up data egress fees and generally requires using platforms from multiple providers.

"Multiplying that across the multiple AI tools and model providers that most enterprises use, costs compound fast," Wiley said.

But by locating vector search, model inferencing and semantic logic within a unified data management environment, those costs get eliminated, according to William McKnight, president of McKnight Consulting.

"Bringing AI to data is significantly less expensive," he told TechTarget.

In fact, McKnight noted that benchmark tests conducted by his firm showed that bringing AI to data cut the three-year total cost of ownership of AI tools in half when compared with fragmented pipelines that required data egress and ETL workloads.

In addition, testing showed that bringing AI to data accelerated development-to-production timelines by over 300%, reduced development complexity by 67%, lowered maintenance complexity by 38% and cut engineering personnel costs in half.

"The cost difference comes down to infrastructure that enterprises don't have to build," Pothukuchi said. "Moving data to AI means standing up pipelines, managing ETL jobs, paying cloud egress fees, and maintaining duplicate storage and access control policies. All of that is ongoing. … When AI runs where the data already sits, those costs disappear."

While costs disappear, security improves. Drastically.

"Each time AI models or agents consume data, you must maintain careful security controls such as encryption, masking and identity management to ensure you don't expose personally identifiable information (PII) to unauthorized users," Petrie said. "It is easier to maintain such controls when you bring AI to the data rather than vice versa."

In fact, while moving data across systems risks exposure at every step since governance has to be repeatedly re-applied, there is 100% perimeter containment when AI is brought to data in a governed data management environment, according to McKnight.

"Extracting proprietary data payloads across network boundaries expands the attack surface, exposes sensitive assets to potential breaches, forces fragmented multi-hop security configurations and compromises strict data sovereignty and regulatory compliance mandates," he said.

And just as benchmark testing showed the cost savings of bringing AI to data, tests conducted by McKnight Consulting quantify the security benefits of keeping data in place when developing and managing AI tools. Integrated, in-place platforms scored four times higher in data security than those that require data movement across external cloud endpoints, according to McKnight.

"Keeping enterprise datasets within native, governed database boundaries with automatic row-level security and encryption directly addresses the governance concerns that … data leaders cite as their primary barrier to AI adoption while eliminating the risks of data leaks, intercepted network payloads and regulatory non-compliance," he said.

Not the be-all-end-all

Although bringing AI to data is a substantial improvement over exporting data to AI models, it does not reduce all the problems preventing many enterprises from moving AI pilots into production.

Even when contextually relevant data and AI models are natively integrated in a secure data management environment, serving agents and other AI applications with appropriate context is challenging, according to Mohan.

He pointed out that AI tools need context from many systems, and even if some of those systems are on one platform, others likely aren't. To enable AI tools to access context spread across systems and platforms without moving data, federated access through governed interfaces such as MCP and zero-copy data sharing are required.

"The winning architecture is one stack with data as the foundation and AI on top," Mohan said. "The real question isn't where AI runs. It's whether it has trusted context when it does."

Pothukuchi similarly noted that while bringing AI to data improves context retrieval, it is imperative that enterprises not only make their data available to AI models but ensure that their data is consistent across all systems so it can be understood by agents and models. If two departments within the same organizations define a word such as "revenue" in different ways, an agent won't know which definition is applicable and will produce confident, but wrong, outputs.

"Bringing a model next to their data is incredibly valuable, but proximity isn't enough," Pothukuchi said. "Companies also have to make their business knowledge legible to AI, things like how metrics are defined, what the exceptions are, and which data sources are authoritative for which questions."

Meanwhile, an AI development and management system's architecture doesn't matter if the underlying data isn't correctly prepared, according to McKnight.

"Bringing AI directly to data solves the architectural tax, pipeline latency, and cloud egress costs, but bringing AI to dirty or uncurated data introduces a far more dangerous problem," he said.

Consequently, human oversight to prevent AI from acting on bad data and data observability to ensure data quality are also critical aspects of any AI workflow.

"The human-in-the-loop control plane and data observability strategy is important," McKnight said. "As enterprise AI evolves from looking up documents taking autonomous action, purely technical prompt engineering loses its competitive advantage to domain expertise and continuous governance."

Eric Avidon is a senior news writer for Informa TechTarget and a journalist with more than three decades of experience. He covers analytics and data management.

Dig Deeper on Data Management