Getty Images/iStockphoto

Why data governance struggles to audit shadow AI

Data loss prevention watches known exits, but shadow AI moves data through agents and SaaS integrations. Gateways, local models and sanctioned AI tools close that gap.

Data governance can't control what it can't see.

Shadow IT bypassed the visibility, accountability and formal controls required for effective governance. Shadow AI presents a similar problem through the ad hoc use of unapproved tools, which can expose enterprise data through activity that conventional governance might not detect. Organizations typically treat it as a data loss prevention issue, but that framing neither captures nor addresses the larger problem. 

Each instance of shadow AI is an act of unapproved context assembly: an employee or agent gathers documents, records or other enterprise data and feeds them to an external AI tool to shape its output. It can bypass governance checks for enterprise data, including whether "it was classified correctly, whether access policies were honored, where the data went, whether it was retained by a third party, and how it influenced the final output," said Nitika Gupta, product leader for security at Snowflake.

Consequently, these governance failures can expose the organization to data breaches, data loss, regulatory noncompliance, inaccurate model outputs and unauthorized agent actions. Examining where existing controls fail and how enterprises have responded shows which approaches can minimize the frequency and severity of these incidents.

Why existing controls miss shadow AI

Several characteristics of shadow AI make it hard to govern with traditional controls, particularly when it's used for context assembly. Conventional governance assumes that organizations can monitor users and their applications. But when an employee uses an unapproved bot or language model outside an organization's purview, "there's no logging or telemetry that goes to the [security information and event management system] that's alerting users, or agents, about it," said Sean Roche, director of value engineering and product marketing at Obsidian Security.

Even controls built for data in transit tend to watch known channels such as email, downloads and endpoints.

With AI agents, however, "data's in motion now, and it connects through OAuth or SaaS integrations," Roche said. "[Data loss prevention] was never made for that. It's looking for chat; it's looking for a download; it's looking for a computer. So, it misses a lot of what happens over the open internet."

The distributed nature of shadow AI also undermines existing governance methods, particularly across the many SaaS applications that run language models, agents and retrieval-augmented generation (RAG). When those applications get access to data in frequently used systems, such as Salesforce, "You need protection in the SaaS app to see that, or you never will," Roche said. "There's no network proxy that's in front of that."

Leakage is only the first failure

Data leakage is often the first clear sign of shadow AI escaping existing controls and increases an organization's risk of a data breach. That risk is particularly acute with agents "because once AI agents become embedded into workflows and decision making, the attack surface expands very quickly," Gupta said.

Agents can also communicate or take actions as though they represent the organization, effectively allowing them to "impersonate the business," Roche said, adding that some of the consequences could result in lawsuits, bad customer service or loss of revenue.

Because shadow AI moves, transforms and uses data in external settings, it also raises data sovereignty questions, including where language models are hosted to where data resides, said Luis Flynn, AI product marketing strategist at SAS.

Shadow AI governance requires more than tools

Organizations have responded to shadow AI with both traditional and newer controls.

"Many enterprises are trying to detect sensitive data leaving the organization, restrict uploads to unsanctioned AI tools, monitor browser or SaaS activity and block the use of certain external services," Gupta said.  

Many organizations focus on shadow AI governance at the prompt level. Tools can dynamically parse prompts and model responses in RAG frameworks to detect, obfuscate or redact sensitive information in real time. But as the industry struggles to standardize its approach to shadow AI, organizations often turn to point tools, such as prompt governance, rather than addressing the problem at the policy level.

It's not uncommon for "people to veer off from the general approach to governance and address it at the tech level and not at the conceptual level," said Terry Dorsey, senior data architect at Denodo.

Such efforts can provide some short-term utility, but they are difficult to sustain and scale. They can also contribute to new data silos when each one handles access control and governance along a single path rather than applying the organization's broader information policies, Dorsey said.

How to govern shadow AI

One way to reshape data governance for shadow AI at the policy level is to rearchitect the data estate to "bring your capabilities closer to your data, instead of the other way around," Flynn said.  

Organizations might, for example, run language models, including open source ones, locally so sensitive data never leaves the organization.

Other approaches include gateways that enforce policies where API calls are made and tools that monitor third-party SaaS applications for user, permission and data governance policy violations. Both also apply to AI agents. The more durable answer is to identify and approve appropriate AI resources already in use and bring them under governance.

Jelani Harper is a data industry analyst and journalist covering data management, AI and enterprise IT for more than a decade.

He is a research lead at Blue Badge Insights, writing analyst reports for GigaOm and articles for VentureBeat and The New Stack.   

Dig Deeper on Data Management