Getty Images/iStockphoto
Where data teams fall short on governance for the EU AI Act
Most enterprises govern AI in siloed, employee-tied pockets, but the EU AI Act assumes formalized enterprise-wide governance. That gap drives the risk data teams now face.
As of Aug. 2, the EU AI Act assesses penalties for the first time against AI oversights stemming from poor data management. The Act places the onus of compliance squarely on data teams and applies to any organization that interacts with EU residents.
A clear problem is that most organizations didn't institute their data infrastructure with such documentation in mind because it wasn't previously required. Now that the Act's obligations for transparency, general purpose AI (GPAI) models and AI literacy are in force, there's a divide between what it penalizes and where most data governance programs actually stand.
By examining which specific data management mechanisms are required and where most data governance programs fall short, data teams can identify concrete actions to rectify these shortfalls.
Data management requisites
Originally, the AI Act mandated high-risk Annex III systems to be fully compliant by August 2, 2026. The May 2026 Digital Omnibus agreement tabled many of those requirements until December 2027. The components that took effect Aug. 2 still demand rigorous data management.
Under Article 50's transparency requirements, data teams must mark AI-manipulated depictions of real people and deepfakes as AI-generated. Chatbots and similar conversational applications must inform people they're interacting with AI, unless it's obvious, and users must be told when biometric categorization or emotion recognition is in use.
GPAI model providers, including teams who substantively tailor foundation models or build applications heavily reliant on them, must keep technical documentation under Annex XI, furnish training data summaries, and confirm the data adheres to EU copyright law under a documented policy. Data literacy is also required of all users and operators of AI systems.
Governance shortcomings
For data teams, the biggest obstacle to EU AI Act compliance is AI governance itself. A 2025 Gartner survey on how organizations handle cybersecurity risk found that nearly 90% lacked AI governance programs. Although that figure represents only one slice of the enterprise, it still indicates how immature AI governance is today. Many organizations run AI governance tied to specific employees or projects. Meeting the Act's transparency and GPAI mandates becomes far easier when the roles, rules and responsibilities for governing AI systems are formalized, documented and spread throughout the enterprise.
Many users and operators of AI systems know their own tooling but not what runs in other departments or serves other customers. Compliance with the Act's GPAI and transparency requirements all but requires data teams to maintain a comprehensive inventory of which AI systems the organization operates, along with their data dependencies and formal documentation of both -- a demand that falls hardest on GPAI model providers.
Noncompliance penalties for GPAI models are €15M or 3% of global turnover, whichever is higher. Systemic risk GPAI models are trained with 10²⁵ FLOPs of compute and require adversarial testing such as red teaming, cybersecurity protections and energy consumption data disclosures. Data teams without complete inventories of AI assets can't meet this requirement.
Siloed, employee-tied governance undermines transparency even more directly when it comes to labeling model outputs as AI-generated. When only individual employees hold that knowledge, it rarely reaches end users, even if they pass it to a manager or colleague. The Act requires that knowledge be readily available to the general public when people interact with these systems, as well as to internal users. Noncompliance penalties are €7.5M or 1.5% of global turnover, whichever is higher.
Remediation
Applying the fundamentals of data governance to their AI instances lets data teams close these gaps. The first step is to classify AI assets fully: the models themselves, their use and users, training data, data models, taxonomies, and data sources. That inventory is what the Act's GPAI documentation and transparency obligations assume already exists. Without it, a data team cannot show which systems are in scope, let alone document them.
Typically, only a few key employees hold this institutional knowledge. However, data teams can formalize it through enterprise architecture to map these systems, and through regex, machine learning, and other statistical methods to discover, classify and tag them.
Next, data teams should categorize the AI assets by the EU AI Act risk tier: minimal risk, limited risk with transparency requirements, high risk, and prohibited. The tier a system falls into determines which obligations attach to it. A limited-risk chatbot triggers Article 50's disclosure rules, while a systemic GPAI model pulls in adversarial testing and energy reporting. These tiers inform next steps, particularly the immediate documentation of GPAI models and public disclosures when people are interacting with AI systems.
The tiers require data teams to update existing governance policies or write new ones. Splitting those policies into two distinct sets helps: one set for the public's interactions with AI systems, another for internal work, such as red teaming or cybersecurity for systemic GPAI models. Legal counsel can tell an organization whether their use of GPAI models classifies it as a model provider, even if it didn't build the model.
In many cases, deployments of dynamic agents can automate parts of the transparency requirements, including watermarking content as AI-generated or triggering workflows to inform end users they're interacting with AI.
Jelani Harper is a data industry analyst and journalist covering data management, AI and enterprise IT for more than a decade. He is a research lead at Blue Badge Insights, writing analyst reports for GigaOm and articles for VentureBeat and The New Stack.