Getty Images/iStockphoto
Why data can't be standardized at the source
Deterministic data controls benefit from early standardization, but context-dependent governance only works at the point of use. The distinction reshapes where each control belongs.
The data industry has spent the last few years borrowing a defining principle from software engineering: shift left. The goal is to catch problems earlier, when they are cheaper and easier to fix, rather than cleaning them up downstream.
Applied to data, that means embedding governance closer to the point of creation. And for deterministic standards, it works. Ownership, classification, schema validation, data contracts, access policies and quality expectations all benefit from being established early. The more organizations automate those controls and embed them into existing workflows, the less governance becomes a separate task.
The problem starts when shift left becomes the all-encompassing governance model. Some rules can be determined at the source. Others depend on how data is eventually combined, interpreted and used.
Shift left reaches its limits when governance depends on context.
Gartner predicts that by 2027, 60% of organizations will fail to realize the anticipated value of their AI use cases due to incohesive data governance frameworks. The problem isn't simply a lack of upfront effort. Traditional governance practices can be too rigid and disconnected from how data is applied in business contexts.
Upstream application developers cannot predict how data will be interpreted months later, how disparate data sets will combine, or how autonomous AI models will act on nuanced business logic in real time.
The question is not whether to shift left or right, but how to make governance follow the data from creation through consumption, placing each control at the specific point where it can be most effective.
To do that, leaders need to account for three operational realities:
1. Context is created where data is used, not where it starts
Systems capture data for a particular purpose, not every meaning that data may take on later. A developer building an application can ensure a customer record has a valid ID, clean formatting and masked payment details.
What the developer cannot predict is the business context required six months later when an analytics team or an autonomous AI agent calculates customer lifetime value across three distinct business units. Business definitions, revenue metrics and operational rules evolve as data moves through the business. You can't govern context that doesn't yet exist, and asking software teams to anticipate every future use up upfront doesn't solve that problem.
2. Engineering velocity stalls under static compliance
Shifting governance left works when controls can be automated and embedded into the workflows teams already use. But when organizations push every regulatory policy, access condition and semantic rule upstream, they risk shifting the governance burden as well.
Data producers are left to interpret policies and make decisions for which they may lack context. The distinction matters: automation embeds governance in the workflow, while manual requirements simply shift governance work from one team to another. Shift left should make governance less manual, not simply move manual work elsewhere.
3. Intelligence operates at runtime, not ingestion
Traditional applications follow fixed, predictable logic. Modern data workflows and autonomous AI models do not. They synthesize unstructured documents, dynamically combine multiple data sources, and make probabilistic inferences in real time. A control applied at the source cannot account for every decision an AI system will make with that data later. An agent can still misinterpret an internal policy, apply the wrong business definition or draw an incorrect conclusion during execution. When governance stops at data entry, companies pay the hallucination tax, and human teams have to verify outputs, correct mistakes and intervene before bad decisions become actions.
Matching the control to the moment
The goal is not to choose between upstream and downstream governance. It is to put each control where it can actually succeed.
At the source, organizations should move the controls that benefit from early standardization, including data ownership, initial classification, structural contracts, schema validation, quality baselines and core privacy requirements. Automate as much of that work as possible so governance becomes part of how data is created, not another process teams have to navigate.
At the point of consumption, controls can account for context that only exists once data is combined, interpreted and put to use. The question changes from whether the data meets a predefined standard to whether it is appropriate for this user, this purpose and this action. This is where organizations need to manage an evolving business context, monitor whether data is fit for purpose and enforce automated guardrails governing how AI systems use and act on data.
Code can be validated before it ships. Data is different. Its meaning, risk and appropriate use evolve as it moves through the business. Governance has to move with it.
Felix Van de Maele is co-founder and CEO of Collibra and was named EY Technology Entrepreneur of the Year in 2019.