What Is Data Governance?

Enterprise data doesn't sit in one place. It's spread across databases, warehouses, data lakes, SaaS applications, cloud storage, APIs, files, collaboration tools, operational applications, analytics platforms — and increasingly, AI systems. As organisations grow, data becomes fragmented across systems and teams faster than anyone updates a diagram of where it all is.
The governance problem that follows isn't really about storing data. It's about understanding and controlling it across all of that spread.
Before an organisation can govern its data, it needs to know what data exists and where it lives.
What is data governance?
Data governance is the system of ownership, policies, controls and accountability an organisation uses to manage how data is created, accessed, used, shared, retained and deleted. Done well, it answers a specific set of questions for any given piece of data:
- What data exists?
- Where is it?
- Who owns it?
- Who can access it?
- How sensitive is it?
- Where has it come from?
- Where is it going?
- Which policies apply?
- How long should it be retained?
- Who changed or accessed it?
- What actions should be allowed?
Data governance shouldn't simply describe how data ought to be managed. Effective governance makes those rules visible, operational and enforceable — not just documented.
Why enterprises need data governance
The practical case for governance tends to come from the same handful of recurring problems:
- Data spread across many systems, with no single view of it
- Unclear ownership — nobody quite sure who's responsible for a given dataset
- Duplicated data, often drifting out of sync with itself
- Sensitive data turning up in places nobody expected
- Excessive access, granted for convenience and never revisited
- Inconsistent retention — some data kept far longer than it should be, some deleted too soon
- Data movement nobody can fully account for
- Regulatory obligations that assume a level of control the organisation doesn't actually have
- Poor data quality undermining decisions built on top of it
- AI systems accessing enterprise data in ways that weren't explicitly planned for
- Difficulty answering what should be basic questions about the organisation's own data
The core pillars of data governance
Put together, modern data governance tends to break down into ten capability areas:
- 01DISCOVERY
- Know what data exists.
- 02CLASSIFICATION
- Understand what the data contains and how sensitive it is.
- 03OWNERSHIP
- Know who is responsible for it.
- 04ACCESS
- Understand and control who or what can use it.
- 05LINEAGE
- Understand where data came from and where it moves.
- 06POLICY
- Define how data may be used.
- 07RETENTION
- Control how long data should exist.
- 08QUALITY
- Understand whether data is accurate, complete and usable.
- 09AUDIT
- Know what happened to the data.
- 10CONTROL
- Be able to act when policy requires it.
These capabilities don't need to come from one universal platform — in most organisations they don't. They're often delivered by several systems working together rather than a single tool that does everything.
Data discovery comes before governance
You cannot govern data you do not know exists. That makes discovery the actual starting point, not an optional first step — across:
- Structured databases
- Warehouses
- Cloud storage
- SaaS
- APIs
- Files
- Applications
Discovery on its own isn't governance. It's what everything else — inventory, classification, ownership, policy — needs in order to have something to act on.
Data classification
Knowing data exists isn't the same as knowing what it actually contains. Classification is what distinguishes, for instance:
- Personal data
- Financial information
- Employee information
- Customer records
- Intellectual property
- Credentials
- Health information
- Confidential commercial information
- Public data
What a dataset is classified as can reasonably influence its access, retention, sharing, security requirements, approval path and deletion rules. (What follows from a given classification varies by organisation, jurisdiction and context — this article isn't offering legal classifications as universal rules.)
Ownership and accountability
Enterprise data should have identifiable responsibility attached to it — a person or team who can actually answer for it, split across a few kinds of responsibility depending on the organisation:
- Business ownership
- Accountable for why the data exists and whether it's still needed.
- Technical stewardship
- Accountable for how it's stored, processed and maintained.
- Security/privacy responsibility
- Accountable for how it's protected and who it's exposed to.
There's no single universal model for how to split this across an organisation. What matters is that if nobody is accountable for a dataset, governance becomes difficult to operationalise— there's no one to make the calls policy depends on.
Data access governance
Knowing data exists is a different problem from knowing who — or what — can actually access it. Access governance covers the full range of things that touch data in a modern enterprise:
- Users
- Service accounts
- Applications
- APIs
- AI systems
- Autonomous agents
The organising principle here is least privilege: access scoped to what a person or system actually needs, not what's convenient to grant. Access governance is a large enough topic to deserve its own treatment — we cover it in more depth in data access governance.
Data lineage and movement
Understanding a dataset well enough to govern it means being able to trace:
- Where it originated
- Where it was copied
- Which systems transformed it
- Where it moved
- Which downstream systems depend on it
Lineage matters across governance, security, quality, privacy, incident response and — increasingly — AI: when something goes wrong, the first useful question is usually where the affected data actually came from and everywhere it went. None of this is trivial to do perfectly in real time across a large, changing environment; it's worth being realistic about that rather than treating it as a solved problem.
Policy needs to become operational
This is where governance tends to break down in practice: the gap between a policy document and an operational control. A policy might state:
“Customer information must not be shared with unauthorised external services.”
Turning that into something enforceable means the organisation can actually identify the data, its classification, the actor requesting access, the destination, the applicable policy and the requested action — and, from that, determine what should happen:
This is architectural framing for what that decision point needs to be capable of — not a claim that Sentinel, or any product, currently performs every one of these actions in production today.
Data governance should operate closer to real time
Traditional governance processes lean heavily on periodic reviews, spreadsheets, static catalogues, manual approvals and retrospective audits. That worked reasonably well when data environments changed slowly. Increasingly dynamic environments — new datasets, new access, changed permissions, data on the move, new AI connections, policy violations — put pressure on governance systems to respond faster than a quarterly review cycle allows.
None of this means governance can or should be instantaneous everywhere; some of it genuinely can't be, and claiming otherwise wouldn't be honest. The more defensible version of the idea is narrower:
Governance becomes more useful when visibility and control move closer to the moment data is actually being used.
We go deeper on what that looks like in practice in real-time data governance.
Data governance in the age of AI
AI doesn't replace the data-governance problem. It makes it more important. AI systems may retrieve enterprise data, summarise it, transform it, infer new information from it, send it to models or providers, expose it through generated outputs, act on it, or move it between systems. That means the old question — which humans can access our data? — now sits alongside a newer one: which systems, models and agents can access our data?
We cover the agent and AI-governance side of this in more depth in AI agent security, shadow AI and enterprise AI governance. Here, the point is narrower: AI is one more — increasingly significant — category of thing that touches enterprise data. Data is still the subject.
A modern data governance lifecycle
Put together, a workable governance process tends to follow a consistent progression:
What should a governed data record contain?
For each dataset, a useful governance record needs real structure behind it — illustrative, not an industry standard:
| FIELD | WHAT IT CAPTURES |
|---|---|
| Dataset | What it's called and how people refer to it. |
| Description | What it actually is, in plain terms. |
| Business owner | Who is accountable for why it exists. |
| Technical owner | Who is accountable for how it's stored and maintained. |
| Location | Where it actually lives. |
| Data classification | What kind of data it is. |
| Sensitivity | How sensitive it is, and why. |
| Source | Where it originated. |
| Downstream systems | What depends on it. |
| Users with access | Which people can reach it. |
| Applications with access | Which applications can reach it. |
| AI systems with access | Which models, assistants or agents can reach it. |
| Retention policy | How long it should exist. |
| Applicable policies | Which policies govern it. |
| Last access | When it was last used. |
| Last modification | When it last changed. |
| Last review | When someone last actually checked the above. |
| Audit history | The record of what's happened to it and what was decided. |
A practical data governance checklist
- Discover enterprise data.
- Maintain an inventory.
- Classify sensitive information.
- Assign ownership.
- Map access.
- Apply least privilege.
- Understand lineage.
- Track data movement.
- Define retention requirements.
- Define deletion processes.
- Establish data policies.
- Monitor access changes.
- Audit important actions.
- Review third-party data access.
- Understand AI access to data.
- Detect policy violations.
- Provide remediation paths.
- Review governance continuously.
From visibility to control
Most organisations sit somewhere on a maturity progression, whether they've mapped it explicitly or not:
A catalogue can tell an organisation what data exists. A control system should help it determine what's allowed to happen to that data next.
Where Sentinel fits
Data governance and control are central to what Porthos Labs is exploring with Sentinel. Sentinel is being designed to give organisations greater visibility into enterprise data: where it lives, what it contains, who or what can access it, how it moves, and which policies apply.
The longer-term product direction is to move beyond visibility toward governed control — helping authorised teams search, understand and act on enterprise data more easily while maintaining permissions, policy and auditability. AI systems and autonomous agents are part of this problem because they are increasingly powerful consumers of, and actors on, enterprise data.
Closing
Data governance shouldn't be a static record of how an organisation hoped its data was being used. It should give organisations an increasingly accurate understanding of their data, and the ability to control what happens to it.
Know your data. Understand it. Control it.