DATA GOVERNANCE

What Is Data Governance?

11 September 2026
Data Governance visual showing enterprise data, policy, security, analytics and cloud systems connected around a governed global data environment.

Enterprise data doesn't sit in one place. It's spread across databases, warehouses, data lakes, SaaS applications, cloud storage, APIs, files, collaboration tools, operational applications, analytics platforms — and increasingly, AI systems. As organisations grow, data becomes fragmented across systems and teams faster than anyone updates a diagram of where it all is.

The governance problem that follows isn't really about storing data. It's about understanding and controlling it across all of that spread.

Before an organisation can govern its data, it needs to know what data exists and where it lives.

What is data governance?

Data governance is the system of ownership, policies, controls and accountability an organisation uses to manage how data is created, accessed, used, shared, retained and deleted. Done well, it answers a specific set of questions for any given piece of data:

  • What data exists?
  • Where is it?
  • Who owns it?
  • Who can access it?
  • How sensitive is it?
  • Where has it come from?
  • Where is it going?
  • Which policies apply?
  • How long should it be retained?
  • Who changed or accessed it?
  • What actions should be allowed?

Data governance shouldn't simply describe how data ought to be managed. Effective governance makes those rules visible, operational and enforceable — not just documented.

Why enterprises need data governance

The practical case for governance tends to come from the same handful of recurring problems:

  • Data spread across many systems, with no single view of it
  • Unclear ownership — nobody quite sure who's responsible for a given dataset
  • Duplicated data, often drifting out of sync with itself
  • Sensitive data turning up in places nobody expected
  • Excessive access, granted for convenience and never revisited
  • Inconsistent retention — some data kept far longer than it should be, some deleted too soon
  • Data movement nobody can fully account for
  • Regulatory obligations that assume a level of control the organisation doesn't actually have
  • Poor data quality undermining decisions built on top of it
  • AI systems accessing enterprise data in ways that weren't explicitly planned for
  • Difficulty answering what should be basic questions about the organisation's own data

The core pillars of data governance

Put together, modern data governance tends to break down into ten capability areas:

  1. 01DISCOVERY
    • Know what data exists.
  2. 02CLASSIFICATION
    • Understand what the data contains and how sensitive it is.
  3. 03OWNERSHIP
    • Know who is responsible for it.
  4. 04ACCESS
    • Understand and control who or what can use it.
  5. 05LINEAGE
    • Understand where data came from and where it moves.
  6. 06POLICY
    • Define how data may be used.
  7. 07RETENTION
    • Control how long data should exist.
  8. 08QUALITY
    • Understand whether data is accurate, complete and usable.
  9. 09AUDIT
    • Know what happened to the data.
  10. 10CONTROL
    • Be able to act when policy requires it.

These capabilities don't need to come from one universal platform — in most organisations they don't. They're often delivered by several systems working together rather than a single tool that does everything.

Data discovery comes before governance

You cannot govern data you do not know exists. That makes discovery the actual starting point, not an optional first step — across:

  • Structured databases
  • Warehouses
  • Cloud storage
  • SaaS
  • APIs
  • Files
  • Applications

Discovery on its own isn't governance. It's what everything else — inventory, classification, ownership, policy — needs in order to have something to act on.

Data classification

Knowing data exists isn't the same as knowing what it actually contains. Classification is what distinguishes, for instance:

  • Personal data
  • Financial information
  • Employee information
  • Customer records
  • Intellectual property
  • Credentials
  • Health information
  • Confidential commercial information
  • Public data

What a dataset is classified as can reasonably influence its access, retention, sharing, security requirements, approval path and deletion rules. (What follows from a given classification varies by organisation, jurisdiction and context — this article isn't offering legal classifications as universal rules.)

Ownership and accountability

Enterprise data should have identifiable responsibility attached to it — a person or team who can actually answer for it, split across a few kinds of responsibility depending on the organisation:

Business ownership
Accountable for why the data exists and whether it's still needed.
Technical stewardship
Accountable for how it's stored, processed and maintained.
Security/privacy responsibility
Accountable for how it's protected and who it's exposed to.

There's no single universal model for how to split this across an organisation. What matters is that if nobody is accountable for a dataset, governance becomes difficult to operationalise— there's no one to make the calls policy depends on.

Data access governance

Knowing data exists is a different problem from knowing who — or what — can actually access it. Access governance covers the full range of things that touch data in a modern enterprise:

  • Users
  • Service accounts
  • Applications
  • APIs
  • AI systems
  • Autonomous agents

The organising principle here is least privilege: access scoped to what a person or system actually needs, not what's convenient to grant. Access governance is a large enough topic to deserve its own treatment — we cover it in more depth in data access governance.

Data lineage and movement

Understanding a dataset well enough to govern it means being able to trace:

  • Where it originated
  • Where it was copied
  • Which systems transformed it
  • Where it moved
  • Which downstream systems depend on it

Lineage matters across governance, security, quality, privacy, incident response and — increasingly — AI: when something goes wrong, the first useful question is usually where the affected data actually came from and everywhere it went. None of this is trivial to do perfectly in real time across a large, changing environment; it's worth being realistic about that rather than treating it as a solved problem.

Policy needs to become operational

This is where governance tends to break down in practice: the gap between a policy document and an operational control. A policy might state:

“Customer information must not be shared with unauthorised external services.”

Turning that into something enforceable means the organisation can actually identify the data, its classification, the actor requesting access, the destination, the applicable policy and the requested action — and, from that, determine what should happen:

EVALUATES
The dataIts classificationThe requesting actorThe destinationApplicable policyRequested action
ALLOW
DENY
RESTRICT
REQUIRE APPROVAL
LOG
REMEDIATE

This is architectural framing for what that decision point needs to be capable of — not a claim that Sentinel, or any product, currently performs every one of these actions in production today.

Data governance should operate closer to real time

Traditional governance processes lean heavily on periodic reviews, spreadsheets, static catalogues, manual approvals and retrospective audits. That worked reasonably well when data environments changed slowly. Increasingly dynamic environments — new datasets, new access, changed permissions, data on the move, new AI connections, policy violations — put pressure on governance systems to respond faster than a quarterly review cycle allows.

None of this means governance can or should be instantaneous everywhere; some of it genuinely can't be, and claiming otherwise wouldn't be honest. The more defensible version of the idea is narrower:

Governance becomes more useful when visibility and control move closer to the moment data is actually being used.

We go deeper on what that looks like in practice in real-time data governance.

Data governance in the age of AI

AI doesn't replace the data-governance problem. It makes it more important. AI systems may retrieve enterprise data, summarise it, transform it, infer new information from it, send it to models or providers, expose it through generated outputs, act on it, or move it between systems. That means the old question — which humans can access our data? — now sits alongside a newer one: which systems, models and agents can access our data?

We cover the agent and AI-governance side of this in more depth in AI agent security, shadow AI and enterprise AI governance. Here, the point is narrower: AI is one more — increasingly significant — category of thing that touches enterprise data. Data is still the subject.

A modern data governance lifecycle

Put together, a workable governance process tends to follow a consistent progression:

DISCOVERFind the data that actually exists across the enterprise.
CLASSIFYUnderstand what it contains and how sensitive it is.
ASSIGN OWNERSHIPRecord who is responsible for it.
MAP ACCESSUnderstand who and what can reach it.
APPLY POLICYDefine how it may be used, shared and retained.
MONITORWatch access and movement on an ongoing basis.
CONTROLAct — restrict, revoke, quarantine or delete — when policy requires it.
AUDITRecord what happened to it and what was decided.
REVIEW / RETIREReassess as usage changes, or decommission cleanly.

What should a governed data record contain?

For each dataset, a useful governance record needs real structure behind it — illustrative, not an industry standard:

FIELDWHAT IT CAPTURES
DatasetWhat it's called and how people refer to it.
DescriptionWhat it actually is, in plain terms.
Business ownerWho is accountable for why it exists.
Technical ownerWho is accountable for how it's stored and maintained.
LocationWhere it actually lives.
Data classificationWhat kind of data it is.
SensitivityHow sensitive it is, and why.
SourceWhere it originated.
Downstream systemsWhat depends on it.
Users with accessWhich people can reach it.
Applications with accessWhich applications can reach it.
AI systems with accessWhich models, assistants or agents can reach it.
Retention policyHow long it should exist.
Applicable policiesWhich policies govern it.
Last accessWhen it was last used.
Last modificationWhen it last changed.
Last reviewWhen someone last actually checked the above.
Audit historyThe record of what's happened to it and what was decided.

A practical data governance checklist

  • Discover enterprise data.
  • Maintain an inventory.
  • Classify sensitive information.
  • Assign ownership.
  • Map access.
  • Apply least privilege.
  • Understand lineage.
  • Track data movement.
  • Define retention requirements.
  • Define deletion processes.
  • Establish data policies.
  • Monitor access changes.
  • Audit important actions.
  • Review third-party data access.
  • Understand AI access to data.
  • Detect policy violations.
  • Provide remediation paths.
  • Review governance continuously.

From visibility to control

Most organisations sit somewhere on a maturity progression, whether they've mapped it explicitly or not:

UNKNOWN
“We don’t know where all of our data is.”
VISIBLE
“We know what data exists and where.”
UNDERSTOOD
“We know what it contains, who owns it and how it moves.”
GOVERNED
“We know which policies and permissions apply.”
CONTROLLED
“We can act when data use violates policy.”

A catalogue can tell an organisation what data exists. A control system should help it determine what's allowed to happen to that data next.

Where Sentinel fits

Data governance and control are central to what Porthos Labs is exploring with Sentinel. Sentinel is being designed to give organisations greater visibility into enterprise data: where it lives, what it contains, who or what can access it, how it moves, and which policies apply.

The longer-term product direction is to move beyond visibility toward governed control — helping authorised teams search, understand and act on enterprise data more easily while maintaining permissions, policy and auditability. AI systems and autonomous agents are part of this problem because they are increasingly powerful consumers of, and actors on, enterprise data.

Closing

Data governance shouldn't be a static record of how an organisation hoped its data was being used. It should give organisations an increasingly accurate understanding of their data, and the ability to control what happens to it.

Know your data. Understand it. Control it.

RELATED RESEARCH
← Back to Research