What Is AI Agent Security?

Most conversations about AI safety still focus on the model: what it was trained on, what it will say, what it will refuse. That question matters, but it is no longer the whole question. AI agents are different from the AI systems that came before them, because they don't just produce answers — they can act. An agent can call tools, query databases, call APIs, write to production systems, trigger workflows, and talk to other agents, with varying degrees of autonomy and varying degrees of human oversight in the loop.
That changes the security question.
It is no longer enough to ask whether an AI model is safe.
Organisations also need to ask:
- What can this agent access?
- What can it do?
- Under whose authority?
- Under what conditions?
- And what happens when it behaves unexpectedly?
What is an AI agent?
An AI agent is a system that uses a model's reasoning to pursue a goal by taking actions, rather than simply returning a response. That distinction matters more than it might first appear.
A language model predicts text. Given an input, it produces an output. On its own, it has no memory beyond the current context, no tools, and no way to affect anything outside the conversation.
An AI assistant wraps a model in an interface — a chat window, a copilot panel — and usually adds some memory and some light integration with other systems, but a person is still driving each step: they ask, it answers, they decide what happens next.
An autonomous or semi-autonomous AI agent goes further. Given a goal, it can plan a sequence of steps, decide which tools to use, execute them, observe the result, and decide what to do next — with a human reviewing the outcome rather than approving every step. In practice, an agent is usually a combination of a few concrete parts: a model, a set of instructions, some memory or context, a set of tools it can call, the permissions attached to those tools, the external systems it can reach, and the decision logic that ties all of that together. Security has to account for every one of those parts, not just the model at the center.
What is AI agent security?
AI agent security is the discipline of controlling what autonomous or semi-autonomous AI systems can access, decide and do across an organisation.
It spans a wider set of concerns than model safety alone, including:
- Identity — knowing which agent is acting
- Authentication — verifying that identity before it acts
- Permissions — what it's allowed to access and do
- Data access — which datasets and records it can reach
- Tool access — which APIs, systems and integrations it can call
- Action controls — what it's allowed to actually execute
- Policy enforcement — evaluating actions against rules before they happen
- Monitoring — observing what agents are doing in real time
- Auditability — a durable record of what happened and why
- Anomaly detection — recognising behaviour outside the norm
- Human approval — a checkpoint for high-impact actions
- Containment — the ability to stop an agent, quickly, when something goes wrong
Why agents create a different security problem
None of this is alarmist — it's a fairly direct consequence of moving from software that responds to software that acts. A few properties of agents make that shift matter:
- Agents can take actions. A wrong output from a chatbot is a bad sentence. A wrong action from an agent can be a deleted record, a sent email, or a changed setting.
- Agents may operate across multiple systems. A single agent can touch a CRM, a data warehouse and an internal API in the course of one task, which multiplies the number of places something can go wrong.
- Agents can inherit powerful credentials. An agent built to “help with finance” is often given the same access a finance employee has — sometimes more, if no one has scoped it down.
- Agents can make decisions from probabilistic model outputs. The reasoning that decides what to do next is not deterministic code; it can be wrong, and it can be manipulated.
- Agents can chain actions together. A small misjudgement at step one can compound by step five, especially when each step feeds the next without a human checkpoint.
- Agents may communicate with other agents. Multi-agent systems introduce trust relationships between agents, not just between a person and a system.
- Agents may operate faster than humans can manually supervise. The volume and speed of agent activity can outpace what a person can review in real time, which is exactly why the controls need to be structural, not just procedural.
The core AI agent attack surface
Put together, these properties create a fairly specific set of risk surfaces worth naming individually:
- Prompt injection
- Malicious or misleading content — in a document, email or web page an agent reads — that manipulates its next action.
- Excessive permissions
- An agent granted broader access than its actual task requires, often inherited wholesale from a human role.
- Sensitive data exposure
- An agent surfacing regulated or confidential data outside its intended context, including to other agents or external tools.
- Tool misuse
- A tool used in a way its designer didn't intend, or invoked with parameters that exceed the task at hand.
- Credential leakage
- API keys, tokens or service credentials an agent holds being exposed through logs, outputs or a compromised tool.
- Unsafe API calls
- Actions sent to external or internal APIs without adequate validation of what's being requested.
- Agent impersonation
- A request or action falsely attributed to a legitimate agent's identity.
- Unapproved agent creation
- New agents or automations stood up outside any inventory or review process.
- Cross-agent trust
- One agent acting on another agent's output without verifying it, propagating an error or manipulation downstream.
- Memory/context poisoning
- Corrupted or manipulated content persisting in an agent's memory and influencing later, unrelated decisions.
- Supply-chain risk from external tools/models
- Third-party models, plugins and tools an agent depends on, each carrying its own trust and update risk.
- Lack of visibility
- The precondition behind most of the above: an organisation that doesn't know which agents exist can't secure them.
Identity becomes fundamental
Once software can act on its own, the first question in almost any incident becomes: which agent did this? Enterprises increasingly need to know, for every agent operating in their environment, which agent is acting, who created it, which user or service authorised it, which credentials it currently holds, which systems it can reach, and what its permissions actually are right now — not what they were assumed to be when it was set up.
Traditional identity and access management gives a useful vocabulary for this — principals, roles, scopes, credentials — but IAM built for human and service accounts doesn't automatically extend to agents that reason about what to do next, spawn sub-tasks, or act on behalf of other agents. What's needed is closer to a distinct concept of agent identity: a durable, inspectable record of what an agent is, what it's allowed to do, and on whose authority — separate from, but consistent with, how the organisation already manages human and service identity.
Least privilege for AI agents
The least-privilege principle isn't new, but it's easy to skip when standing up an agent quickly: give it only the access it actually needs to complete its task, nothing more.
Take a simple, concrete example: an invoice-processing agent. It reasonably needs to:
- Read invoices
- Extract fields
- Compare purchase orders
But it should not automatically receive permission to:
- Change supplier bank details
- Approve large payments
- Export the customer database
Each of those additional permissions might be individually justifiable for some other agent, some other task. The point of least privilege is that they shouldn't come bundled in by default just because the agent happens to operate in “finance.”
Policy enforcement must happen before action
There's an important difference between monitoring what happened and controlling whether an action is allowed to happen. A log entry after the fact tells you what an agent did. It doesn't stop it from doing it. High-risk actions need a real-time policy decision made before execution, not a report generated afterward.
Consider an agent that attempts to export a large volume of customer data to an external service. Before that export happens, a security and control layer needs to be able to evaluate:
- The agent's identity
- The requested action
- The classification of the data involved
- The destination
- The agent's current permissions
- The applicable policy
- The broader risk context
— and, from that evaluation, decide whether to allow, deny, require approval, restrict, or simply log the action. This is an architectural description of what that decision point needs to be capable of — not a claim that any particular product, including Sentinel, currently performs every one of these checks in production today. It's also, more broadly, the decision point an AI control plane is meant to provide.
Human approval still matters
None of this means every action needs a human in the loop — that would defeat much of the value of automation. It means autonomy should be proportional to risk. Low-risk, repetitive tasks can reasonably run with more autonomy. High-impact, hard-to-reverse actions warrant a human checkpoint, including things like:
- Deleting production data
- Changing payment details
- Granting permissions to another agent or user
- Exporting sensitive datasets
- Deploying code
- Modifying infrastructure
- Communicating externally on behalf of the company
Human-in-the-loop controls work best as a targeted checkpoint on the actions that genuinely warrant one, not a blanket requirement that turns an agent back into a suggestion box.
Visibility and auditability
You can't govern what you can't see. That starts with a maintained inventory: which agents exist, which models back them, which tools they can call, what permissions they hold, which data sources they connect to, what actions they've taken, and which policies apply to them.
On top of that inventory, an audit trail needs to be able to answer, for any action:
- Who or what acted?
- What did it access?
- What action was attempted?
- What policy applied?
- Was it allowed?
- Was human approval involved?
- What happened afterward?
That record is what turns “we think our agents behaved correctly” into something an organisation can actually demonstrate — to itself, to auditors, or to a regulator.
Shadow AI and unmanaged agents
Most of the controls above assume the organisation knows an agent exists. That assumption doesn't always hold. Teams stand up agents and automations directly against convenient APIs, outside any central inventory, for entirely reasonable reasons — speed, convenience, a tool that just works. The result is often called shadow AI: AI systems and agents operating inside an organisation that security and governance teams don't know about, and therefore can't secure, audit, or apply policy to. An organisation cannot secure AI systems it doesn't know exist — which makes discovery a precondition for everything else in this article, not an optional extra. We cover this in more depth in What Is Shadow AI?.
What a secure enterprise agent architecture looks like
Pulling the ideas above together, a secure agent architecture generally sits as a control layer between the AI systems making decisions and the enterprise data and actions they touch:
None of these layers is a nice-to-have bolted on afterward — identity, permissions and policy need to be evaluated before an action executes; observability, enforcement and audit need to hold regardless of which model or agent framework sits above them.
A practical AI agent security checklist
For teams starting this work, a reasonable practical checklist looks like:
- Maintain an inventory of agents.
- Give every agent a defined identity.
- Map tools and data access.
- Apply least privilege.
- Separate development and production permissions.
- Control credentials.
- Evaluate risky actions before execution.
- Require approval where impact is high.
- Monitor agent behaviour.
- Record auditable actions.
- Review permissions continuously.
- Detect unmanaged/shadow AI.
- Prepare a containment/kill mechanism.
- Treat external tools and models as supply-chain dependencies.
AI agent security and governance are related, but different
The two terms get used interchangeably, but it's worth separating them. Governance is about rules, accountability, policy and oversight — the decisions an organisation makes about what AI should and shouldn't be allowed to do, and who is responsible for that. Security is about access, behaviour, execution, protection and containment — the mechanisms that actually enforce those decisions in a running system. They overlap heavily, and a mature program needs both, but governance without enforcement is a policy document, and security without governance is enforcement with no rules to enforce. We cover the governance side in more depth in What Is Enterprise AI Governance?.
Where Sentinel fits
This is the problem space Porthos Labs is exploring with Sentinel enterprise AI security and control. Sentinel is being designed as an enterprise AI security and control layer for discovering AI systems, understanding their access and permissions, applying policy, monitoring activity, and creating greater control over autonomous systems.
Closing
As AI moves from generating information to taking action, the security model has to change with it. The central enterprise question will increasingly become not simply “Which AI models are we using?” but “Which intelligent systems are acting inside our organisation, what authority do they have, and how do we remain in control?”