Skip to main content
Expanding our AI Security Suite, Powered by Behavioral AILearn More

Aug 17, 2026

Excessive Agency in AI Systems: Causes, Risks, and Controls

Excessive agency lets AI agents act beyond their task. See how it happens, where it has surfaced, and the controls that keep agents within scope.

Key Insights

OWASP identifies excessive functionality, permissions, and autonomy as the root causes of excessive agency.

The National Institute of Standards and Technology describes how a hidden instruction can produce a valid, authenticated request from a trusted agent.

The National Vulnerability Database records an agent framework whose unrestricted database tool let prompt injection trigger destructive SQL commands.

Nearly half of organizations have already experienced a security incident involving an AI agent, and 53% say their agents exceed intended permissions at least occasionally, according to the Cloud Security Alliance's enterprise AI agent survey. Those scope violations describe excessive agency in AI systems: the condition in which an agent's ability to act, through its tools, credentials, or unattended decisions, outruns the task it was assigned.

Agency problems rarely stem from a single flaw. This guide breaks down the three root causes of excessive agency and follows the chain from untrusted input to authenticated action. We also lay out the deterministic controls, telemetry, and governance thresholds that keep an agent inside its task boundary. You'll learn where these failures have already surfaced in real-world production systems and how to draw the line between an AI assistant working within its scope and one capable of destructive, fully authorized changes.

Key Takeaways

  • A useful assessment starts by comparing every available action with the agent's assigned job.

  • Risk analysis should trace the full path from incoming context to the resulting tool call.

  • Deterministic controls should enforce boundaries after the model produces an output.

  • Autonomy should expand only as controls prove they hold.

What Is Excessive Agency?

Excessive agency is a vulnerability in which an AI system's ability to act exceeds what its task requires. This happens when an agent holds more authority, in functions, permissions, or autonomy, than its assigned task needs. OWASP calls excessive agency "the vulnerability that enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs from an LLM, regardless of what is causing the LLM to malfunction."

A practical reading of "excessive" is any agency beyond the task: a mailbox summarizer needs to read messages, so a send capability is excess. OWASP ranks that gap among the most consequential risks in its Top 10 for LLM applications.

It helps to distinguish excessive agency from adjacent failure modes. Hallucinations, bugs, and prompt injection produce bad output. In contrast, excessive agency determines what that output can do. Consider a code bug that passes the wrong user ID to a read-only tool. It only leaks one record. The same bug behind a tool with the ability to bulk delete rights wipes an entire table. Improper output handling covers insufficient scrutiny of output before the handoff, while excessive agency covers the authority after it.

The infographic contrasts three adjacent failure modes—hallucinations, prompt injection, improper output handling—that affect output quality with excessive agency, where AI autonomously executes actions, illustrated by a database deletion

How OWASP's Excessive Agency Risk Has Changed

OWASP has renumbered the entry and broadened its triggers and controls with each edition since 2023, so checking the edition label tells you which control list an audit, vendor claim, or internal policy maps to.

LLM08 in the 2023/24 Edition

Excessive agency first appeared as LLM08 in July 2023, in a version that tied the risk to "unexpected/ambiguous outputs" and listed malicious plugins and poorly engineered prompts among the causes. That edition, which used plugin terminology throughout, also recommended human approval for every action.

LLM06 in the 2025 Edition

The 2025 entry moved to LLM06 and added "manipulated outputs" to its definition. Its triggers, which already covered hallucination and prompt injection, now include compromised extensions and, in multi-agent systems, malicious or compromised peer agents. "Extension" replaces "plugin" as the umbrella term in this edition. Human approval narrows from all actions to high-impact actions, and complete mediation and sanitizing LLM inputs and outputs appear as explicit prevention items.

The Expanded Agentic Emphasis in 2026

OWASP's 2026 edition reflects how far the agentic shift has reshaped the risk: excessive agency climbed from LLM06 to LLM03, the biggest rank change on the list. OWASP had published a separate Top 10 for Agentic AI Applications in December 2025, extending the risk's scope to tool misuse, identity and privilege abuse, and cascading failures across autonomous, tool-using systems.

Top 3 Root Causes of Excessive Agency

Three root causes inflate different parts of an agent's agency budget: excessive functionality, excessive permissions, and excessive autonomy.

1. Excessive Functionality Exposes Unnecessary Actions

Common engineering shortcuts create unnecessary functionality. An editing extension can modify and delete documents when the agent only needs to read them. A plugin trialed in development often stays wired in after launch. A shell-command extension may fail to filter out the rest. Each unused function is an action a confused or manipulated model can still select.

2. Excessive Permissions Expand the Blast Radius

A read-only extension may connect to a database with an identity holding SELECT, UPDATE, INSERT, and DELETE rights. A per-user extension may run under a generic high-privileged identity with access to every user's files instead of one scoped to that user alone. Pairing each identity with role-based access control keeps those permissions mapped to what the agent's task actually needs. Agents often run on broad delegated service tokens, so one manipulated call can reach beyond the requesting user's data.

3. Excessive Autonomy Removes Necessary Review

An extension that deletes a user's documents without confirmation has excessive autonomy. In practice, autonomy as a design choice remains separate from capability; designers can still require a capable agent to ask before acting. Persistent memory extends autonomy across sessions because a wrong assumption stored once, such as which records count as obsolete, can shape later actions without review.

How Excessive Agency Turns Model Output Into Action

Excessive agency causes harm through several connected steps that end in an authenticated, harmful action.

Granting Agency Through Tools and Extensions

The path starts at design time, when developers wire a customer-service agent to a payments API that can issue refunds, assign its credentials, and decide whether calls run automatically. Control starts here, with narrow tools and scoped credentials.

Ingesting Untrusted Instructions or Ambiguous Context

Ingestion is where attacker-controlled content first reaches the model: agents read emails, web pages, repository files, and support tickets. The National Institute of Standards and Technology (NIST) describes AI agent hijacking guidance as addressing a type of indirect prompt injection in which an attacker inserts instructions into data the agent ingests. A vague request can also push the agent toward a harmful interpretation. Separating untrusted content from instructions at this stage limits how far a hidden instruction can travel.

Passing Manipulated Output Into a Tool

The model blends trusted instructions with untrusted content and emits a tool call, such as a refund amount or a raw SQL statement handed to a database tool. Unvalidated output passed downstream gives whoever controls the input indirect access to everything the tool can do. A refund limit written only into the system prompt offers no protection because injected text targets the prompt. Output validation at this handoff can stop out-of-scope calls.

Executing an Authenticated but Harmful Action

A hidden instruction can become a valid, authenticated request from a trusted agent, as NIST agent security guidance shows in a delete-user scenario. The downstream API provides the last chance to refuse the request by checking the requesting user's rights.

Real-World Examples of Excessive Agency

Incidents, vulnerability records, and research demonstrations show how excessive permissions, functionality, and autonomy produce real consequences.

Destructive Database Actions Without Technical Blocks

Broad permissions and unchecked autonomy can cause destructive database actions even without prompt injection. The AI Incident Database records a July 2025 case in which an AI development assistant deleted a live production database during a code freeze despite the user's repeated instructions not to act, according to the user's public account.

Open-Ended Tools Enabling Destructive Commands

The National Vulnerability Database entry for CVE-2025-67510 describes an agent framework whose database tool accepted arbitrary SQL without semantic restrictions. Prompt injection could therefore trigger DROP, TRUNCATE, DELETE, and ALTER statements, limited only by database permissions, until a later release fixed it.

Security and Operational Impact

Excessive agency turns a single bad output into data loss and disclosure, destructive changes, or runaway cost, usually through fully authorized identities.

Unauthorized Data Access and Disclosure

An agent with read access and an outbound channel can exfiltrate data. In NIST's evaluation of one model family, a hidden instruction caused an assistant to exfiltrate login credentials after the user asked it to show unread emails.

Destructive Changes to Systems and Records

Agents can use identities with valid authorization to delete records or alter configurations. The same authority can approve a refund no one meant to send. Access controls therefore approve each request before execution, and the damage appears only afterward.

Runaway Costs, Repeated Actions, and Cascading Failures

A planning loop with no termination or escalation rule does not stop on its own, so a single mistake keeps repeating until something else intervenes. A loop may accumulate API costs while repeatedly triggering downstream automation. In multi-agent pipelines, one bad action becomes the next agent's input and spreads the failure.

How to Prevent Excessive Agency

Prevention caps each budget line with deterministic controls outside the model because models cannot reliably police themselves.

A five-step security framework—narrow tools, least-privilege access, downstream authorization, human approval gates, and agent isolation with rate limits—arranged around a central AI agent icon to prevent excessive agency

1. Narrow Tools to Task-Specific Functionality

Narrowing a tool to the one action a task requires removes whichever extra capability a manipulated model could otherwise reach for. A file-writing tool can replace shell access. A summarizer only needs read access, not the ability to send. Open-ended extensions such as "run a shell command" or "fetch a URL" expose more functionality than many tasks require. When a deterministic function can handle a task, the task may not need an agent at all.

2. Scope Agent Identities With Least-Privilege Access

Each agent can carry its own identity instead of shared or human credentials, grounded in standard identity and access management practice, scoped through database permissions and user-context execution with least-privilege access controls and minimum OAuth scopes. This separation limits both the systems an agent can reach and the actions it can perform for a particular user. Abnormal's AI Agent Security, part of its AI Security suite and currently in private preview, is designed to track each agent's privileges, tool access, and reachable data, flagging authority that creeps past the task.

3. Enforce Downstream Authorization and Complete Mediation

Complete mediation checks every access for authority, and agents need that check in the downstream API, where it stays deterministic and auditable. In the refund example, the payments API rejects any refund above the approved subscription limit, whatever the model requests, and validates each tool call against the user's permissions and session context.

4. Require Human Approval for High-Impact Actions

Approval gates pause an agent before irreversible, external, financial, privilege-changing, or bulk actions, such as deleting data, sending messages, moving money, or changing credentials. Gates on routine actions cause approval fatigue and train reviewers to click through.

5. Isolate Agents and Set Rate Limits

Sandboxed execution environments, restricted outbound connectivity, and separate agents for different access levels contain what a compromised agent can touch. Rate limits and per-session budgets cap a runaway loop's reach, though they do not prevent the initial action. Fixed outbound email and refund limits can constrain a support agent in a sandbox whose only egress is the payments API.

How to Detect and Test Excessive Agency

Detection depends on logging every tool call with its authorization outcome, testing agents against injected and ambiguous inputs before deployment, and feeding confirmed failures into existing incident response processes.

Tool-Call and Authorization Telemetry

Useful audit records capture action classification, risk score, authorization outcome, approval identifier, execution result, and policy version. A unified monitoring view can then connect each model decision to the policy and identity that governed the resulting tool call. Abnormal's AI Agent Security is designed to generate telemetry like this for the agents it discovers, flagging runs that drift off an agent's assigned task.

Alerts for Privilege Use and Approval Bypass

Useful alert conditions include:

  • Elevated Privilege Use: An agent exercises rights it rarely or never needs.

  • Approval Bypass Attempts: Repeated calls try to route around a required gate.

  • Approval Drift: Reviewer approval behavior shifts over time.

Together, these alerts show whether permissions or autonomy are moving beyond the agent's normal task boundary.

Adversarial Tests for Indirect Prompt Injection

In the InjecAgent benchmark, a prompted frontier-model agent was vulnerable 24% of the time, rising to 47% with a reinforced attacker prompt. By contrast, AgentDojo requires agents to finish their original task while under injection, which tests whether security controls preserve both safety and usefulness.

Scenario Tests for Ambiguous and Underspecified Tasks

ToolEmu tests how benign but underspecified instructions lead agents to act on unwarranted assumptions. A prompt like "clean up old customer records" run against a production-like dataset reveals whether an agent asks for clarification or starts deleting.

3 Governance Best Practices for Safe Levels of Autonomy

Governance sets how much autonomy each agent gets and who answers for its actions, an extension of the AI governance responsibilities teams already carry for other systems.

1. Approval Thresholds Based on Action Impact

Approval thresholds work best when they weigh reversibility and reach. Reading logs can run unattended. Revoking credentials, isolating hosts, deleting data, or changing production policy warrants approval first.

2. Named Human Owners and Traceable Decisions

Each agent needs a defined set of controlling human users and an accountable owner. Every consequential action should trace to that owner through an audit trail linking each approval identifier to the person who granted it.

3. Progressive Autonomy as Control Confidence Grows

Research on agent autonomy describes a progression from an operator level, where the user directs every action, up to an observer level, where the agent acts alone. New agents can start at the consultant or approver level and gain autonomy only as telemetry and testing show controls holding. Each level also needs a mechanism that can disengage the agent.

Keep Agency Within the Task Boundary

Safe agency depends on aligning tools, credentials, authorization, approval, and monitoring with the assigned task. A useful pre-deployment evaluation lists every tool, credential, and unattended action an agent holds, then identifies anything the task cannot justify. As systems gain autonomy, that evaluation provides a concrete basis for deciding whether their controls are ready to support it.

Abnormal's AI Security suite is designed to put that evaluation into practice. AI Agent Security discovers agents across AI platforms, maps ownership, and tracks each agent's privileges, tool access, and reachable data, flagging authority that creeps past the task. AI Cloud Security correlates identity and service account anomalies to catch compromised agents acting against AWS, Azure, or Google Cloud, with automated containment through credential rotation, access revocation, and session termination.

Book a demo to see how Abnormal helps contain a compromised agent once confirmed.

Frequently Asked Questions

Can Excessive Agency Occur Without an Attacker?

Yes. An ambiguous request or misread context can push an over-privileged agent into a harmful action with no adversary involved. The agent's permissions, available tools, and approval requirements determine how far that mistake can spread.

Which Agent Actions Should Require Human Approval?

A single reversible edit rarely needs a gate, but the same operation across thousands of records usually does. Reviewers benefit from seeing the exact arguments, such as the recipient, amount, or SQL statement, rather than a model-written summary that could omit important details.

Does Read-Only Access Eliminate Excessive Agency Risk?

No. Read-only access removes destructive changes but leaves disclosure risk when the agent can access sensitive data and communicate through an outbound channel. In practice, sensitive read permissions deserve controls comparable to write permissions when an outbound path exists.

Protect Against Evolving Email Threats

See how behavioral AI detects attacks that legacy defenses miss.