AI Agent OAuth Scopes and Permission Boundaries for Tool-Calling
Delegated tokens alone won't stop agents from accessing data they shouldn't at execution time.

An agent asked to summarize one engineering project holds a service credential that can read many. Both project reads are credential-valid under the terms of the token, but only one of them belongs to the job it was actually given. The credential defines what is reachable, while what belongs to this assignment is a separate, narrower question. A scope defines a perimeter around a category of resource, not a boundary around a specific task, so a token scoped to "read engineering projects" authorizes every project inside that category with equal force, whether the agent needs one of them or all of them.
This gap is structural. Research framing the problem as pre-action authorization draws a hard line between what an agent can do, which is a credential boundary, and what it should do in a given moment, which is a policy boundary that depends on the specific task, the specific arguments, and the specific state of the world at the time of the call. OAuth scopes only ever answer the first question. Enforcement today happens at one of two points: inside the model, through alignment training, or inside the application, through validation logic written ad hoc by whoever built the integration. Neither produces a policy-based enforcement layer, and neither leaves behind a verifiable audit record that shows why a given call was allowed.
Model alignment is a probabilistic defense. A security boundary that fails even a small fraction of the time against a motivated adversary is not a boundary. This holds regardless of how tightly the token itself is scoped: a well-scoped token does not stop an agent from sweeping every record it is permitted to reach, loading all of that returned text into the model's context window, or acting on instructions that arrive hidden inside that data and steer the next call it makes. Scope breadth and the authorization gap are separate variables, and narrowing one does not narrow the other.
How credential inheritance creates a wrong-sized permission envelope
The most common way this gap gets exploited in production systems is not an exotic one: agents receive, or are allowed to forward, the user's own session token instead of a credential purpose-built for the task at hand. An agent holding that same token has no equivalent restraint built in. It can sweep every record the token can reach in a single pass and pour the returned text directly into its own context, with no human pause between retrieval and use.
That inheritance also creates a direct leakage path. The "Agents of Chaos" red-teaming study, cited in the OAP research, found that agents with persistent memory and shell access leaked sensitive identifiers after a single verb in a prompt was reframed. A narrower, purpose-built token issued for a specific task would have prevented the leak structurally, independent of whether the model's alignment held up under the rephrasing.
Credential inheritance also leaves open a second surface that most teams overlook even after they scope tokens properly at setup. Scoping the initial token correctly does not close every door an agent can open over the course of a task.
What delegated authorization flows deliver and where they stop
Delegated authorization flows represent a genuine improvement over inherited tokens, and the industry's movement toward them deserves to be taken seriously on its own terms. The canonical pattern has a user complete an OAuth flow once during setup; that short-lived user token is then exchanged, using OAuth 2.0 Token Exchange under RFC 8693, for a scoped agent token the agent can use later without needing the user present for every action. The agent's token carries the user's identity as a separate claim rather than standing in for the user's own session, which closes the direct leakage path described above: there is no raw user credential for a prompt injection to extract, because the agent never held one.
WorkOS's agent permissions checklist states the rule that makes delegation sound rather than merely convenient: a delegated agent's effective permissions are the intersection of what the agent is permitted to do and what the user can currently do, never the union of the two. Autonomous agents, such as a nightly cleanup job or a scheduled triage process, need a different revocation model than delegated agents acting inside a session for a signed-in user: the former gets revoked by an administrator or during organizational offboarding, the latter loses access the moment the user it acts for does.
These flows rest on standards with real institutional weight behind them. OAuth 2.0 itself (RFC 6749), Token Exchange (RFC 8693), Rich Authorization Requests (RFC 9396, published May 2023), and Resource Indicators (RFC 8707) are all published IETF RFCs, not vendor specifications. OAuth 2.1, which consolidates many of these patterns, remains an Internet-Draft as of September 2026, currently at revision 16, with an IESG submission milestone set for December 2026, which shows the standards body still treats this as an active, unfinished area of work.
None of this, however, reaches into the moment a tool call actually executes. Token exchange, intersection logic, and permission ceilings all operate at setup time or at token-issuance time. A well-scoped, properly delegated agent token can still be used against a resource the task never called for, in a sequence the designer never anticipated, or steered by an instruction smuggled into the data the agent just retrieved. Capability-based tokens push the setup-time model as far as it can go, issuing a token for one narrow purpose, such as sending email to a single specified domain, that expires shortly after issuance. That improves both the blast radius of a compromise and the quality of the audit trail. It is still a credential-level control. It makes no decision about the call itself, at the moment the call is made.
Why enforcement must happen at the tool call boundary
The gap that delegation and capability tokens leave open can only be closed by moving the authorization decision to the moment of the call itself, evaluated synchronously, before the call is allowed to execute. A credential issued at setup time cannot account for the specific arguments of a given call, the state of external systems at that moment, the cumulative effect of several calls taken together, or an instruction injected into the agent's context between issuance and execution. All of those only exist at the instant the call is about to happen.
The OAP research formalizes this as pre-action authorization: the system intercepts the tool call synchronously, evaluates it against a declarative policy, and produces a cryptographically signed audit record, with a measured median decision latency of 53ms. Enforcement has to sit entirely outside agent reasoning, as a separate system the agent cannot talk its way around because it cannot address it directly.
The clearest analogy is the Zero Trust model already deployed for human access to enterprise systems. The operational rule that falls out of this is simple to state: authorization gets checked on every tool call, not once per session. A session-level check, no matter how well it was configured at the start, cannot see a permission that changed mid-session, a context that shifted after the first few calls, or a prompt injection that arrived in call four.
Three categories of decision exist only at call time and nowhere else. Semantic gating catches a call that falls entirely within an agent's role and touches no protected data, yet still needs a human to approve it because it runs after business hours or crosses a cumulative limit the agent has been approaching across several prior calls. Exposure control catches the case where a record the agent is fully permitted to retrieve still contains a secret or a hidden instruction, requiring the boundary to filter fields or redact content before that text ever enters the model's context. Query narrowing lets the boundary reshape the arguments of a call before it reaches the backend, so that an operation the agent is authorized to run in general returns only the records that belong to its current assignment, not everything the underlying credential happens to reach.
The OAP adversarial evaluation gives this a concrete test. Under a restrictive pre-action policy, a comparable population of attackers achieved a 0% success rate across a large number of attempts. That result is the clearest available evidence that pre-action enforcement is not another probabilistic safeguard layered on top of an already-probabilistic model. It behaves as a deterministic gate, which is the property a security boundary actually needs.
What a pre-action enforcement layer checks on every tool call
A pre-action enforcement layer answers five questions on every single tool call, before that call is permitted to run: who is the agent, who is it acting for, what is it allowed to do in general, whether this specific action is allowed right now, and whether the response coming back can be returned safely. Each question maps to a distinct check. Audit and revocation ask whether the system can later prove what happened and, if necessary, undo the access that was granted.
OBPE describes this as a staged pipeline. The boundary first authorizes the typed operation and the resource it targets, then narrows the query or shapes the arguments before the backend ever sees the request, then executes the backend call, then filters the records and fields that come back, then redacts any matching content or masks specific values in the response before it reaches the agent. Each of those five stages can independently halt the exchange or modify it. A call can be authorized in principle and still arrive at the agent stripped of the fields it was never meant to see.
A governance rule keeps agent-specific policy from quietly becoming more permissive than the data owner intended. OBPE proves that agent policy cannot widen the owner's ceiling as a formal property of the system. Irreversible or high-impact actions deserve the same treatment regardless of how the five-question check resolves: anything that sends money, deletes a record, or grants new access should require an out-of-band human approval before it proceeds. Runaway loop controls sit at this same layer. Rate limits, budgets, and circuit breakers that stop an agent from retrying a failed call indefinitely and burning through quota or money are themselves authorization decisions the boundary makes, not separate operational settings bolted on afterward.
Scope Design for Slack and GitHub
A pre-action enforcement layer can only make decisions as fine-grained as the scopes the underlying API exposes to it. Where an API defines coarse scopes, the enforcement burden shifts almost entirely onto the application's middleware. Where an API defines fine-grained scopes, the token itself starts carrying real policy meaning, which gives the enforcement layer more to check against before a call ever reaches the backend.
Not every API offers that kind of granularity. An org-level admin token with broad, undifferentiated access pushes essentially all of the enforcement burden onto the application layer: the intersection logic and the query-narrowing described earlier have to compensate entirely for distinctions the token itself cannot express. Kalshi's authentication model sits at the far end of this spectrum. Kalshi authenticates requests through RSA key-pair signatures rather than OAuth scopes, so the permission boundary lives at the level of the key pair itself, with no scope tokens available to narrow further. Policy enforcement for a Kalshi-connected agent has to happen entirely at the application layer as a result, and Kalshi's demo environment gives developers a way to test that enforcement logic without risking real money while they build it.
Testing any of this requires more than a static response fixture. Verifying that scope enforcement actually works means simulating a sequence: adding a Slack channel, posting a message to it, then attempting to read that channel's history with a token that was never granted history scope. The correct behavior is a 403 on that third call specifically, which only happens if the test environment remembers the state created by the first two calls.
Testing enforcement logic against stateless mocks gives false confidence
A stateless mock answers every request the same way regardless of what happened before it, which is exactly the property that makes it blind to the authorization failures a pre-action enforcement layer exists to catch. The Slack example above only produces a meaningful test if the simulator remembers that the channel was created and that a message was posted to it before the history request arrives. A mock that returns a fixed response to every call will return the same answer to that third request whether the token has history scope or not, and whether the channel exists or not, because it was never built to track the sequence of calls that led up to it.
This matters because the failures pre-action enforcement is designed to prevent are, by definition, cumulative and sequential. Semantic gating depends on an agent's actions across a session adding up to something that crosses a threshold. Query narrowing depends on the state of a task at the moment a specific call is made. A credential that was valid for the first nine calls in a sequence can still be the wrong credential for the tenth, and a test built on a stateless response cannot surface that. Verifying an enforcement layer honestly requires testing it the way it will actually be used: against a system that carries state forward from one call to the next, the same way the agents it is meant to govern do.
Sources
- If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary
- Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents
- Runtime Authorization for Resources Acquired by AI Agents
- Authorization Architectures for Tool-Using AI Agents
- AC4A: Access Control for Agents