Testing OAuth Token Refresh and Expiry Scenarios Locally

Token expiry goes unexercised in tests built on real timers.

Senior Correspondent · · 10 min read
Cover illustration for “Testing OAuth Token Refresh and Expiry Scenarios Locally”
OAuth and SSO Flows · October 7, 2026 · 10 min read · 2,179 words

An access token expires on a timer that most test suites never wait for, so code that looks correct in development can fail the moment it meets real traffic. It's a scheduled event every integration will hit, repeatedly, for as long as it runs. A local test that completes in a few seconds will never touch that boundary, so the test passes while the part of the system that handles expiry and refresh goes completely unexercised.

The scheduling varies by provider, and none of it is something a development team gets to set. Google access tokens expire after 3,600 seconds, one hour, and Google refresh tokens carry six distinct expiration conditions of their own, the sharpest being an automatic seven-day expiry that applies whenever the app's consent screen sits in "Testing" status. Slack's MCP Server, when token rotation is turned on, issues access tokens with a limited lifetime and gives back no refresh token at all, so the only way to recover from expiry is a full browser re-authentication, not a background refresh call.

Putting two or three of these providers into a single integration makes the problem concrete. An agent or automation that touches Slack, Google, and GitHub at the same time is running against three clocks that don't share a start time, a duration, or a renewal mechanism. Waiting for three independent clocks to line up, or trying to catch each one in isolation, turns expiry testing into something no team does regularly, and it mostly doesn't get done.

The failure modes that appear when tokens expire mid-run

Diagram: Two Silent Ways Token Expiry Breaks an Agent. Visualizes: Show two distinct failure paths that occur when a token expires mid-run, contrasting their visibility and danger.

Token expiry rarely throws an exception that stops a program in its tracks. It degrades output quietly, which is what makes it hard to catch in a test suite built around checking for errors.

Two patterns appear most often in token-expiry failures. The first is a silent skip: a call made with an expired token returns null or an empty payload, and the calling code interprets that as "no data found" rather than "the credential failed," so the run finishes and reports success with a gap in its output that nobody flagged. The second is more specific to LLM-driven agents, and it's arguably the more dangerous of the two. The model receives error text back from the failed call, treats it as information to reason around rather than a hard stop, and generates a plausible next step that routes past the broken tool. The agent finishes the task, the output reads as coherent, and it's wrong in a way that's difficult to trace back to a token that expired three steps earlier.

Revocation needs to be tested as its own condition, separate from ordinary expiry, because the two don't behave the same way. A failed access token is recoverable: refresh it and retry the call. A failed refresh token exchange cannot be recovered by the application. The user has to re-authorize from scratch. How much exposure a given run has to that second failure mode depends on how long the run lasts: a twenty-minute task is unlikely to cross paths with a revocation event, but a multi-hour or multi-day agentic run has a much wider window in which a revoked grant can land mid-task.

Parallel test execution has a failure mode that single-threaded testing never surfaces. When multiple workers in a CI shard share one OAuth grant, a refresh under token rotation isn't just wasteful when two callers try it at once, it's actively dangerous. The second caller presents a refresh token the first caller already invalidated, and some authorization servers read that pattern as a sign of token theft, responding by revoking the entire grant. That takes down every worker sharing the grant, not just the one that triggered it. The risk scales directly with how parallel the suite is: more shards produce more overlapping expiry windows, raising the chance two workers collide on the same refresh at the same moment. A long-running agentic process can hit all of this at once: crossing several providers' expiry boundaries over its lifetime, landing on a mid-run revocation, and triggering a rotation race if sub-agents happen to share credentials.

Why waiting for real timers and live sandboxes doesn't work

The obvious fix, just let the token expire for real and watch what happens, fails for reasons that have nothing to do with effort or discipline.

Time is the first constraint. Access tokens measured in minutes or a single hour don't fit into a CI pipeline that needs to run in seconds, and a development loop built around waiting an hour for a token to die isn't one a team will run more than a handful of times. So either expiry testing gets skipped from routine runs, or it gets scheduled so rarely that regressions slip through between checks.

Live credentials carry risk even inside a sandbox environment. Google's own guidance on OAuth integration says credentials belong only in secure storage, never hardcoded and never committed to a repository, yet testing expiry against a live sandbox tends to mean exactly that: a real token with real scope sitting in a CI configuration file so a test can watch it expire. If that test logs the token at any point, for instance printing it before and after a refresh call to confirm the exchange worked, that credential now lives permanently in CI logs.

Provider-side sandbox behavior adds a layer of confusion on top of the risk. Zendesk enforces its default token TTLs on any client with no usage, or no usage in the past three months, and a sandbox client that only gets exercised occasionally for expiry tests can quietly cross that inactivity line. The test then fails for a reason that looks like a bug in the test itself, when the real cause is infrastructure drift the team never touched.

Stateless mocks remove the credential risk but trade it for a different gap. A mock that returns a hardcoded 401 after a fixed number of calls can confirm that a client handles a single failed request, but it has no concept of what rotation actually does: issuing a new token while invalidating the old one, tracking which token is current, or reproducing the exact race condition where two callers try to refresh the same grant at once. Recorded cassettes have the same problem from another direction. A cassette captures a provider's token endpoint behavior, format, TTL, and response structure, as it existed at the moment of recording. When the provider changes any of that, as Zendesk did in its February 2026 rollout enforcing new default TTLs, the cassette keeps returning the old behavior and the test keeps passing against conditions that no longer exist in production.

What a stateful local simulator needs to model

Testing expiry and refresh in a way that means anything requires a simulator that treats the entire token lifecycle as state it owns, tracking not just what each call returns but how tokens relate to each other across calls over time.

Control over the clock is the foundation everything else depends on. The simulator has to let a test move time forward on command, so a token with a sixty-minute lifetime can be pushed past expiry inside a test that finishes in milliseconds. If you don't have that, you're stuck choosing between waiting on real timers or testing with tokens that never actually reach the condition under test.

Rotation needs to behave the way a real rotation-enforcing provider behaves, not as a single swapped value. When a test calls the refresh endpoint, the simulator has to issue both a new access token and a new refresh token while invalidating the old refresh token. If a later call is made with that invalidated refresh token, it should come back with the specific error a real provider returns, invalid_grant, not a generic 401 that tells the client nothing about why the call failed.

Revocation needs to exist as its own triggerable condition, separate from expiry. A test should be able to mark a refresh token revoked independent of its expiry status, reproducing what happens when a user pulls access or an admin revokes a grant directly. The simulator's response to that revoked token has to differ from its response to a token that simply expired, because a client is supposed to handle those two cases differently: one calls for a refresh attempt, the other calls for sending the user back through authorization.

Race conditions require real concurrency handling inside the simulator, not just sequential responses stacked up. The simulator needs to process concurrent refresh requests in the correct order, letting the first caller succeed while it rejects the second caller's now-invalid token, so a test can confirm that a client's own mutex or queuing logic actually prevents the stampede. The new token set, access token, refresh token, and expiry together, needs to publish as a single atomic update. A simulator that updates the access token but still serves the old refresh token to a racing request will let a client's concurrency bug pass every test cleanly, defeating the purpose of testing it.

Fault injection belongs in the design from the start, not bolted on later. The token endpoint should accept injected latency, 429 responses, and 503 responses independently of whatever resource endpoints sit behind it, so a test can confirm that when the auth server is slow or temporarily unavailable, the client doesn't corrupt its own stored token state. On a Slack simulator project, a pull request review caught that moving a route out of the authentication middleware group had silently removed fault injection from that route. So the fix mounted fault injection directly on the specific route, and a test was written to catch exactly that regression again.

Finally, the simulator needs to start from realistic state, not empty. If a simulator has no tokens loaded at startup, every single test has to run a full authorization flow before it can even begin testing expiry, which slows every run and ties every test to the happy path regardless of what it's actually trying to check. Pre-seeded tokens with a configurable issue time and TTL let a test start exactly where it needs to, at the expiry boundary itself or already past a revocation, without reconstructing the setup from scratch each time.

Running expiry and refresh scenarios against specific APIs locally

No single expiry test fits every provider, because each one handles the token lifecycle differently, so the test has to be built around that provider's specific behavior.

Slack's MCP Server issues access tokens with roughly a one-hour lifetime and returns no refresh token at all, so the correct test confirms the client detects expiry and triggers full re-authorization. A simulator built for Slack has to model the absence of a refresh token as its own distinct state, not as a missing field to work around. Slack also enforces a short window for exchanging an authorization code before it expires, and a simulator needs to enforce that same window so a test can catch a client that waits too long to complete the exchange.

Stripe offers a useful model for what a mature approach to this looks like. A stateful fake Stripe server that remembers state across calls, so that an action taken in one call actually affects what a later call returns, is the right pattern to build toward for OAuth as much as for billing. Stripe's own sandbox environment supports moving time forward deliberately to trigger subscription state changes and the webhooks tied to them, which is the same clock-control capability an OAuth simulator needs for expiry and rotation testing.

Testing OAuth refresh in CI without a login stampede

Parallel test suites turn a single expired access token into dozens of simultaneous refresh requests if every worker shares one OAuth grant. That stampede wastes capacity on its own, and it gets worse under refresh token rotation, where one worker can replace the shared refresh token while another worker is still mid-attempt with the copy that just became invalid. Two concurrent refreshes against a rotating grant can read to the authorization server as token theft, and the server's response is to revoke the whole grant. So that takes down every worker sharing it, and a timing issue turns into a full CI failure.

The fix is to treat token acquisition as shared infrastructure with an explicit concurrency contract, not as something each test or worker handles independently. Cache a usable access token so that workers aren't each requesting their own. Allow only one refresh per OAuth grant at a time, so concurrent attempts queue. Publish the full new token set (access token, refresh token, and expiry together) as a single atomic update, so no worker ever sees a half-updated state. Where the authorization server and test environment support it, give separate workers separate grants entirely, because the safest shared token is one that isn't actually shared. A CI suite built on these four rules can run fully parallel, trigger real refresh and rotation behavior, and still finish without a single worker ever colliding with another over the same credential.

Diagram: Four Rules for Parallel CI Without a Token Stampede. Visualizes: Visualize the four concrete rules that prevent a refresh stampede in a parallel CI suite: (1) Cache a usable access token so workers don't each request their own.

Sources

  1. Best Practices

More in OAuth and SSO Flows