Rate Limiting Agent API Calls Across Tenants
Token costs matter far more than request counts when agents share API quotas.

Advertisement
When every GET /products costs about the same, counting requests per minute is a fair way to estimate server load. That assumption held up for most of the REST API era, and it is the reason rate limiting has historically meant counting HTTP calls. Agent workloads break that assumption at the root, because the real cost of a call is measured in tokens and compute time, not in the fact that a request happened at all. A single inference call to a large language model can burn through thousands of tokens and tie up a GPU for several seconds, yet at the HTTP layer it looks exactly like a throwaway health check that costs nothing. That gap between what a request looks like and what it actually costs is the whole problem. Three things make it worse. Agentic prompts, built from memory, tool results, and conversation history, can swing from 200 tokens to 80,000 tokens within the very same tenant session, so no flat per-request ceiling can describe both ends of that range honestly. On top of that, one user action can set off dozens of model calls, vector searches, and calls to outside APIs, so a single gap in enforcement at the front door turns into hundreds of downstream operations within seconds. Agents also plan, retry, call tools, pull in context, and ask a model to check its own intermediate answers, so one user request can quietly become many model calls, each one small enough to slide under any per-request ceiling on its own. The most common failure is a loop: an agent retries, appends more context at each step, and because that context keeps growing, the token cost of each retry grows faster than the number of retries does.
Shared quotas and cascading tenant outages
Rate limits from API providers apply to an account, not to a single agent or a single tenant sharing that account. That means one runaway loop inside one tenant's workload can burn through a quota that every other tenant on the platform depends on, all at the same time. Picture a modest handful of agents, each firing off one call per second. That pace alone is enough to hit the OpenAI Tier 1 rate ceiling for gpt-4o inside of 60 seconds, before a single one of those agents has even finished its first task. Once a group of agents hits that wall together, the failure compounds on itself: each one retries after the same fixed delay, so they all come back at once, produce another burst, get hit with another wave of 429 errors, and repeat the cycle without end. There's a cost tucked inside that storm that's easy to miss. Every failed call that had already built up partial context still gets billed somewhere upstream, even though the agent gets no usable output and, because the gateway rejected the request before any generation happened, incurs no token charge on that specific call. Multiply that across customers sharing one OAuth app or one API key, and one customer running a bulk export can starve every other customer's agent conversation on the same key. No amount of good behavior at the agent level fixes this, because each agent only sees its own traffic. It has no way to know what the other agents sharing its quota are doing right now, so only something sitting above all of them, watching the whole account at once, can solve the aggregate overuse across tenants.
The fixed-window failure mode operators most commonly overlook
Fixed-window counters are usually the first thing a team reaches for, but they are the most dangerous choice for agentic traffic, because they can be hit twice in a row at the edge of a time window in a way that overwhelms GPU capacity almost instantly. The mechanics are simple enough to miss. A tenant can send a full window's worth of calls in the final second of minute N, then send another full window's worth in the first second of minute N+1. Both windows technically respect their own limit, but the tenant has just sent double the intended rate inside a two-second span. In the REST API world, where requests are cheap and roughly the same size, that kind of boundary burst barely registers. In an AI workload, the same burst can saturate GPU capacity instantly, a consequence that simply doesn't exist when the thing being limited is a cheap, uniform HTTP call. Rules built to stop that kind of burst tend to overcorrect: a 2026 analysis found that overly static rate limiting rules block an estimated 41% of legitimate AI agent traffic, while limits that ignore token size are easy to get around just by sending lots of small requests that each individually stay under the ceiling but add up to the same total load.
Token-aware rate limiting as the minimum viable first layer
Counting tokens instead of requests is the threshold below which a rate limiter provides no meaningful protection for LLM workloads. The underlying mechanism is the token bucket: a bucket holds a supply of tokens that refills at a steady rate up to some fixed cap, every call draws tokens out of that bucket, and an empty bucket means the call either waits or gets rejected. This also builds in a sensible amount of burst tolerance, since a tenant that's been idle for 10 seconds will have built up credit and can fire off a short burst without tripping the limiter. On its own, that's still just a request limiter with a clever refill schedule. The change that actually matters is setting the cost of each call to its estimated token count instead of a flat charge of one unit per call, which turns the bucket from something that counts calls into something that measures compute. A simplified version of that consume logic looks like this:
function consumeTokens(bucket, estimatedTokens):
refill(bucket) // add tokens accrued since last check, capped at bucket.capacity
if bucket.available >= estimatedTokens:
bucket.available -= estimatedTokens
return ALLOW
else:
return REJECT_429
Where that bucket lives matters as much as how it works. A single bucket shared across an entire user account is too blunt an instrument, because one misbehaving script ends up blocking all of that user's legitimate work alongside it. A bucket keyed to the combination of user, repository, and model gives enough resolution to isolate the problem without collateral damage. Most production setups need more than one layer of these buckets at once: a ceiling on total tokens per minute across an entire organization, a separate ceiling on tokens per minute for each individual agent, and a ceiling on requests per minute for each authenticated user. Where this logic runs also matters. Putting it inside each agent's own code means every team ends up writing the same logic over and over, discovering its edge cases independently and inconsistently. Put it in a gateway that sits in front of every agent, so it gets written once, with every workload behind it inheriting the same protection automatically. When the limiter does reject a call, a 429 response paired with a Retry-After header is the right signal to send, because most agent frameworks already know how to read that response as a cue to pause, wait, and retry later, which is exactly the behavior a gateway wants to encourage.
Circuit breakers and the leaky bucket as the second and third enforcement layers
A token-aware bucket controls how much volume gets through, but it has no way to notice a pattern, and it can't protect a downstream service from a burst that technically passes the bucket's check but still overwhelms a system built for steady, limited traffic. Those are two separate problems, and they need two separate tools. Circuit breakers exist to catch behavior that a volume check would miss. Instead of watching how many calls come through, a breaker watches signals like how fast costs are rising, how often the same prompt repeats, the error rate, and whether the context attached to each call keeps growing. Each of those is a sign that something is looping or misbehaving rather than that a tenant happens to be busy. When a breaker trips, the point isn't to cut a tenant off cold. A well-built fallback chain moves traffic from the primary model to a cheaper model, then to a semantic cache, and only then to a 503 response if nothing else can serve the request. A 503 at the end of that chain is a reasonable outcome. A budget-destroying loop left unchecked is not.
The leaky bucket solves a third, different problem. Incoming requests sit in a queue, and a consumer drains that queue at a fixed rate, with no burst allowance. That makes it the right tool specifically for protecting downstream services that have no room to absorb a spike, like a single-threaded embedding model or a third-party API with its own strict rate limit. The leaky bucket trades some latency for the upstream agent, since every request sits in that queue even when the system is lightly loaded, in exchange for a perfectly smooth, predictable rate downstream. Put together, the three layers each guard a different point in the pipeline. The token bucket throttles volume right where a tenant's traffic enters the system. Circuit breakers watch for bad patterns further along in the pipeline, after volume has already been shaped. The fallback chain keeps the experience usable while a breaker is open. None of the three is trying to eliminate every possible failure. The goal is to keep the blast radius of any single failure contained. A platform-wide cost ceiling sits above all of this as a last backstop: a single tenant that exceeds its own tier hits its own bucket first, and the platform-wide breaker only steps in when several tenants misbehave at once. Some operators are starting to add a further layer on top of these three, using models that look at behavioral patterns to tell a legitimate traffic spike apart from abuse, and using forecasts built from queued agent tasks to throttle intake before a queue has a chance to overflow.
Isolation at the tenant boundary: why per-tenant state and sandboxing belong in the same design
Rate limiting per tenant and running code per tenant are really the same decision, just made visible at two different layers of the system. An operator who gets the isolation boundary right at the rate limiter but leaves it loose at the execution environment has left the isolation incomplete. A compromised or misbehaving agent can still reach into another tenant's data no matter how carefully the quota is enforced, because the quota was never the thing standing between the two tenants in the first place. Per-tenant state is the starting assumption the rest of the architecture has to be built on. The same logic that says each tenant needs its own bucket also says each tenant needs its own execution boundary, and container-based isolation alone tends to fall short once agents are executing code that was generated rather than written by a developer who can be held to a review process. Treating the bucket boundary and the sandbox boundary as two separate concerns leaves a gap exactly where the system is most exposed: between where traffic is measured and where it actually runs.


