SubscribeSign In
Agent to Product

Retry Logic and Backoff Strategies for LLM API Failures

Rate limits are the dominant LLM failure mode, requiring exponential backoff with jitter.

Staff Writer · · 10 min read
Cover illustration for “Retry Logic and Backoff Strategies for LLM API Failures”
Agent Reliability · October 2, 2026 · 10 min read · 2,144 words

Advertisement

ORBITAnalytics built for editors.

A document summarization flow dies at 2 AM because Claude's API starts returning 529 "Overloaded" errors, and every user who opens the app that morning stares at a spinner that never resolves. That scenario is the plain fact behind this piece: LLM API failures are not random noise that a blanket retry loop can absorb, they are structurally distinct failure shapes, and each one demands its own response. Traditional REST error handling only has to answer one question, did the code fail. Check the HTTP status, catch the exception, alert on latency: those failure modes are deterministic and show up where you're looking. LLM systems break that model because they can fail while a call returns 200 and quality quietly erodes behind a clean status code.

Five properties of LLM systems create error-handling demands that standard software never had to deal with. Probabilistic outputs can fail without signaling failure at all: a hallucinated policy, a fabricated number, an invented API endpoint will return HTTP 200 and sail through a latency check. Tool calls can succeed and still be wrong, an agent that calls the right tool with hallucinated parameters completes the call just fine, and the wrong CRM record or the email sent to the wrong address is already live before anything flags an error. Rate limits, not server crashes, are the dominant transient failure: traffic analysis from February 2026 found rate limits accounted for the majority of all LLM API errors. And quality degradation from a context window overflow, a prompt injection, or a model checkpoint change can erode output with no API error appearing anywhere to catch it.

None of this is theoretical. OpenAI has suffered outages running many hours at a stretch, and Anthropic logged a substantial number of major incidents across a 90-day window in early 2026, with a median resolution time over an hour. A product built on a single LLM provider, with no retry logic and no fallback path, is signing up for significant downtime over the course of a year. Every section that follows answers a specific piece of this problem."

The six failure shapes

Six distinct failure shapes exist, and each one calls for its own fix. A rate limit is transient and happens when a client exceeds requests-per-minute or tokens-per-minute; the right move is to retry with exponential backoff and jitter, or reroute to another provider, and to honor the Retry-After header whenever the response includes one. An overload signals provider-wide capacity strain rather than a problem with one account; the right move is a circuit breaker paired with a fallback, because retrying just adds more load to a provider that's already buried. A connection error, a network failure, is genuinely transient and calls for retry with jitter. Context exceeded is not transient at all: sending the same request again returns the same error every time, so the fix is to truncate the input or switch to a model with a larger context window. And a quality failure, where the API hands back a 200 but the output is hallucinated, refused, or malformed, needs a validation layer sitting in front of tool execution rather than a retry keyed to a status code.

One rule holds across all of this: never retry a 4xx error that isn't a 429. A 400 bad request, a 401 auth failure, a 413 payload too large, none of these resolve themselves with time. Retrying them burns throughput and hides the actual bug that needs fixing.

The quality-failure row does the most damage in production, by a wide margin. A Superface study found that AI agents attempting simple CRM tasks failed as often as 75% of the time across repeated runs, and the cause wasn't API errors, it was hallucinated actions, schema violations, and tool misuse, the kind of failure that HTTP monitoring never catches. Andrew Wheeler documented a case from production in February 2026: Google Maps grounding silently returned the message "I cannot find any google maps data right now," with no programmatic error anywhere for a retry loop to catch.

Classification has to come before any of the mechanics: retry, reroute, circuit-break, truncate, or validate. Each of the sections below covers exactly one of those five responses, matched to the failure shape that requires it.

Diagram: Six LLM Failure Shapes, Five Responses. Visualizes: Visualize the six distinct LLM failure shapes mapped to their correct responses.

Exponential backoff with jitter as the baseline for transient failures

For transient failures, the right retry mechanism is exponential backoff with full jitter. That's not a matter of elegance, it's because naive retry (hitting the same endpoint on a fixed schedule) compounds rate-limit failures into a thundering-herd cascade instead of resolving them.

The mechanism behind thundering herd is simple to trace. A provider returns 429, every client retries on the same fixed schedule, and they end up synchronized: all of them hit the just-recovered limit at the same instant, triggering another 429, and the cycle repeats. OpenAI's own documentation states that unsuccessful requests still count against rate limits, so every failed retry consumes quota that a successful call would have used more productively. In a multi-agent setup, this amplifies faster: several agents making calls per second can exhaust a quota in well under a minute, then all retry at once, re-exhausting the limit the moment it recovers.

The fix, in practical terms drawn from production as of 2026, has these knobs. Base delay runs a second or two. The multiplier doubles the wait on each attempt. Jitter adds a random value between zero and the base delay, breaking the synchronization that causes the stampede. Max delay caps out at a minute or two. Max attempts runs several rounds for background work. For 429s specifically, read the Retry-After header first and honor it; fall back to the backoff policy only when that header is missing, and add jitter regardless, since even a small random spread across concurrent requests is enough to prevent the synchronized stampede. For 5xx errors, the provider isn't telling you a specific wait time, so exponential backoff capped at a sane ceiling is the right default. For interactive, user-facing requests, keep the early retries in the hundreds of milliseconds and cap the attempt count low. Users abandon a stalled request long before a background job would.

The cost of skipping this step isn't abstract. One developer reported burning through $87 in GPU hours before landing on a working retry approach, because unbounded retries with no backoff turned a routine rate-limit event into a billing problem. None of this requires custom engineering to fix. Python libraries like tenacity and backoff implement jittered exponential backoff as a one-line decorator, so the mechanism costs almost nothing to adopt once you know to reach for it. Token-aware scheduling goes one step further: put a queue in front of LLM calls and process work with a bounded worker pool, so a short prompt and a large context aren't competing for the same slice of quota on equal footing.

Circuit breakers and the overloaded-provider problem

Backoff solves the problem of waiting the right amount of time before trying again. A separate problem remains unsolved: recognizing that a provider is not coming back anytime soon, when the correct move is to stop trying it for now. That's the job of a circuit breaker: it detects sustained degradation and halts the hammering before the failure spreads further than it already has.

A circuit breaker trips after a set number of consecutive failures, opens for a reset window, and during that window sends traffic to a queue instead of attempting live calls. Once the reset window closes, the breaker moves to a half-open state and lets exactly one probe request through. If that probe succeeds, the circuit closes and normal traffic resumes. If it fails, the window resets and the wait begins again.

Claude's 529 "Overloaded" error is the clearest case where this matters. A 529 signals capacity strain across the whole provider, not a limit tied to one account, so retrying it with exponential backoff just adds more load to a system that's already saturated. The correct response is to trip the circuit and reroute, not to wait longer and try the same door again.

Circuit breakers that only watch HTTP status codes miss the failure mode that does the most damage: silent quality degradation, where outputs grow less coherent over a series of calls while every one of those calls still returns 200. A circuit breaker worth building tracks validation-gate failure rates alongside HTTP error rates, and trips when the validation failure rate crosses a threshold, even while the status-code stream looks perfectly healthy.

The Evaluator-Optimizer agent pattern shows how badly this goes without a breaker in place. When the evaluator's quality bar can't be met, the system doesn't fail cleanly, it loops, retrying indefinitely in pursuit of a bar it will never clear. A single runaway task built this way can burn through a substantial share of an API budget in one afternoon.

Fallback chains: routing around failures rather than waiting them out

When a provider is down, or rate-limited past any reasonable recovery window, waiting isn't the answer. Routing the request to a provider that can serve it right now is. Building that routing well means separating three fundamentally different scenarios that most tutorials flatten into a single "fallback" bucket.

Three distinct categories do the actual work. Transient-error fallback handles a 429 or a 5xx by routing that one request to a secondary provider, with the primary resuming once its circuit closes. Capacity fallback handles a 529 overload by routing to a provider with throughput to spare, and the ordering of that chain should reflect which providers actually have higher ceilings. Google's paid tier, for example, offers substantially higher tokens-per-minute for input than OpenAI's published throughput, which makes it a reasonable higher-capacity rung in a fallback chain. Context-window fallback handles a context-exceeded error by routing to a model with a larger context window, not to the same model under a different account, since the same model will hit the same wall. LiteLLM's typed fallbacks make this distinction explicit in configuration, routing general errors (429s and 5xx included) to one backup path and context-window errors to a different one.

Tier differences between providers change how a fallback chain should be ordered. Anthropic's Tier 1 limits run tight, while Tier 4 reaches substantially higher throughput, so the right fallback ordering for a Tier 1 account looks different from the right ordering for a Tier 4 account on the same provider.

Vespper, an incident response platform, hit this problem directly when its OpenAI account was deactivated without warning. Using LiteLLM's fallback feature, the team built automatic switching between providers, and in the process of implementing it, found and fixed a bug in LiteLLM's fallback handling, which they contributed back to the open-source project. FutureAGI's 2026 evals found that agent runs with a configured fallback chain saw a notably lower rate of task abandonment during simulated provider outages, largely because the agent could still complete its planning step on a fallback model even while the primary reasoning endpoint stayed down. Gateway layers like LiteLLM, Portkey, Bifrost, Cloudflare AI Gateway, Vercel AI Gateway, and Kong AI Gateway all support fallback on failure.

One fallback costs nothing at all: semantic caching. If a semantically similar query was answered recently, serve that cached answer, with no API call, no added latency, and no dependency on an external provider at that moment. It is the cheapest fallback available, and it belongs ahead of every other option in the chain.

Idempotency and compensating actions in multi-step agent workflows

Everything above deals with a single call. Multi-step agent workflows raise a harder problem: a retry that behaves correctly at the HTTP layer can still do real damage, because the question isn't whether the call succeeded, it's whether the side effects from the steps before it are safe to run a second time.

Trace the mechanism through a concrete case. A five-step workflow fails at step 4 and the system retries from step 1. Steps 1 through 3 run again. If any of those steps carries a side effect, a charge processed, a record created, an email sent, that side effect fires twice. The API layer sees a clean retry and a successful completion. The CRM sees a duplicate record it was never supposed to have, and the customer sees two charges on a statement that should carry one.

This is the reason idempotency keys matter more in agent workflows than in a typical REST retry policy, which only answers whether the code failed. An idempotency key lets a step recognize that it already ran and skip re-executing the side effect, even when the surrounding workflow retries from an earlier point. A refund action undoes a charge. A delete action undoes a record creation. A cancellation email undoes a confirmation that should never have gone out.

Retry logic answers the question of whether a single call eventually succeeds.

Sources

  1. LLM Error Handling and Fallback Strategies for Production
  2. AI Error Handling Patterns 2026: Circuit Breakers, Retries & Fallbacks for LLMs
  3. ry-ops

More in Agent Reliability