Webhook Patterns for Agent Lifecycle Events
Webhooks replace polling to eliminate latency and wasted compute in agent systems.

Advertisement
Long-running agent jobs, batch runs, hour-long sessions, background responses, break polling as a strategy. Poll on a fixed interval and you pay compute during the dead time, then still miss the moment something actually happens because your check landed thirty seconds too early. Webhooks fix the shape of the problem: the source system fires the moment there's news, and the agent wakes up only when there's real work waiting.
The latency gap is not subtle. Polling gets you somewhere between 30 seconds and 5 minutes of lag depending on interval. Webhooks land under a second. For a chat agent or a live transfer flow, that's the difference between feeling instant and feeling broken.
A developer thinking about where agent systems actually fail in production makes the point that the model isn't the bottleneck, event routing is. Everyone spends their attention on model selection, but the failures cluster around tool definition quality, retry and fallback design, and how events get routed once they fire. That last piece is exactly what webhook architecture is responsible for getting right.
Think about what a single user goal produces in a 2026-era pipeline. Retrieval fires, a tool call goes out, a fallback model kicks in because the primary is overloaded, a guardrail check runs post-generation, an evaluator scores the output. Each of those is an event, and each one might need to notify something downstream. Polling doesn't scale against that kind of fan-out; you'd need a separate poll loop per concern, checking constantly, mostly finding nothing. Events scale fine, because nothing runs until something happens.
The trigger/callback split (the two jobs every agent webhook is doing)
Every webhook in an agent system is doing one of two jobs, and they point in opposite directions.
A trigger is an external event flowing into the agent to wake it up. Stripe reports a dispute. GitHub reports a failed CI run. A calendar sends a push notification. Something in the outside world happened, and the agent needs to act.
A callback is a completion event coming back from work the agent already kicked off asynchronously. A batch job finishes. A transcription job returns. A sub-agent that got delegated a task reports back that it's done.
Five patterns appear in production agent systems, and each one sits cleanly on one side of that split:
- Provider callbacks: callback, inbound
- Event-triggered agents: trigger, inbound
- MCP async results: callback, inbound
- Live-session pushes: trigger, inbound
- Agent-to-agent and platform notifications: callback, outbound
None of the AI framing changes the delivery mechanics that actually govern behavior, as the same signature verification, retry behavior, and duplicate risk apply regardless. A checkout.session.completed event that wakes up an agent behaves exactly like one that updates a row in a billing database. Same signature verification, same retry behavior, same duplicate risk. The model call sitting in the middle of the flow is new. The plumbing around it isn't, and treating it like it needs reinventing is how teams end up building fragile, bespoke infrastructure for a problem webhooks solved a decade ago.
Knowing which of the five patterns applies tells you what to build. A trigger needs a router and a filter. A callback needs a hydrate-and-reconcile step. Get that distinction wrong early and the fix later touches every handler in the system.
How major providers have converged on thin-payload webhook design, and what that requires from you
OpenAI, Google Gemini, and Anthropic Claude built their agent webhook systems somewhat independently, and they landed on nearly the same design. That convergence deserves close reading before diving into where each one diverges.
The shared contract:
Payloads are thin. The delivery tells you an event type and a resource ID, not the actual result. You call back into the provider's API to fetch the full object. This is deliberate: if it shipped the full result in the payload and a retry fired, you'd risk processing stale data.
Delivery is at-least-once, not exactly-once, and all three providers say so explicitly. Duplicates are not a bug you might hit, they're a certainty you plan around. Dedupe on a delivery ID.
There's no ordering guarantee. Lifecycle events can and do arrive out of sequence. The right mental model is: each event is a nudge to go check current state, not a trustworthy record of what state you're in now.
Deliveries are signed and time-bounded, using HMAC or JWT signatures with roughly a five-minute freshness window, closing off replay attacks.
Retries are generous but slow, running 24 to 72 hours of exponential backoff. That protects the provider's infrastructure from a struggling endpoint, but it also means that if a bug in your handler goes unnoticed for a day, recovery is going to take a while once you fix it.
Where the details diverge, and where they bite:
OpenAI follows the Standard Webhooks spec, covering five event families: Background Responses, Batch, fine-tuning, evals, and Realtime SIP. Retries run up to 72 hours. There's no first-party CLI for local webhook testing, nothing like stripe listen, so local development requires your own tunneling setup.
Gemini also uses Standard Webhooks, but runs two separate configuration models with two separate signing schemes. Static endpoints get HMAC-SHA256. Per-job dynamic endpoints get RS256 JWTs, verified against Google's JWKS endpoint. Results don't land in the payload at all; they're made available externally once the job completes.
Claude's Managed Agent product covers agent session events and vault-credential lifecycle events, with SDK verification helpers shipped in seven languages. It also auto-disables a misbehaving endpoint: after roughly twenty consecutive delivery failures, Anthropic switches it off until someone manually flips it back on.
The thin-payload choice isn't a style preference, it's an instruction. Acknowledge the delivery fast, hydrate the full object asynchronously, and reconcile against the provider's API for anything that might have slipped through. A provider webhook is a notification, not a system of record, full stop. Verification, deduplication, tolerance for out-of-order delivery, respecting the retry window, rotating your signing secret, building a local-dev loop that actually works: all of that sits on your side of the line, not the provider's.
The ingest architecture that every pattern plugs into
Every pattern above, no matter which of the five it is, plugs into the same pipeline: external service, webhook endpoint, queue, worker pool, with a dead-letter queue catching anything that exhausts its retries.
The endpoint itself should do almost nothing. Accept the payload, verify the signature, return 200 OK, enqueue the work. The endpoint should accept the payload, verify the signature, return 200 OK, and enqueue the work, with nothing more required of it. Everything heavier happens somewhere else, asynchronously.
Calling an LLM directly inside your webhook handler races the provider's timeout clock. Retell AI's webhook spec enforces a 10-second timeout with up to 3 retries. Other providers run similar windows. LLM inference alone can regularly exceed a 10-second window. Put a model call in the request path and you'll fail deliveries that had nothing wrong with them.
What happens at each stage, in order:
Signatures get verified first, always, before anything else touches the payload. Every provider does this a little differently, so centralize that logic in one place rather than rewriting it per handler; that's how verification bugs sneak in.
Timestamps outside a roughly five-minute window get rejected outright, closing the door on replay attacks.
Deduplication happens before the event ever hits the queue. Use the provider's event ID as the dedupe key, store it in Redis with a TTL longer than the provider's retry window, and check it on the way in.
Once it clears those checks, the event gets enqueued and the endpoint returns a 2xx. SQS, RabbitMQ, Redis Streams, BullMQ, whatever queue you're running, it absorbs the actual work from here. The endpoint's job ends the moment it acknowledges.
Workers pull from the queue and process with their own exponential backoff, and anything that permanently fails routes to a dead-letter queue instead of vanishing.
HTTP has no memory, but agents need to remember what conversation or task they're in the middle of. Pass a session_id in the payload, have the worker pull history from Redis or wherever state lives, do the work, and write the new state back before it finishes.
If the queue's full, return 503 with a Retry-After header. Don't drop the event and hope the provider's retry logic bails you out.
Keep a full audit trail somewhere durable: event ID, source, payload hash, verification result, processing result, retry count. That's the difference between a postmortem that takes an hour and one that takes a week of guessing.
Cap payload size and reject anything oversized before it does damage downstream. Emit a span per webhook, tagged with source, event type, and result. Observability here isn't a nice-to-have. A misfired agent doesn't just print a stack trace, it takes an action in the world.
Idempotency is the load-bearing pattern (how to implement it correctly)
Picture the failure mode directly: an agent processes a Stripe payment_intent.succeeded event and kicks off provisioning. Works fine in testing. Ships to production, and Stripe's retry logic fires the same event multiple times because the first acknowledgment got lost somewhere. Now provisioning has run three times, and nobody notices until the API bill or the resource count looks wrong.
At-least-once delivery is the standing guarantee from OpenAI, Gemini, and Claude alike. That means duplicates aren't an edge case to shrug off, they're built into the contract you signed up for the moment you registered a webhook.
The dedupe key has to be the provider's event ID, not a hash of the payload and not a timestamp. Payload hashes miss duplicates when the payload has any field that changes between deliveries. Timestamps are worse, they don't identify the event at all.
In Redis, this looks like:
Set the key with NX (only if it doesn't already exist) and a TTL longer than the provider's full retry window, so 24 to 72 hours depending on the provider. On a cache hit, return the cached response immediately and skip re-enqueuing entirely. On a miss, enqueue the work, and once the worker finishes, write the result back to that same key so any later duplicate gets served the cached outcome instead of running the job again.
At the queue layer, add a second line of defense. BullMQ and similar systems support deduplication directly: set a custom jobId equal to the provider's event ID, and a duplicate job gets rejected outright if the original still exists, or use BullMQ's dedicated deduplication option with a deduplicationId. That means the queue itself can catch a duplicate before any worker ever touches it. This isn't a replacement for the Redis check, it's defense in depth, catching what slips past the first layer.
The pattern runs both directions. Outbound webhooks, the ones an agent sends to other systems, need the same protection. Attach an event ID header to every delivery so whatever's downstream can dedupe your calls the same way you dedupe inbound ones.
Idempotency is what makes retries safe to have at all. Strip it out and retries stop being a reliability feature, they become a multiplier on every side effect your agent produces.
Lifecycle event taxonomy (which events to subscribe to and what each one demands)
Not every lifecycle event deserves a handler, and not every one deserves the same handler. Subscribing to everything just because it's available is how ingest pipelines drown in noise. Filtering which events even reach your system, like the webhook_events field in Retell AI's agent configuration, is a production decision, not an afterthought.
High-signal events that need real handling:
request.completed means the job the agent kicked off has a result waiting. This is what triggers the hydrate-and-reconcile step, pulling the full object from the provider's API.
A guardrail-blocked event means a guardrail intercepted the model's output before it went anywhere. This needs an actual handling path, not a log line that nobody reads until something breaks downstream.
A model-fallback event means the primary model was unavailable and a standby took its place. Depending on the system, that might mean reconciling agent state that assumed the primary model's behavior.
An eval-failed event means a post-generation evaluator rejected the output. This is the fork in the road: retry the generation, or escalate to a human.
Voice and chat agents make the sequencing concrete. Retell AI's documented event types show the shape clearly: call_started or chat_started opens a session with basic metadata only. call_ended or chat_ended closes it out with the full call object, minus analysis. call_analyzed or chat_analyzed arrives once analysis finishes, carrying the full data including the analysis object, and this is usually the event that actually drives CRM updates, archiving, and alerting. transcript_updated fires turn by turn, carrying the transcript data. Transfer events (transfer_started, transfer_bridged, transfer_cancelled, transfer_ended) each carry their own payload shape and need their own handler.
Because there's no ordering guarantee across any of this, a call_ended can genuinely arrive before your handler finishes processing call_started. Build every handler to treat its event as "go check current state" rather than assuming it's step three of a sequence that started with step one.
Ignore this layer long enough and you get the auto-disable behavior mentioned earlier: Anthropic's roughly twenty-failure threshold isn't arbitrary punishment, it's what happens when an endpoint goes dark and keeps failing silently until the provider shuts it off rather than keep retrying into a void.
Retell AI's pattern lets you register webhooks at the account level or the agent level, and agent-level registration takes precedence, fully replacing account-level delivery for that specific agent. Account-level only fires for agents that never set their own. Use both layers on purpose, not because you forgot which one you configured.
The relevance-filter problem (why routing logic belongs outside the LLM)
The naive version of this pipeline routes every incoming event straight to the agent and lets the LLM figure out what matters. That doesn't solve the polling problem, it just relocates it downstream, and now you're paying LLM latency and LLM cost just to decide whether an event was worth looking at in the first place.
The fix is a thin router sitting in front of the agent. It watches the incoming event stream, checks each event against a set of registered filters, and only wakes the agent up once it's found a real match, handing over that event plus whatever context the agent actually needs.
In the canonical pipeline, this router sits right after the queue. It looks at event type and routes accordingly: a Stripe event goes to the payment-handling agent, a GitHub event goes to issue-triage, a Slack event goes to whatever's drafting notifications. The router only needs the event type, whatever metadata rides on the envelope, and a session or correlation ID if one exists. It doesn't need the full payload; that gets hydrated later, after routing, by whichever worker actually picks up the job.
This separation matters more for agents than for ordinary backend systems, because a misfired agent doesn't just throw an error into a log file, it takes an action. It sends an email, moves money, closes a ticket. Putting a fast, deterministic layer in front of the LLM to make the routing call keeps the blast radius of a misclassification small, instead of letting a confused model decide on its own what an ambiguous event means.
Register the watch and the filter once, at the infrastructure layer, and let the agent only ever see events that already matched something specific. It's the same instinct behind the thin-payload design the providers converged on: do the minimum necessary at the moment of delivery, push everything else downstream to run asynchronously.
The dead-letter queue is the router's safety net. Anything that exhausts its retries or matches no registered handler lands there for a human to look at, instead of disappearing without a trace.
MCP's Tasks primitive and the gap it leaves in inbound webhook handling
The MCP spec dated 2025-11-25 introduced Tasks, a "call-now, fetch-later" primitive for tool calls that can't return an answer right away. Instead of blocking, the server hands back a taskId, and the client polls through a defined set of states.
Work through 2026 pushed this further, moving toward MCP servers that are stateless and horizontally scalable over Streamable HTTP.
Tasks solve a real problem, but only one side of it: how does the agent handle a tool call that isn't going to answer synchronously? That's a client-side concern, and Tasks answers it cleanly.
What Tasks doesn't touch is the server's half of the same story. Something still has to reliably receive the inbound webhook announcing that the work is finished, verify it came from where it claims to have come from, deduplicate it against retries, and correlate it back to the original tool call that kicked the whole thing off. None of that comes free with Tasks. It's still entirely on the developer building the MCP server.
That gap is exactly where Pattern 3 from the trigger/callback split lives, the MCP async results pattern. A third-party API sitting behind an MCP server finishes its work and fires a callback. The MCP server has to catch that webhook and handle it using the same discipline covered throughout this piece: verify the signature, check it against a dedupe key, enqueue it rather than processing inline, and only then resolve the pending task back to the agent that's been waiting on it. Tasks gives the agent a clean way to wait. It doesn't give the server a clean way to be told the waiting is over, and that piece still has to be built by hand.
Sources
- How Developers Connect Webhooks to AI Agents, MCP Servers, and LLM Tools
- Webhook Patterns for AI Voice Agents: Idempotency, Retries, and Security
- Webhook-Driven Agent Architecture | 2026 Guide
- Retell webhooks overview - Retell AI
- Event-Driven AI Agents: Patterns That Scale
- Webhook Integration for AI Agents: Event-Driven Notifications and Callbacks
- hookdeck.com
- hookdeck.com


