SubscribeSign In
Agent to Product

Alerting on Agent Idle Timeouts and Stuck Sessions

Distinguish stuck agents from idle ones by tracking progress, not just liveness.

Staff Writer · · 9 min read
Cover illustration for “Alerting on Agent Idle Timeouts and Stuck Sessions”
Agent Reliability · October 5, 2026 · 9 min read · 2,066 words

Advertisement

ORBITAnalytics built for editors.

Idle timeouts and stuck agent sessions are the most common production failure that monitoring tools miss entirely, because a stuck agent and a healthy idle agent produce the exact same signal from the outside: nothing. Effective alerting on this failure mode requires knowing the two distinct ways agents get stuck, the signals that separate a stall from legitimate silence, and the timeout instrumentation that turns absence-of-progress into something a dashboard can actually show.

Idle timeouts and stuck sessions as a monitoring blind spot

A green dashboard tells an operator nothing about whether an agent is actually working. The process responds to health checks. The channel adapter accepts messages. The model provider shows an open connection. Every probe returns healthy, yet a workflow can stop making progress and never finish on its own, a failure none of those checks are built to detect.

Standard application monitoring was built for a different class of failure. It catches crashes, it catches latency spikes, it catches memory and CPU exhaustion. None of that tooling has a concept of "absence of progress," which is the actual failure signature of a stuck agent. A model request waits forever. A tool call never returns. A polling loop can also spin indefinitely, checking a condition that will never become true. None of these throw an exception. None of these raise an exception or appear in an error log. The user just sees a channel that stopped answering.

This is a monitoring architecture problem rather than a model quality problem. The agent itself may be working exactly as designed when it gets stuck, waiting correctly on a tool call that happens to hang, retrying correctly against a provider that happens to be degraded. The failure lives in the fact that nothing in the stack is watching for the stall itself. Platforms like Agent37, which host agents for multiple customers simultaneously, face this blind spot acutely: a single stuck agent session that produces no output looks identical to a healthy idle agent to all standard infrastructure monitors, yet it silently breaks the customer's workflow while consuming platform resources the whole time.

The two distinct stuck states that alerting must handle separately

Treating "stuck" as a single state is the root cause of most missed alerts. There are two mechanically different ways an agent gets stuck, and they produce opposite observable signatures. An alert built to catch one is blind to the other.

The first is quiet-stuck: the agent produces no output. A step hangs, a request never resolves, and nothing happens, visibly or otherwise. If the system doesn't track explicit session state and a stall flag, you can't tell a working idle agent from a quiet-stuck one. When a step never settles, it doesn't throw an error or log a warning. It simply never completes, and nothing downstream knows to care.

The second is busy-but-not-progressing, and it's the more dangerous of the two because it defeats the obvious fix. An agent stuck in this state generates continuous activity, so it fools any watchdog built around the assumption that activity equals health. Retries, polling loops, and tool-call cascades all register as motion. Each new attempt resets the progress clock, so a stall detector tuned to catch silence never fires, because there is no silence to catch.

The OpenClaw community documented this failure directly. On two separate incidents and different channels, two Discord channel sessions got stuck retrying a model call for hours, and this billed a significant sum. The existing stall detection was keyed on whether the run had been quiet for a set number of minutes. Each retry reset that clock, so the session was never flagged as stalled even as it burned hours and money. A tool-misuse cascade follows the identical pattern: one bad tool call triggers another, the agent looks busy the entire time, and the loop keeps running until some external limit eventually cuts it off.

The signals that distinguish a stuck session from a legitimately idle one

Diagram: Five Liveness States That Make Stuck Sessions Visible. Visualizes: The article argues that collapsing agent liveness into a binary alive-or-dead check erases the distinction that matters, and that useful monitoring requires at least five…

The right question for monitoring to ask is whether a specific step has made progress within the window it was expected to take, not whether a process is alive. Asking that question instead of checking liveness is what makes the two stuck states detectable.

You can't catch quiet-stuck sessions with external process-level probes alone, so the agent has to report its own internal state through a session object. The process can be alive and fully responsive while one of its steps is hung indefinitely. A stall flag should mark any running session that hasn't produced a heartbeat within a defined window, so that silence becomes a prompt to check rather than an assumption of health. Liveness needs at least five distinct states to be useful: active, idle, stuck, sleeping, and unresponsive. An idle agent isn't a problem. A stuck agent is. Collapsing all five into a binary alive-or-dead check erases exactly the distinction that matters.

Catching busy-but-not-progressing sessions means decoupling the signal from activity. Progress has to be measured by forward movement in the task, not by the presence of any output or API call, a distinction that becomes critical when a managed hosting platform like Agent37 needs to tell apart a customer's legitimate long-running workflow from a genuinely stalled session that should trigger an alert. Retry count and retry rate work as progress-independent signals: a session that has made many model calls without advancing its goal is stuck no matter how recently the last call fired. Token consumption per session works as a cost-visible proxy, where a single session's spend crossing a set fraction of the daily budget counts as an incident signal on its own, whether or not any error has fired. An agent that has sat "active" on the same task for several multiples of its expected duration, with no status change, should be flagged regardless of how much API traffic it's generating.

Long-running workflows complicate this picture because they introduce a legitimate idle state that looks exactly like quiet-stuck from the outside. Real enterprise workflows spend most of their time dormant, waiting on a human signature, a shipping confirmation, or an approval gate, so the monitoring layer has to know whether silence is expected or not, a distinction that purpose-built agent hosting platforms have to bake into their observability infrastructure rather than bolt on afterward. The fix is an explicit state machine with named checkpoints. If an agent reads its current step from session state instead of replaying chat history to guess where it left off, it can't skip a step or hallucinate progress, and monitoring can then watch for steps that fail to transition rather than trying to interpret silence after the fact.

Instrumenting a timeout budget that makes stuck steps visible

Diagram: Five Layered Timeout Budgets That Name the Breaking Point. Visualizes: The article contrasts one top-level run timeout (which buries the stall) against five independently scoped budgets that each name the layer that broke: provider request…

A single top-level timeout on the whole agent run is the wrong unit to instrument against. One large ceiling lets a nested operation, buried three layers deep, consume the entire budget while everything else waits on it, and by the time that ceiling fires, it reports a generic run failure instead of naming the layer that actually broke.

The better shape is a set of smaller budgets, each scoped to one layer of the run, each bounded independently so a timeout at any layer names that layer directly. A provider request budget separates the first-token timeout from the total run timeout, because silence before any output appears is a meaningfully different signal than a stall partway through a stream. A tool budget caps each individual tool call at something shorter than the full agent run, so one slow external API can't freeze an entire turn. A polling budget makes async media or generated-content jobs return a handle, instead of holding the turn open while a job finishes somewhere else. An auth budget expires OAuth and device-code flows on a defined schedule, so a stale authorization window can't block a fresh attempt. A delivery budget caps channel retries, so a retry loop on the delivery side can't quietly eat the run's entire time allowance.

OpenClaw 2026.6.1 applied exactly this principle across its provider and plugin request paths. Retries, OAuth and device-code lifetimes, media downloads, local service probes, and generated-content polling loops are all now bounded before they can hang a run, and agents and other tool-backed runtimes recover more cleanly from interrupted tool calls, stale session bindings, compaction handoffs, and media delivery retries as a result.

Debugging an unknown timeout follows a consistent order once budgets exist at every layer. Identify which layer went silent first. Check whether that layer has its own budget. If it only inherits the full run timeout, that's the design gap to fix. Decide whether a retry or a fallback actually helps in that spot, rather than assuming one does. Keep logs that name the specific timer, the provider or plugin involved, and the recovery path that was tried. Simply increasing the top-level timeout is a reliable way to turn a one-minute defect into a twenty-minute defect, because it hides the stall for longer without fixing what caused it.

Cost and trust consequences of missing instrumentation

Missing this instrumentation produces two costs, and they compound rather than stay separate. One is direct billing: stuck sessions accrue real charges because compute keeps running until some explicit limit fires, and without per-layer bounds, that limit is either the full run ceiling or nothing. The other is trust damage from workflows that look like they succeeded while they were quietly broken the entire time.

The OpenClaw community case is the clearest documented example of the billing side. A Discord session retrying a model call for hours, across two separate incidents, billed a significant sum with no alert ever firing, because stall detection reset on every retry and the meter ran uninterrupted the whole time. Production agents need alerting on token spend that fires close to real time, when a session crosses a set percentage of its daily budget, rather than a report that arrives at the end of the day after the charge has already landed.

The trust side is less visible but costs more. What actually damages confidence in an agent product isn't a crash. A plausible-looking non-answer does the damage: a workflow the user believed was running, that was silently stuck the entire time, with a dashboard showing zero errors throughout.

Long-running workflows carry this risk in a more acute form, because their legitimate idle periods make it genuinely hard to tell when silence has crossed from expected dormancy into a real stall. HR onboarding can span two weeks. Invoice disputes stall for days waiting on a vendor's reply. Sales prospecting sequences stretch across a month of touchpoints. If the monitoring layer has no explicit dormancy gates and state-machine checkpoints, you can't tell a workflow that's correctly waiting from one that's broken and waiting for nothing. An individual agent build can't just bolt this on afterward. The platform hosting the agent carries that responsibility.

Per-user agent isolation and the alerting surface for operators

When every customer's agent runs inside a shared application environment, stuck-session alerting stops working right when you need it most. One session's stuck state becomes indistinguishable from another tenant's legitimate idle period, and a restart that fixes one customer's stuck agent risks interrupting a different customer's active work in the same environment.

Per-user isolation solves this at the architecture level, not the alerting level. One isolated, persistent sandbox per customer means a stuck session in one sandbox can be identified, restarted, and billed on its own, without touching any other tenant's state. Isolation also cleans up cost accountability: compute tracked per sandbox means a runaway session's charges attach to one customer instance instead of getting averaged across a shared pool where the signal disappears into noise.

Agent37's Cloud API can provision one isolated sandbox per user with a single API call. Each sandbox carries its own filesystem and a hard disk quota, so a stuck-step watchdog running in one sandbox can't interfere with another tenant's session state even if both are experiencing problems at the same time. That isolation is also what makes a richer alerting surface possible. Instead of monitoring one shared agent process and hoping a global health check catches a problem buried inside it, an operator can surface per-instance heartbeat state, per-instance token spend, and per-instance step-transition history, and alert on one specific customer's stuck session without that signal getting masked by everyone else's healthy traffic running alongside it.

Sources

  1. Host any agent harness | Agent 37
  2. AI Agent Sandbox Pricing & Provider Comparison | Agent37

More in Agent Reliability