Health Check Endpoints for Hosted Agent Instances
Complete health checks must verify the LLM, tools, memory, and output quality.

Advertisement
A hosted agent instance running and a hosted agent instance working are two different claims, and most health checks only test the first one. This article lays out what a complete health check for a hosted agent must verify: the LLM API it depends on, the tools it calls, the memory it reads and writes, and the quality of what it actually produces.
Why web service health checks fail for agents
A conventional liveness probe asks one question: is the process alive? For a stateless API, that question is enough on its own, because a stateless API has nothing else going on. If the process is up and answering requests, it's doing its job.
An agent is a different kind of thing. It's a reasoning engine, a tool executor, a stateful memory system, and often a communication gateway across several channels at once, all running inside what looks, from the outside, like a single container. That container can report itself as perfectly alive while the LLM API behind it is unreachable, while its memory state has quietly corrupted, while a tool credential expired an hour ago, or while the agent is looping on the same failed action over and over. None of that trips a liveness probe, because a loop, a corrupted memory state, or an expired credential leaves the container reporting itself as alive.
Web services fail loudly. A broken endpoint returns a 5xx, and the system watching it reacts immediately. Agents can fail in a much quieter way: a plausible-sounding response comes back to the user, generated from stale context, backed by a broken tool, or produced by a model that silently got swapped in for the one it was built to use. There's no server error status to catch, because nothing actually crashed. An agent workflow can break because a model endpoint is unavailable, because retrieval comes back empty, because a tool times out, because an upstream API changed its contract, because authentication expired, because the model handed back invalid parameters, because an external system rejected an action, or because the agent got stuck repeating itself. Every one of those is a real failure, and none of them appears as a container-level problem.
That gap has a direct operational cost. An operator whose health check returns 200 while any of those conditions is true believes their customer's agent is working fine, and it isn't. The health check said yes, and the agent said nothing useful to the person waiting on it. Closing that gap means building a health check that verifies the full chain the agent depends on, including the box it runs in.
The stateful/stateless split determines which checks are even relevant
Before deciding what to check, an operator has to know what kind of agent they're running, because stateless and stateful agents need different things verified.
Stateless request-response agents, the kind that classify a document or score a form and hand back a result, don't carry session state between calls. There's no history to protect.
Stateful session-based agents are a different story. Coding assistants, and personal or internal-facing agents like OpenClaw or Hermes that remember what a user said five turns ago, need liveness probes plus something more: checks on session continuity. Is the memory store reachable? Is the session state intact? If a user drops off mid-conversation and comes back, can the agent find them and pick up where it left off? A health check that never looks at that state store has left an entire layer of the system unverified, no matter how green the rest of the dashboard looks.
Event-driven, asynchronous agents, built for long-running workflows or multi-step pipelines, add a third dimension on top of both: queue health. Is the queue actually being consumed? Are workers processing at the rate they're supposed to? Are results landing where they're meant to land?
Most production systems blend these patterns rather than picking one cleanly, so a health check strategy has to work across all three rather than assume a single shape. Every hosted agent, regardless of pattern, leans on a model, on tools, on memory, and on its own output quality, and that dependency chain is what stays constant underneath the differences. That's the chain the rest of this piece works through, layer by layer.
Layer one: verifying LLM API availability and model integrity
The LLM API is the most critical dependency an agent has, and it's also the one operators are most likely to assume is fine without ever checking, simply because it's external, run by someone else, and usually reliable.
A check worth trusting has to confirm a few distinct things. Model availability: is the specific model version the agent was built around still being served, or has it been quietly deprecated or swapped for something else? And a functional response check: does a minimal probe request come back with a structurally valid answer, confirmed by more than a 200 status code sitting on top of garbage?
That last point about model integrity is where most of the real risk hides. An agent built around one model's function-calling format can get routed, mid-incident, to a model that formats tool calls differently, and start producing malformed calls with no container-level error anywhere to flag it. The infrastructure looks fine. The agent is broken.
The correct response to LLM unavailability is to mark the agent degraded or unhealthy and hold it out of serving traffic until the primary dependency comes back, not to reroute it silently to a different model without checking that model can actually do the job. A managed agent hosting platform can prevent silent model substitution specifically by enforcing failover rules that check capability compatibility before switching, and by logging which model version actually served each customer's request, so operators know, rather than guess, what model answered a given call.
Layer two: tool integration health and credential freshness
Tool integrations break more often, in production, than almost anything else in an agent's stack, and conventional health checks skip this layer.
A handful of failure patterns occur constantly here. OAuth tokens expire, and the agent's next tool call comes back with a 401 that whoever's running the thing may never see clearly. Permission scopes narrow without notice, so a token that once had write access gets rotated down to read-only, and the agent doesn't find out until it tries to act and fails.
Every API connection an agent depends on needs to be checked against the production environment before launch, and every token needs its scope confirmed as correct, but a one-time check at deployment isn't enough. That verification has to run continuously, on an ongoing basis after go-live.
Scoping tool permissions tightly matters for more than health monitoring. The OWASP Agentic Skills Top 10 names the skill and plugin execution layer as the place where most agent runtime risk clusters. Keeping permissions at the minimum level each tool actually needs shrinks both the damage a compromised credential can do and the room an attacker has to abuse a tool through a prompt injection.
Layer three: memory store availability and session-state continuity
For a stateful agent, the memory store carries as much weight as the model itself. A container reporting perfect health while its memory backend is unreachable or corrupted is an agent that greets every returning customer like a stranger.
A memory layer check needs to confirm storage reachability first: is the Redis instance, the vector database, or whatever persistent store sits behind the agent answering within a reasonable latency window? From there, it needs to confirm read/write integrity: the agent must actually write a state entry and read the same entry back correctly, beyond the store simply responding to a ping. Session continuity comes next: if a user drops off mid-conversation and reconnects, does the agent load the right context, or has that state been evicted, corrupted, or orphaned somewhere along the way? And memory expiration policy has to line up with reality: do the TTLs on short-term storage match how long a session is actually expected to run, or are sessions quietly expiring while users still expect the agent to remember them?
Designing this probe takes some care, because a slow read that times out at the probe level can mark a perfectly healthy agent as unhealthy, creating a false alarm that's arguably worse than missing a real one. The consequence of skipping this layer altogether is specific and visible to the customer: someone whose agent loses session state mid-workflow doesn't see an error message. They see an agent that's forgotten them. That's a product failure, not a line in an infrastructure log.
This is exactly where per-user isolation architecture becomes part of health check design rather than a separate concern. Stateful, session-based agents that carry context across interactions, the kind of architecture Agent37 supports through per-user persistent sandboxes, need especially strict session-continuity checks, because routing a user to the wrong instance or losing their memory state mid-conversation breaks the entire experience. In a multi-tenant setup where each customer runs their own agent instance, a memory failure in one sandbox has to stay contained to that sandbox. The health check has to be scoped per instance, not rolled up into one aggregate number across every tenant, or a single customer's bad day looks like a platform-wide outage that isn't actually happening.
Layer four: reasoning-loop detection and output-quality probes
An agent can pass every infrastructure check available and still be quietly broken: looping on a task without making progress, inventing tool parameters that don't exist, or handing back a response that's perfectly formed and completely wrong. No liveness probe and no readiness probe catches any of that, because nothing crashed.
Two distinct failure modes live in this layer. The first is the reasoning loop: the agent re-runs the same sequence of tool calls without getting anywhere, burning through inference budget and the user's patience without ever throwing an error a container could detect. The second failure mode is quieter: output quality drifting down over time, where responses pass format checks but fail on facts, on completeness, or on whether they actually solve what the user asked. Nothing about that failure looks wrong until a user complains.
What's changed is that evaluation probes built into agentic workflows can now check semantic quality, factual grounding, completeness, whether an answer is sufficient, during live inference rather than only in a test suite run before deployment. NIST's guidance describes these probe results being accumulated into a machine-readable audit trail that helps assess agent actions and outputs after the fact. MLflow's guidance on production agent observability frames the same idea around tracking faithfulness, drift, and hallucination rates over time, building a feedback loop rather than treating evaluation as a single gate passed once and forgotten.
Turning that into a health check means two concrete things. Loop detection needs a maximum-step counter on each task: if an agent blows past that threshold without reaching a clean stopping point, that is a degraded signal worth surfacing, and it trips a circuit breaker rather than a crash. Quality probes need a small canonical task with a known correct answer, run on a schedule, with the agent's response checked against that expected answer by a structured evaluator rather than a person reading it over.
There's a compliance layer sitting on top of all this too. As of August 2, 2026, the EU AI Act's transparency rules are enforceable for any system that interacts with people, requiring disclosure that users are talking to an AI and labeling of AI-generated content under Article 50. Structured logging of agent decisions and tool calls is a legal requirement under Article 12, but only for systems classified as high-risk, and standalone high-risk obligations under Annex III don't take effect until December 2, 2027. For agents that fall outside that high-risk category, this kind of logging stays a matter of good engineering practice rather than law, which doesn't make it optional in any practical sense.
Composing the layers into a single health check endpoint
A complete health check for a hosted agent is one response that pulls the status of every dependency layer into a single structured answer an orchestrator can actually act on.
That response needs to distinguish at least three states: healthy, degraded, and unhealthy. Healthy means every layer passes and the agent can take on and finish user requests without qualification. Degraded means something non-critical is impaired, a tool integration running slow, a quality probe coming back borderline, but the agent can still serve traffic; the right move here is to alert, not to pull the instance. Unhealthy means a critical layer is down outright: the LLM API unreachable, the memory store inaccessible, a credential expired. At that point the agent shouldn't serve requests at all, and the orchestrator needs to route around it or restart it.
Container orchestration platforms like ECS and Kubernetes already support readiness and liveness probes so the orchestrator knows which instances are fit to take traffic, and a structured multi-layer response maps onto that distinction cleanly: readiness, can this agent serve a request right now, corresponds to the full chain passing end to end, while liveness, is this agent alive at all, corresponds to just the process and memory layers holding up.
Circuit breakers belong above this health check, not folded inside it. When a dependency layer keeps failing, the circuit breaker is what decides to degrade gracefully instead of letting the system fail completely. The health endpoint's job is to report the state accurately, layer by layer, model, tools, memory, output quality, so whatever sits above it, a circuit breaker, an orchestrator, a human on call, has the real picture instead of a single green checkmark standing in for four separate things that were never actually checked.
Sources
- Building Production-Ready AI Agents in 2026
- Hosted Hermes and OpenClaw agents behind one API
- Health Endpoint Monitoring Pattern - Azure Architecture Center
- Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems
- AI Agents Push Humans Out of the Loop
- The Hitchhiker's Guide to Agentic AI: From Foundations to Systems


