SubscribeSign In
Agent to Product

Migrating a Single-Instance Agent Prototype to Multi-Tenant

Scaling an agent prototype requires rearchitecting isolation, state, and credentials first.

Staff Writer · · 12 min read
Cover illustration for “Migrating a Single-Instance Agent Prototype to Multi-Tenant”
Multi-Tenant Architecture · September 9, 2026 · 12 min read · 2,625 words

Advertisement

ORBITAnalytics built for editors.

A single-instance agent prototype works because it cheats. One memory store, one credential set, one execution thread, one list of tool permissions. That works fine until a second tenant shows up, and then it doesn't degrade politely. It collapses: sessions bleed into each other, credentials cross wires, tool calls land in the wrong customer's context entirely.

Gartner projects that over 40% of agentic AI initiatives will be canceled by 2027, pointing to escalating costs as a leading driver. Most of that failure starts exactly here, in prototypes that got customers before they got re-architected. Blaxel's production deployment guide breaks the resulting damage into five categories, and they don't stay in their lanes. Latency compounds across multi-step workflows until users feel it before any dashboard does. Weak isolation turns a single bad prompt into an incident response call. State management has no shared benchmark for agent workloads, so teams find their limits the hard way, under real load. A 99.99% uptime target leaves about four minutes of downtime a month, which is nothing once flaky tools start needing retries. And token costs vary meaningfully between cached and standard pricing, a gap that compounds quickly across conversational workloads.

These five don't sit in isolation. Latency causes retries. Retries drive up token spend. Rising spend puts pressure on the reliability budget. Treat the five as one system, because that's how they behave in production. None of this is a feature to bolt on later. It's a full architecture decision, and it has to happen before any multi-tenant code gets written.

What actually changes when you go from one agent to many

Four things break independently when tenant count goes from one to two, and each needs its own fix.

State management is the first casualty. A prototype keeps state in one process, maybe one shared database table. Multi-tenant needs state that's isolated per customer, persists through restarts, and can be recovered without touching anyone else's data.

Isolation boundaries come next, and prototypes usually have none. Code, memory, and tool execution all share a single process. Multi-tenant needs real walls between tenants at the compute layer, the data layer, and the network layer, enforced, not assumed.

Credential scoping is the third piece. One API key covering every tool call works fine for a demo. It falls apart the moment two different customers need the agent to act as themselves, with their own OAuth-connected Gmail or Slack or GitHub account, and no developer standing between them and a shared token.

Then there's provisioning logic. In the prototype, a founder starts the thing by hand. That doesn't scale past a handful of paying customers, and it shouldn't have to: provisioning needs to trigger automatically the moment a new customer signs up.

Vamshidhar Pandrapagada's multi-tenant infrastructure blueprint names the underlying mistake plainly: cramming the inbound API handler, the LLM client, the session store, the tool execution engine, and the scheduling logic into one service. It works fine on stage. It falls over at scale. The fix is decomposition into five services that deploy independently: a Gateway for ingress, an Orchestrator for coordination, a Scheduler for anything time-based, a Memory service for state, and a Sandbox for isolation.

Salesforce's BYOP platform is the cautionary tale here. A shared, monolithic planner locked code, infrastructure, and deployment together so tightly that a regression in one agent could take down every other agent on the platform. More than 100 engineers ended up bottlenecked behind a single centralized CI/CD pipeline, where any deployment needed coordination across the whole team just to ship safely.

Order matters, too. Isolation and state come first. Credential scoping can't be layered on top of provisioning logic that's already locked in, because the credential store has to be keyed the same way the state store is.

Choosing an isolation model before writing any code

Three models actually get used in production, and AWS's prescriptive guidance plus the Axonius case study lay them out clearly.

The silo model gives each tenant a dedicated runtime instance. Isolation is about as strong as it gets: each session runs in its own compute context, no shared process state anywhere. Authorization gets simple too, since IAM policies (or the equivalent) applied to the runtime become the only enforcement point needed, no application-level routing logic required. Each tenant's runtime can even run a different agent version or model configuration. The cost shows up in provisioning latency at onboarding, in the operational load of running many separate runtimes, and in hard scale limits (AWS defaults to a quota of 1,000 agents per account, adjustable on request). Silo fits regulated industries best, cybersecurity, healthcare, finance, where data residency rules and per-tenant customization justify paying for dedicated infrastructure.

The pool model puts every tenant on one shared runtime, separated structurally rather than physically. Tenants authenticate through OAuth 2.0 (Amazon Cognito is a common choice), JWTs carry a tenant claim, and the runtime validates the token before the agent code routes tool calls accordingly. It's operationally simple, one runtime to deploy and watch, and onboarding is fast since there's no infrastructure to provision. The tradeoff is that tenant separation now lives entirely in application code. A routing bug isn't just a bug anymore, it's a security boundary failure. Pool fits agencies serving SMBs with similar needs, or early multi-tenant products that haven't hit demand for per-tenant customization yet.

The bridge model splits the difference: shared runtime, but every tool call routes through a gateway with two enforcement layers stacked on top of each other. Cedar policy rules handle deterministic access control. Lambda interceptors handle dynamic validation, including token exchange through sts:AssumeRole. Enforcement happens at the infrastructure level, so even sloppy agent code can't cross a tenant boundary, and the setup gives centralized governance with real audit trails. The cost is complexity: VPC connectivity, gateway configuration, interceptor functions, tenant mappings, Cedar policies, all of it has to be built and maintained. Bridge is the right call for most agencies and SaaS operators. It gets the efficiency of shared infrastructure without inheriting shared security risk.

Axonius picked silo, and for good reason: dedicated VPCs per customer was already their model, and cybersecurity asset data is sensitive enough to demand it. They built microVM-level session isolation, JWT authentication, and IAM role tagging for cost allocation, and cut what was estimated as an eight-week build down to 10 days.

A practical move for operators serving a mixed customer base: offer isolation as a pricing tier. Pool for the standard plan, silo for enterprise or regulated accounts. Either way, this decision has to happen before state schema design starts, because the isolation model dictates where and how tenant state actually lives.

Redesigning state management for per-tenant persistence

Prototype state usually lives as one shared object or one database table, no tenant key attached, no expiration policy, no path to recover it if something breaks.

Multi-tenant state needs three things the prototype never had to think about. Every state record needs a tenant identifier attached and enforced on both read and write, since a session ID by itself isn't a safe boundary. State also has to survive restarts and resume exactly where it left off. Salesforce's BYOP platform handles this with Redis-backed session management through an SDK that persists the state machine, the action queue, and any pending confirmations, so individual teams never have to hand-roll their own Redis client, serialization logic, or TTL handling. And every tenant needs its own cleanup policy: orphaned state from a customer who churned six months ago is still costing money to store, so expiry rules need to get defined at schema design time, not discovered later as a surprise line item.

The WaiiPlanner build inside Salesforce's BYOP platform is a good look at what this actually takes in production: an 8-state ReAct machine spanning 7 intent domains, where every state transition persists across user interactions, async jobs get polled with streamed progress updates during operations that run several minutes, and any write action requires explicit confirmation before it executes.

Memory deserves its own warning. Composio's analysis of agent failures in production names bad memory management, what it calls "Dumb RAG," as one of three leading causes of failure, alongside brittle connectors and polling-based architectures. And even inside a pool model, sharing one database doesn't mean sharing data: Microsoft's research on multi-tenant AI systems notes that row-level security exists as an option but isn't typically used as the default isolation method, which is worth knowing before assuming the database is handling it for you.

Whatever schema gets decided here constrains the next step directly. A tenant's credential store has to be keyed the same way as their state store, or the two systems end up fighting each other.

Scoping credentials and tool permissions per tenant

The prototype's credential setup is usually one API key covering every tool call, meaning the agent talks to Gmail or Slack or GitHub as the developer, never as the actual customer.

Multi-tenant flips that. The agent has to act as the right person for every single tool call, and no developer should be manually storing or rotating anyone's tokens to make that happen.

Composio's connected accounts model is the pattern most teams reach for. An operator generates a Connect Link, the user authenticates against whatever service they're connecting (Gmail, Slack, GitHub, Notion, the list runs long), and Composio stores that OAuth credential under the user's own connected account. From there, Composio handles token refresh and makes the actual API calls, so the developer never touches an individual token directly. Composio supports over 1,500 integrations through MCP or direct API access, with both the Python package composio and the TypeScript SDK @composio/core actively maintained. Internal APIs that aren't public integrations can get wrapped into the same managed auth layer too, through Composio's custom-actions support.

Under the bridge model, credential enforcement gets even tighter: a REQUEST interceptor pulls the JWT, looks up which tenant it belongs to, and exchanges it for a short-lived, tenant-scoped IAM credential through STS AssumeRole. Even an agent that's misbehaving can't reach past its own tenant's credentials, because it physically doesn't have access to anyone else's.

The underlying rule is minimum permission. Tool access boundaries need to get defined before provisioning starts, and read-only should be the default for anything that doesn't strictly need write access. Every write action needs to be idempotent, no exceptions, because duplicate tickets, duplicate refunds, duplicate emails, or duplicate deletes are the kind of mistake that's expensive to explain to a customer.

Composio holds SOC 2 and ISO 27001:2022 certification for its managed cloud service, which matters for operators who need to show enterprise customers their supply chain is accounted for. And third-party tool marketplaces carry their own risk that belongs in this conversation, not off to the side. A Koi Security audit of OpenClaw's skill marketplace found 341 malicious entries out of 2,857 ClawHub skills scanned, with broader scans turning up more than 800 additional suspicious entries in the same stretch. Vetting third-party integrations isn't a separate task from credential scoping. It's part of the job.

Building a provisioning layer that runs without a founder in the loop

The prototype's provisioning story is usually a founder SSHing into a box, starting the agent by hand, setting environment variables one at a time. That works for the first three customers. It stops working at the fourth or fifth.

Multi-tenant provisioning has to fire automatically the moment a customer onboards, with no hand-configuration in the loop. The right unit here is one isolated, persistent agent per customer, created the same way every time.

A provisioning call needs to do a specific set of things, in order: spin up an isolated runtime, container, microVM, or managed sandbox, scoped to the new tenant; apply that tenant's configuration, model choice, tool access list, memory backend, connected accounts; register the tenant in a central agent registry that tracks which agents exist, which versions they're on, and who has access; and return a tenant-scoped endpoint, since every agent needs its own address rather than sharing one.

Client-level deployment guidance lays out the architecture that supports this: an agent registry, a tenant management layer handling authentication, authorization, resource quotas and billing, and an execution environment that stays tenant-aware throughout. Speed matters at this layer more than it seems like it should. Reported benchmarks put microVM boot times around 125 milliseconds with memory overhead of 5 MiB per instance, well within production tolerance. Blaxel Sandboxes go further, resuming from standby in under 25 milliseconds with full memory state, running processes, and the filesystem intact, then dropping back to standby within 15 seconds of going idle.

Anyone on the silo model needs to plan around AWS's default 1,000-agent quota per account well before hitting it, since adjusting it takes lead time. And provisioning has to bake in resource quotas and request throttling per tenant from day one, or a single noisy customer can degrade service for everyone else. Salesforce's BYOP platform got its per-request overhead down to as low as 5 milliseconds across thousands of active sessions and tens of thousands of daily requests, because each reasoning engine scaled on its own rather than competing for shared capacity.

One more thing worth building in early: the provisioning API needs to support the operator's own branding. Tenants should see the operator's product on screen, not the infrastructure sitting underneath it.

What multi-tenant observability requires that prototype logging ignores

Prototype logging catches application errors and calls it a day. Production agent monitoring has to catch tool calls, model calls, retries, handoffs between agents, and the actual finish reason behind every completion, none of which standard APM tooling was built to see.

Distributed tracing for agents needs to cover four things at minimum. Every tool call needs a tenant ID attached so a failure can get isolated without sifting through every other tenant's sessions. Multi-step workflows need latency visible at each individual step, since latency that compounds across five steps disappears completely in an aggregate metric. Token cost needs tracking per tenant, because the gap between cached and standard pricing adds up meaningfully for a conversational agent, and it has to map to billing and quota enforcement. And retries need their actual failure reason logged, not a flat success-or-fail flag, since retries both inflate cost and hide the root cause if nobody's looking.

Salesforce built observability and distributed tracing into BYOP's architecture from the start, not after scaling problems forced the issue. That's the order it should happen in, every time. Axonius's approach to cost allocation, tagging each agent runtime with its tenant identity, is worth copying too: it makes infrastructure spend attributable per customer, which means accurate billing and early warning before a cost anomaly turns into a real problem.

Gartner's list of reasons agentic AI projects get canceled includes escalating costs, unclear business value, and inadequate risk controls, poor governance chief among them. Observability that can't answer which tenant, which tool call, which cost, in that order, is a governance gap. That's not a minor detail. It's the difference between a project that survives its first audit and one that doesn't.

The managed hosting path: what operators skip by not provisioning infrastructure themselves

Everything above assumes an operator is building and running the provisioning layer directly, and that's the right call when control, compliance, or deep customization actually demand it.

For operators whose real constraint is time-to-market and a small team's bandwidth, managed per-user sandboxing removes the provisioning, monitoring, and recovery work from the critical path entirely. The isolation models, the credential scoping, the observability requirements described above don't go away. They just get handled by infrastructure built for exactly this problem, freeing the team building the agent to spend its time on the agent, not the plumbing underneath it.

Sources

  1. Axonious: Multi-Tenant AI Agent Deployment with Secure Isolation Using Amazon Bedrock AgentCore - ZenML LLMOps Database
  2. Client-Level AI Agent Deployment Strategies
  3. How to Deploy AI Agents to Production: A Complete Guide
  4. Building a Multi-Tenant AI Agent Platform Handling 7K+ Sessions
  5. How to Deploy Multi-Tenant AI Agent Infrastructure That Actually Scales | by Vamshidhar Pandrapagada | Medium
  6. The 2025 AI Agent Report: Why AI Pilots Fail in Production and the 2026 Integration Roadmap | Composio
  7. blaxel.ai
  8. aws.amazon.com

More in Multi-Tenant Architecture