SubscribeSign In
Agent to Product
Cost ModelingLong read

LLM Token Cost Forecasting for Agent Workflows

Why agent workflows cost vastly more than token prices suggest.

Senior Correspondent · · 10 min read
Cover illustration for “LLM Token Cost Forecasting for Agent Workflows”
Cost Modeling · October 10, 2026 · 10 min read · 2,248 words

Advertisement

ORBITAnalytics built for editors.

Per-token prices for frontier models have fallen roughly 280 times over two years, and enterprise AI spend has risen 320% over that same period. The two numbers are connected: cheaper tokens did not shrink the bill, they funded a lot more agent activity than anyone budgeted for.

Why falling token prices have not made agent workflows cheaper

Cheaper tokens changed behavior before they changed costs. Once a team could run a model for a fraction of the old price, the obvious move was to run it more: more steps per task, longer context windows, more ambitious workflows that would have looked wasteful two years ago. Consumption grew faster than price fell, and that gap is visible on the invoice. Inference used to make up 70 to 80% of total run cost for a typical mid-size AI product two years ago. Today that share has dropped to somewhere between 30 and 45%, because the rest of the bill grew up around it: hosted vector search, observability tooling, human review layered on top of agent output. A team could tune its inference spend to the last penny and still watch the total bill climb, because inference is no longer the main event. A forecast built around "what does a token cost today" starts wrong and gets wronger, because consumption, not price, is the variable nobody has under control. Agents make that variable especially hard to bound, which is the subject of everything that follows.

What The Agentic Loop Multiplier Is

A chatbot answers a question with one call to the model. An agent does not work that way. It makes a call, reads the result, decides on a next step, makes another call, and keeps going until the task is done or it gives up. Every one of those calls re-sends the full context built up so far, so the token bill does not add up step by step, it compounds.

Tool access adds a cost that has nothing to do with which tool gets used. If an agent has five tools available, the prompt for every single call carries all five tool schema definitions, whether that call touches one tool or none. A task that only ever needs a calculator still pays for the search tool, the database lookup, and the two others sitting unused in the prompt, call after call.

Then there is retrying. When a tool throws an error, returns something malformed, or a verification step fails, the agent does not just try again cheaply. It re-sends everything it has accumulated up to that point, adds the error, and generates a fresh attempt. Two or three of those retry cycles on a single sub-task can cost more tokens than the rest of the task combined.

Stanford's Digital Economy Lab found that agentic tasks consume 1,000 times more tokens than ordinary code reasoning or code chat tasks. Input tokens drove that cost, not output tokens. The model spends its budget re-reading everything it has already said and done, every single step.

The size of the multiplier tracks the shape of the agent. A light agent with two or three tools is near the bottom of the range, while a full agentic system running planning loops and multi-step reasoning is near the top. Knowing where an architecture falls on that range, before a single price gets plugged into a spreadsheet, tells a forecaster more than any token price will. A simple chat request might run a few hundred tokens. A fraud-detection task involving a transaction lookup, a risk score, a retry, and a case comparison runs many times that across its tool calls, for what the end user experiences as one answered question.

Why execution paths are stochastic and forecasting models built for deterministic requests fail

The agentic loop multiplier does more than run high. It runs unpredictably, even on the exact same task run twice. That unpredictability is the deeper problem, more than the raw size of the multiplier itself.

Stanford's Digital Economy Lab ran identical agents on identical tasks and found token costs varying by up to 30 times between runs of the same task. The same bug might get fixed in three tool calls one time and take a long, winding path the next, and there is no way to know in advance which path a given run will take. It produces a genuinely different outcome each time, from the same model on the same input.

The same study found that frontier models are bad at predicting their own token usage before a task even starts. The correlation between a model's self-estimate and what it actually spends is weak to moderate at best, and the bias runs in one direction: models tend to underestimate. Worse, task difficulty as judged by human experts barely predicts token cost either, so the instinct to say "this looks simple, it won't cost much" is not a reliable input to any forecast. A manager who sets a budget based on the median run will find that a meaningful share of runs blow past that median by a wide margin, simply because the distribution has a long, expensive tail built into how agents work.

That changes what a forecast should even try to model. The useful move is to model a distribution of loop lengths, set a hard ceiling, like a maximum number of turns, and size the budget to the shape of that distribution. The goal shifts from predicting one number to bounding a range.

Context bloat as a slow-moving cost that compounds silently over a deployment's lifetime

Run-to-run variance is the fast half of the problem. The slow half occurs over the life of a deployment, not within a single task.

As an agent accumulates memory entries across sessions, the token cost of every later call grows along with it. That growth is not steady from the start. It stays mild for a while and then gets steep as the number of stored entries climbs. Production systems have shown context windows reaching sizes far beyond what the team budgeted at launch, often within a few weeks of going live. A cost structure that looked fine in a prototype can look completely different once real users have been feeding the system for a month.

Background inference makes this worse. Monitoring agents, document watchers, and compliance surveillance systems do not wait for a user to ask something, they run continuously, consuming tokens against every event that crosses their path. In most 2024 enterprise deployments, this kind of workload barely registered. By 2026, it accounts for a meaningful and growing share of the monthly inference bill, and it cannot be throttled back without weakening the function it was built to provide.

A single field in the system prompt, labeled "Current Date & Time," changed on every turn. That one changing field invalidated the prompt cache on every request. Instead of the cheap cache reads the system was designed around, the full 170,000-token context got fully reprocessed at standard pricing on every call, with cache reads sitting at zero the entire time. The team had not made an architectural mistake. A single misplaced dynamic field was enough to erase the savings the whole caching strategy depended on.

The lesson for forecasting is direct: context hygiene is an input to the forecast itself, not a later optimization step. A forecast that does not account for how context grows over weeks and months, and how a single misplaced field can collapse expected cache savings, will drift away from reality not in a year but within weeks of launch.

The seven variables a valid agent cost forecast must track

Most teams build a cost forecast around four numbers: active users, how often they open a session, and the input and output tokens each session uses. That gets a forecast for a chatbot roughly right, and it gets an agent forecast badly wrong, because it leaves out three variables that matter just as much as the first four.

Retry rate is the first omission. Every retry cycle re-sends accumulated context plus an error and asks for a new attempt, so a workflow with a high retry rate costs meaningfully more than its happy-path estimate suggests, even before anything else changes. Cache hit rate is the second. How often a call actually benefits from prompt caching, rather than reprocessing the full context at standard pricing, decides whether the architecture gets the savings it was designed around or quietly loses them the way the OpenClaw system did. The agent loop multiplier is the third, and the one teams skip most often. It captures how many LLM calls a single user request actually triggers, and where an architecture lands on that range depends on how many tools it carries and how many planning or reasoning steps it runs before producing an answer. A light agent with two or three tools is near the bottom of that range; a full planning-and-reasoning system is near the top. Setting that variable correctly, based on the architecture itself, has to happen before any price gets plugged into the model. Put those three together with the four everyone already tracks, active users, session frequency, input tokens per session, and output tokens per session, and the forecast finally has a shape that matches how an agent actually spends money: not as a single number per task, but as a set of multipliers stacked on top of a baseline, each one capable of doubling or tripling the final bill on its own.

Three durable ratios that keep a forecast valid when price tables go stale

A forecast built directly on today's price table expires the moment prices change, and prices change often. The fix is to build the forecast on ratios that stay roughly stable even as the prices underneath them move, and treat the current price as a single number plugged in at the end, not as the foundation the whole model rests on.

The first ratio is the gap between flagship and economy-tier models. Flagship models cost far more per token than economy-tier alternatives, and that gap is wide enough that routing even a majority of requests to economy-tier models produces a sizable drop in blended cost. A Q1 2026 enterprise analysis found this kind of saving available to teams willing to route intelligently instead of defaulting to the most capable model for every call. What matters for a forecast is the ratio between the prices on either end of that gap, since that ratio holds even after the next round of price cuts.

The second ratio is the agentic premium over simple retrieval. Agentic workflows consume several times the tokens of a simple RAG setup, a single retrieval call followed by one generation step. Any estimate built from a chatbot prototype or a single-call demo needs a large multiplier applied before it means anything for a true multi-step agent, and that multiplier is the ratio, not a fixed token count that will be outdated by the next model release.

The third is the retry multiplier, and its effect on cost per completed task is large enough to change a budget. A simulation of agent costs found that a 5x retry multiplier pushed the cost of a successful task from $5.73 to $28.65. That is a five-fold difference on a forecast, driven entirely by how often the system has to try again before it succeeds.

Prices change almost weekly, and any table of model prices a team builds today will be wrong within a matter of weeks. The ratios underneath that table, the gap between model tiers, the premium agents carry over simple retrieval, the multiplier retries add to cost per task, do not expire the same way. A forecast built on those three ratios needs one update at each new model release: the current price. Everything else in the structure holds.

Diagram: The Retry Tax: How One Variable Multiplies Cost Per Task. Visualizes: Show the stark cost jump a 5x retry multiplier produces on a single agent task: from $5.73 (happy path) to $28.65 (5x retry multiplier applied) — a five-fold increase…

What real agent deployments cost

Real deployments keep failing in the same few places, and each failure traces back to one of the mechanisms already laid out. The retry tax appears when a workflow's retry rate runs higher than the happy-path estimate assumed, and the simulated jump from $5.73 to $28.65 per completed task under a 5x retry multiplier is what a moderately retry-heavy workflow looks like once the multiplier is applied honestly. Context bloat becomes visible when a team forecasts cost at launch based on a fresh, empty memory store and gets blindsided weeks later when accumulated entries have pushed context size far past what the original budget assumed. Cache invalidation occurs exactly the way it did in the OpenClaw case: a single dynamic field in a system prompt invalidates a cache that the entire cost model was quietly depending on, turning what should have been cheap cache reads into full reprocessing of a 170,000-token context on every call. The single-model default raises costs when a team routes every request, simple or complex, to the most capable flagship model available, ignoring the savings a flagship-to-economy ratio would apply.

None of these are exotic failures. They are the direct, traceable consequences of the same four or five variables a conventional forecast leaves out: the loop multiplier, the retry rate, the cache hit rate, and the pace at which context grows over the life of a deployment. A forecast that tracks those variables from the start, and builds on the ratios between model tiers and workflow types rather than on today's price table, is the only kind that survives contact with how agents actually run.

Sources

  1. How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks - Stanford Digital Economy Lab
Filed underCost Modeling

More in Cost Modeling