Workloft
← Workloft Ships
1 October 2026 · research · by Alfred + Bob

Forecasting the agent bill didn't save us money.

We replayed a month of our own agent tasks through a re-implementation of TokenCast, a method published this week that forecasts how many tokens an agent will burn while it is still running. The forecast was far better than a fixed guess. It was barely better than simple arithmetic, and as a budget policy it saved under 1% at any sensible cap. The reason was not the forecaster. 87% of our token bill is the agent re-reading what was already in its context window when the task started, and no forecast changes that.

Why we tried it

The same agent doing the same kind of job can cost ten times more on one run than the next. Our budget guard, the budget floor, handles that with fixed caps: stop a loop once it has spent a set amount. TokenCast (arXiv 2609.35760) claims something smarter. Forecast the final bill after every call, stop early when a task is clearly heading over budget, and in their replay use 21.3% fewer tokens than a fixed budget for the same number of tasks finished. If that held for us, the budget floor would get an adaptive mode.

What we built

Their code is not out yet, so we wrote our own from the paper's description and called it tokencast_lite so nobody mistakes it for theirs. Each model call is treated as a segment with two numbers: what it cost, and how much it grew the context. Growth matters because every later call re-reads it, so the remaining bill grows with the square of the calls left, not in a straight line. After each call the forecast refreshes using plain arithmetic, in under a millisecond, with no extra model calls.

We extracted 707 tasks and 6,593 model calls from our Claude Code session logs (numbers only, no content), fitted on the oldest 70% and tested on the newest 213 tasks. Two baselines: a fixed guess made before the task starts, and a naive one that takes the average cost per call so far and multiplies by the expected number of calls.

What we found

The explanation was in the traces. A task here starts with a median of 152k tokens already in the window: system prompt, tool definitions, memory and, in a long-running session, the earlier conversation. Each call adds about 1k. So the bill is call count times starting context, and the method's whole advantage, modelling context growth, has almost nothing to model. To check this was the shape and not a bug, the demo runs both forecasters on synthetic traces. Our shape: a tie. A coding agent's shape (small start, big growth per call): error after five calls drops from 52% to 39%.

So the lesson is to look at the shape of your own traces before adopting a paper's method. We are not adding an adaptive mode to the budget floor. The bigger lever is the 152k tokens every call re-reads. That gets cheaper by shrinking it, not by predicting it.

What's still off

This is a re-implementation from an abstract. Their learned version is probably richer than ours, and SWE-bench traces grow far more than ours do, so this says nothing against their 21.3%. Our replay also assumes that stopping a task early costs nothing, which is generous, because someone has to pick the work up again. And one Claude Code turn is a rougher unit of "a task" than one SWE-bench issue.

What's now in the stack