We replayed a month of our own agent tasks through a re-implementation of TokenCast, a method published this week that forecasts how many tokens an agent will burn while it is still running. The forecast was far better than a fixed guess. It was barely better than simple arithmetic, and as a budget policy it saved under 1% at any sensible cap. The reason was not the forecaster. 87% of our token bill is the agent re-reading what was already in its context window when the task started, and no forecast changes that.
Why we tried it
The same agent doing the same kind of job can cost ten times more on one run than the next. Our budget guard, the budget floor, handles that with fixed caps: stop a loop once it has spent a set amount. TokenCast (arXiv 2609.35760) claims something smarter. Forecast the final bill after every call, stop early when a task is clearly heading over budget, and in their replay use 21.3% fewer tokens than a fixed budget for the same number of tasks finished. If that held for us, the budget floor would get an adaptive mode.
What we built
Their code is not out yet, so we wrote our own from the paper's description and called it
tokencast_lite so nobody mistakes it for theirs. Each model call is treated as a
segment with two numbers: what it cost, and how much it grew the context. Growth matters
because every later call re-reads it, so the remaining bill grows with the square of the
calls left, not in a straight line. After each call the forecast refreshes using plain
arithmetic, in under a millisecond, with no extra model calls.
We extracted 707 tasks and 6,593 model calls from our Claude Code session logs (numbers only, no content), fitted on the oldest 70% and tested on the newest 213 tasks. Two baselines: a fixed guess made before the task starts, and a naive one that takes the average cost per call so far and multiplies by the expected number of calls.
What we found
- Forecasting works. Three calls in, the fixed guess is off by a median of 62%. Both forecasters are off by 35 to 38%.
- The clever part does not.
tokencast_litebeat the naive sum by three points after three calls, and lost to it after ten. - The budget saving is small. 9.5% fewer tokens at a cap so tight that 40% of tasks get stopped. Under 1% at every cap that lets roughly 70% or more finish. Weighted by price, 2% at best.
The explanation was in the traces. A task here starts with a median of 152k tokens already in the window: system prompt, tool definitions, memory and, in a long-running session, the earlier conversation. Each call adds about 1k. So the bill is call count times starting context, and the method's whole advantage, modelling context growth, has almost nothing to model. To check this was the shape and not a bug, the demo runs both forecasters on synthetic traces. Our shape: a tie. A coding agent's shape (small start, big growth per call): error after five calls drops from 52% to 39%.
So the lesson is to look at the shape of your own traces before adopting a paper's method. We are not adding an adaptive mode to the budget floor. The bigger lever is the 152k tokens every call re-reads. That gets cheaper by shrinking it, not by predicting it.
What's still off
This is a re-implementation from an abstract. Their learned version is probably richer than ours, and SWE-bench traces grow far more than ours do, so this says nothing against their 21.3%. Our replay also assumes that stopping a task early costs nothing, which is generous, because someone has to pick the work up again. And one Claude Code turn is a rougher unit of "a task" than one SWE-bench issue.
What's now in the stack
extract.py: turns Claude Code session logs into per-task token traces, numbers only.tokencast_lite.pyandreplay.py: the forecaster, two baselines, forecast error at checkpoints and a budget replay.- On GitHub, with a demo that runs with no data and no keys. Point it at your own logs before you buy anyone's token forecast, ours included.