Paper 1 of 5
TokenCast: Forecasting Token Consumption During LLM Agent Execution
Problem
Token consumption of LLM agents varies widely across runs and grows as context is repeatedly read, making total consumption hard to predict; naive predictors that ignore context growth and compositional effects fail to capture this dynamic.
Approach
TokenCast learns a composable cost representation for each execution segment that records its own token consumption and the context growth it introduces. Adjacent segment representations are composed to yield a cumulative estimate that accounts for extra input cost when earlier context is re-read. The forecast is refreshed as new evidence arrives without additional LLM calls, incurring a mean cumulative prediction time of 32.8 ms per run. The method combines segment triples, cross‑fitting, and cost weighting within a LightGBM predictor. Interval calibration improves coverage and interval score.
Result
TokenCast achieves an average MAE reduction of 14.5% across 96 benchmark‑model‑prediction‑point combinations. On SWE-bench Verified with GPT-5.4 it lowers MAE from EGTP’s 74.6 tokens to 38.9 tokens at In‑call Update (47.9% reduction) and from TRAIL’s 115.0k tokens to 80.0k tokens at Task Update (30.4% reduction). Mean cumulative prediction time is 32.8 ms per run, and calibrated intervals cover 82.0% of outcomes versus 52.7% for Self‑Prediction.
Why it matters
Developers of LLM agents and system designers who need accurate token‑budget forecasting should consider TokenCast to reduce waste and improve budget control.
Method details
- Evaluated on six agent LLMs: GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, DeepSeek-V4-Pro, Qwen3.8-27B, Llama-3.2-3B-Instruct.
- Benchmarks: SWE-bench Verified, Search-R1, MMLU-Pro, LongBench-v2.
- Collected 11,712 execution traces from 240 tasks.
- Baselines: TRAIL, EGTP, TIE, and Self‑Prediction.
- Base predictor architecture: LightGBM, selected as the lowest‑error base.
- Update frequency: every call yields 32.8 ms mean cumulative prediction time and 19.7 forecasts per run.
Numbers
- MAE reduction average 14.5% across 96 combinations
- In‑call Update MAE 38.9 tokens vs EGTP 74.6 tokens
- Task Update MAE 80.0k tokens vs TRAIL 115.0k tokens
- Mean cumulative prediction time 32.8 ms per run
- Interval coverage 82.0% vs Self‑Prediction 52.7%
- TokenCast uses 21.3% fewer tokens than a fixed‑budget policy
Limitations
The paper does not establish effectiveness for LLM agents beyond the six evaluated models or for tasks outside the four benchmarks used.
TokenCast reduces MAE relative to the strongest comparator by 47.9% at In-call Update, from EGTP’s 74.6 to 38.9 tokensFound in the source text, word for word.
Picked because: Provides a concrete token‑consumption predictor for LLM agents, enabling cost‑aware scheduling and budgeting in production deployments.