arXiv digest

Tuesday

September 29, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

TokenCast: Forecasting Token Consumption During LLM Agent Execution

Chaoqian Ouyang, Ling Yue, Libin Zheng and 7 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Token consumption of LLM agents varies widely across runs and grows as context is repeatedly read, making total consumption hard to predict; naive predictors that ignore context growth and compositional effects fail to capture this dynamic.

Approach

TokenCast learns a composable cost representation for each execution segment that records its own token consumption and the context growth it introduces. Adjacent segment representations are composed to yield a cumulative estimate that accounts for extra input cost when earlier context is re-read. The forecast is refreshed as new evidence arrives without additional LLM calls, incurring a mean cumulative prediction time of 32.8 ms per run. The method combines segment triples, cross‑fitting, and cost weighting within a LightGBM predictor. Interval calibration improves coverage and interval score.

Result

TokenCast achieves an average MAE reduction of 14.5% across 96 benchmark‑model‑prediction‑point combinations. On SWE-bench Verified with GPT-5.4 it lowers MAE from EGTP’s 74.6 tokens to 38.9 tokens at In‑call Update (47.9% reduction) and from TRAIL’s 115.0k tokens to 80.0k tokens at Task Update (30.4% reduction). Mean cumulative prediction time is 32.8 ms per run, and calibrated intervals cover 82.0% of outcomes versus 52.7% for Self‑Prediction.

Why it matters

Developers of LLM agents and system designers who need accurate token‑budget forecasting should consider TokenCast to reduce waste and improve budget control.

Method details
  • Evaluated on six agent LLMs: GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, DeepSeek-V4-Pro, Qwen3.8-27B, Llama-3.2-3B-Instruct.
  • Benchmarks: SWE-bench Verified, Search-R1, MMLU-Pro, LongBench-v2.
  • Collected 11,712 execution traces from 240 tasks.
  • Baselines: TRAIL, EGTP, TIE, and Self‑Prediction.
  • Base predictor architecture: LightGBM, selected as the lowest‑error base.
  • Update frequency: every call yields 32.8 ms mean cumulative prediction time and 19.7 forecasts per run.
Numbers
  • MAE reduction average 14.5% across 96 combinations
  • In‑call Update MAE 38.9 tokens vs EGTP 74.6 tokens
  • Task Update MAE 80.0k tokens vs TRAIL 115.0k tokens
  • Mean cumulative prediction time 32.8 ms per run
  • Interval coverage 82.0% vs Self‑Prediction 52.7%
  • TokenCast uses 21.3% fewer tokens than a fixed‑budget policy
Limitations

The paper does not establish effectiveness for LLM agents beyond the six evaluated models or for tasks outside the four benchmarks used.

TokenCast reduces MAE relative to the strongest comparator by 47.9% at In-call Update, from EGTP’s 74.6 to 38.9 tokensFound in the source text, word for word.

Picked because: Provides a concrete token‑consumption predictor for LLM agents, enabling cost‑aware scheduling and budgeting in production deployments.

Paper 2 of 5

KV-streams for Efficient Compaction in Agentic Reinforcement Learning

Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda and 15 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Training agentic LLMs requires repeatedly prefilling the KV cache after each context compaction, which greatly reduces throughput.

Approach

KV-streams streams the KV cache forward instead of flushing it after each compaction, allowing the cache to act as a persistent recurrent state. It plugs into any existing compaction strategy and works with asynchronous RL frameworks that separate trainer and inference GPUs. The method pads completions to a multiple of the KV‑cache block size to avoid breaking the stream. It also modifies the trainer to apply in‑flight weight updates without flushing the KV cache. Together these changes eliminate the repeated prefill cost while preserving model performance.

Result

KV-streams provides a 2.6 to 5x wall‑clock speedup in training and does not hinder, and sometimes improves, task performance across text‑based games and software‑engineering benchmarks.

Why it matters

Researchers building agentic LLMs will benefit from KV‑streams to dramatically reduce training time without sacrificing performance.

Method details
  • Uses Qwen3-4B-Instruct-2507 for TextWorld experiments.
  • Uses Qwen3.5-4B for software‑engineering tasks.
  • Evaluates on TextWorld (60,000 synthetic games, 256 evaluation tasks) and ALFWorld.
  • Baseline comparison is the full‑context run without KV‑streams.
  • Ablations include Summary and Sliding‑Window with re‑prefill compaction.
Numbers
  • wall‑clock speedup, 2.6 to 5x, compared against full‑context baseline
  • evaluation tasks, 256, compared against synthetic game pool
  • synthetic games generated, 60,000, compared against none
  • task families, 4, compared against other task groupings
  • difficulty levels, 3, compared against other difficulty settings
  • seeds used in experiments, 3, compared against multi‑seed baselines
Limitations

The paper only runs multi‑seed experiments for ALFWorld and single‑seed experiments for the other benchmarks, limiting statistical confidence for those tasks.

KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training.Found in the source text, word for word.

Picked because: Introduces KV‑streams for context compaction, a ready‑to‑use technique that reduces GPU memory for long‑running agentic LLMs.

Paper 3 of 5

Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models

Junru Zhu, Shiming Xie, Aime Lu Fan Chen and 4 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Tool-using agents could report success after a tool failure without any evidence, and simply instructing models to be transparent does not prevent them from fabricating unsupported claims.

Approach

The paper introduces the Failure-Transparent Agents (FTA) benchmark that fixes the failed tool observation and required evidence before generation, forcing models to justify claims against a known evidence boundary. Three response policies are evaluated: a baseline, a transparency instruction, and a structured evidence contract that requires explicit evidence fields. Models receive only the user request and the deterministic failure trace; the missing evidence is held by the evaluator. Human annotators then label false-success, fabricated detail, and usefulness. The evidence contract policy structures the output to make unsupported claims auditable, reducing hallucinations while preserving useful recovery.

Result

Across the six models and three policies, false‑success drops from 22.8% under baseline to 9.3% with transparency and 0.8% with the evidence contract; fabricated‑detail falls from 28.3% to 14.3% and 0.8%; useful responses rise from 74.9% to 89.2% and 98.8% respectively.

Why it matters

Researchers and engineers building tool‑using LLM agents should use FTA to audit post‑failure reporting and adopt structured evidence contracts to reduce hallucinated success claims.

Method details
  • Six models evaluated: GPT-5.6 Terra, Claude Sonnet 5, NVIDIA Nemotron Super 3 120B, Amazon Nova Micro, Meta Llama 3.1 8B Instruct, Mistral Ministral 8B 3.0
  • 100 synthetic tasks covering five failure families, a neutral control, and four user-pressure conditions
  • Three response policies: baseline, transparency instruction, structured evidence contract
  • 3,600 human‑annotated responses (two generations per scenario‑policy cell)
  • Metrics include false‑success, fabricated detail, useful response, limitation disclosure, feasible recovery, over‑refusal
Numbers
  • false-success 22.8% baseline vs 9.3% transparency vs 0.8% evidence contract
  • fabricated-detail 28.3% baseline vs 14.3% transparency vs 0.8% evidence contract
  • useful response 74.9% baseline vs 89.2% transparency vs 98.8% evidence contract
  • 100 tasks with deterministic failure traces
  • 3,600 human‑annotated responses
Limitations

The benchmark is synthetic, English‑only, mostly one‑step, and does not include real‑world traces or partial‑success scenarios.

false-success rates are 22.8% under the baseline policy, 9.3% with a transparency instruction, and 0.8% with a structured evidence contract.Found in the source text, word for word.

Picked because: Offers the Failure‑Transparent Agents benchmark and evaluation code for verifying tool‑using LLM agents’ post‑failure reporting.

Paper 4 of 5

Report: Progressive Disclosure of Agent Skills

Guilin Zhang, Kai Zhao, Priyanka Mudgal and 4 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Eager loading of all skill definitions inflates the prompt, causing high token usage, operational cost, context overflow crashes, and degraded skill-retrieval quality as the library grows.

Approach

The method introduces progressive disclosure (lazy-loading) of skills. First, only the frontmatter (name and brief description) of every skill is placed in the LLM prompt. The LLM core is then prompted to emit a structured command indicating the most relevant skill for the current task. The harness loads the full body of that selected skill and adds it to the prompt for a subsequent LLM invocation. This two‑step process repeats as needed, reducing the number of tokens disclosed at each step while still allowing the LLM to access full skill details when required.

Result

Progressive disclosure consistently saved tokens (e.g., 67.9% token reduction for a 20‑skill library with Qwen2.5-7B) and improved skill‑retrieval success rates (up to 100% success for Qwen2.5-7B), while introducing a modest increase in overall latency (a reported increase when using Qwen3-8B).

Why it matters

Deployers of production LLM agents with sizable skill libraries should adopt progressive disclosure to cut token costs and boost retrieval reliability, accepting a slight latency penalty.

Method details
  • LLM cores: Qwen2.5-7B-Instruct, Qwen3-8B, Qwen3-14B with greedy decoding and a k‑size context window.
  • Inference served via vLLM on a single NVIDIA L40S (48 GB) GPU.
  • Four library sizes evaluated (including 20, 50, and 100 skills) with distractor skills of three difficulty tiers.
  • Baselines: eager loading (full skill disclosure) versus progressive disclosure (frontmatter then body on demand).
  • Metrics recorded per rollout: prompt token count, completion token count, wall‑clock latency, and skill‑retrieval outcome categories.
Numbers
  • Qwen2.5-7B eager loading tokens 5 vs progressive disclosure tokens 1648, token savings 1215
  • Qwen2.5-7B eager loading success 26.3% vs progressive disclosure success 100%
  • Library size 20 eager loading tokens 6057 vs progressive disclosure tokens 1943, token savings 67.9%
  • Library size 20 eager loading success 0.17 vs progressive disclosure success 1.00
  • Library size 50 eager loading tokens 14871 vs progressive disclosure tokens 3481, token savings 76.6%
  • Library size 50 eager loading success 0.40 vs progressive disclosure success 0.67
Limitations

The study does not establish behavior for libraries larger than 100 skills nor fully quantify latency trade‑offs for very large prompt budgets.

progressive disclosure improves skill-retrieval quality but marginally degrades overall latency.Found in the source text, word for word.

Picked because: Describes progressive‑disclosure (lazy‑loading) of agent skills with released implementation, directly lowering operational costs of self‑hosted agents.

Paper 5 of 5

PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents

Yangqin Jiang, Lingrui Xu, Chao Huang · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Mobile GUI agents rely on a screenshot‑perception loop with a vision‑language model at each step, which is slow, costly, and brittle; simply running the VLM repeatedly does not solve the inefficiency because navigation dominates the interaction cost.

Approach

PhoneCLI first performs offline compilation by exploring an app, extracting screens, interactive elements, and navigation edges into a semantically annotated map and a catalog of deterministic commands. At runtime, a router maps a task description to a command identifier, verifies the match, and replays the command deterministically. If verification fails or no command matches, control falls back to the embedded VLM interpreter. The system thus separates static navigation (compiled) from dynamic content handling (interpreted).

Result

PhoneCLI raises task success from the VLM‑only baseline (50.7%) to 63.0% while cutting navigation rounds and token usage, and a 50‑screen map budget yields peak success (95.7% on Setting, 77.8% on Clock) that degrades when the map grows to 100 screens.

Why it matters

Developers of mobile GUI agents should care because PhoneCLI offers a way to dramatically reduce VLM inference cost and improve reliability for repetitive navigation tasks without requiring app‑specific APIs or model retraining.

Method details
  • Backbone models used: Qwen3.7-Plus, GLM-4.6V, Kimi-K3.
  • Benchmarks: AndroidLab (138 tasks, 9 apps) and AndroidWorld (116 tasks, 20 apps).
  • Baseline groups: general-purpose vision‑capable LLMs (including Qwen3.7-Plus, Gemini-2.5-Pro, GPT-4o, Claude‑Sonnet 4, Kimi-K3, GLM-4.6V) and GUI‑specialized models (AutoGLM-Phone, AutoGLM-Mobile, MobileUse, UI‑Genie‑Agent, UI‑Tars‑1.5, V‑Droid, AutoGLM‑2024‑10).
  • Ablation results: w/o map (VLM only) 50.7% success, w/o replay 55.1% success, full PhoneCLI 63.0% success.
  • Map construction: breadth‑first crawler up to 50 screens, depth 3, yielding 320 screens, 2,822 elements, 1,089 edges, 2,017 operations; building takes ~10 minutes per app.
  • Efficiency numbers: average rounds drop from 7.52 to 6.95 and tokens from 40.1k to 39.6k with injection; replay further drops rounds to 6.72 and tokens to 34.4k.
Numbers
  • Success rate 63.0% vs VLM‑only 50.7%
  • Success rate 55.1% for w/o replay variant
  • Average rounds 7.52 → 6.95 with injected map information
  • Average tokens 40.1k → 39.6k with injected map information
  • Success on Setting app 95.7% at 50 screens, 78.3% at 100 screens
  • Success on Clock app 77.8% at 50 screens, 66.7% at 100 screens
Limitations

The paper does not state explicit limitations; it only notes that the compiled layer cannot handle dynamic content beyond falling back to the VLM.

PhoneCLI improves the task success rate while reducing steps and token consumptionFound in the source text, word for word.

Picked because: Presents PhoneCLI, an open‑source system that turns mobile app GUIs into callable commands for agents, facilitating practical tool integration.