arXiv digest

Tuesday

September 22, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Who Pays for the KV Cache? Attributing Shared AI Inference Spend Across Kubernetes and LLM Provider Bills

Timothy Urista · abstract · pdf

quote verifiedfigures checkedread: abstract onlycs.DC

Problem

Organizations lacked a unified ledger for AI inference spend across Kubernetes allocations and LLM provider bills, causing large portions of cost to be unowned; simply labeling pods does not fix it because owner labels on only leader pods leave most GPU spend unattributed.

Approach

The paper introduces unalloc, an open‑source tool that merges cost data from OpenCost, LiteLLM, OpenAI and Anthropic into a single exact ledger and reports unowned spend. It defines metering rules that decide cost attribution inside a shared inference server, comparing token‑based and time‑share meters. The tool is evaluated on five case studies ranging from a vLLM‑style simulator to real PyTorch transformer serving and distributed parallel inference. Results are positioned against recent Shapley‑based energy attribution methods. The approach reveals how attribution breaks at system seams and how fallback keys can reassign cost.

Result

When only LeaderWorkerSet leader pods are labeled, 66% of the deployment's GPU bill is unowned; applying a natural fallback key assigns 61% to a Helm chart name and reduces the headline unallocated share to 4%. On an NVIDIA H100, a token meter gives a retrieval‑heavy tenant 12‑14 percentage points more of the bill than an equal time‑share meter at every load, while GPU utilization stays at 97‑99% across loads of 2 to 16 requests per second. Reading a single page of a billing API reports only a quarter of the total spend.

Why it matters

Operators of shared inference clusters and organizations paying for LLM API usage need accurate cross‑system cost attribution to allocate spend correctly and avoid large unowned expenses.

Method details
  • unalloc integrates OpenCost, LiteLLM, OpenAI and Anthropic cost streams
  • Case study 1 uses a vLLM‑style serving simulator with paged KV memory and prefix caching
  • Case study 2 serves a multi‑tenant trace with a real KV cache on PyTorch transformer
  • Case study 3 runs tensor‑ and pipeline‑parallel inference on torch.distributed
  • Case study 4 uses the unmodified CLI against mock provider APIs
  • Four downstream use cases are also evaluated
Numbers
  • 66% unowned GPU bill when only leader pods labeled
  • 61% assigned to Helm chart name by fallback key
  • 4% headline unallocated share after fallback key
  • 12-14 percentage points more cost for retrieval‑heavy tenant with token meter
  • 97-99% GPU utilization across loads of 2 to 16 rps
  • one quarter of spend reported by reading one page of billing API
Limitations

The meters used are not ground truth and the study relies on synthetic OpenCost allocations rather than observed billing data, so exact true attribution is not established.

owner labels set only on LeaderWorkerSet leader pods leave 66% of that deployment's GPU bill unownedFound in the source text, word for word.

Picked because: Introduces unalloc, an open-source tool that integrates Kubernetes cost data with LLM provider billing to attribute inference spend, giving engineers actionable cost-visibility for self-hosted deployments.

Paper 2 of 5

Harness-Zero: Harness Distillation via Agent-as-Harness

Haoran Ye, Yuxing Lu, Haonan Dong and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Agent harnesses boost performance but their gains are tied to the harness at deployment, forcing a model to use a suboptimal shared harness or manage many specialized ones, and the obvious fix of simply attaching the best harness fails because the harnesses differ in action space and information, preventing direct supervision.

Approach

Harness-Zero introduces an agent-as-harness that, guided by an optimized harness, intervenes on student responses to translate them into the target harness's action space, creating compatible training demonstrations. These corrected trajectories are then used to fine‑tune the model via LoRA supervised fine‑tuning, internalizing the specialized harness behavior while keeping a fixed minimal target harness at deployment.

Result

After fine‑tuning, Harness-Zero raises macro‑average task success from 23.3% to 44.3%, surpassing the 41.7% achieved when the specialized harness remains attached, and agent-as-harness outperforms code-as-harness with an average of 81.1% versus 78.1%. It also recovers harness‑induced behaviors at an average rate of 82.3% across 28 patterns.

Why it matters

Researchers building agent systems can internalize domain‑specific harness benefits into model weights, reducing dependence on external scaffolds and simplifying deployment.

Method details
  • Base model Qwen3.5-9B fine‑tuned with LoRA rank 32 for two epochs using the Tinker recipe.
  • Training data: 487 SpreadsheetBench rollouts, 282 AppWorld rollouts, 500 USPTO rollouts.
  • Datasets: SpreadsheetBench Verified (300 train, 100 eval), AppWorld (147 train, 168 test_normal), USPTO Retrosynthesis (500 train, 100 test).
  • Baselines: code-as-harness, base model without distillation, base model with the specialized harness attached (41.7% success).
  • Ablations show harness‑guided review is the most effective supervision source and that weaker harnessing agents can make review harmful.
Numbers
  • macro‑average task success, 23.3%, base model without distillation
  • macro‑average task success, 44.3%, after Harness‑Zero SFT
  • macro‑average task success, 41.7%, with specialized harness still attached
  • agent‑as‑harness vs code‑as‑harness, 81.1% vs 78.1% average success
  • behavior recovery rate, 82.3%, average across 28 patterns
  • training rollouts, 487 SpreadsheetBench, 282 AppWorld, 500 USPTO
Limitations

The method relies on a capable harnessing model; weaker models make review harmful and increase collection cost, and it does not fully eliminate the need for a harness.

Harness-Zero improves the base model's macro-average task success from 23.3% to 44.3%, even exceeding the 41.7% it reaches with that harness still attached.Found in the source text, word for word.

Picked because: Presents Harness-Zero, a method to distill and compress LLM agent harnesses into lightweight, deployable components, enabling practical tooling and verification of agent pipelines.

Paper 3 of 5

Trajectory-Aware Benchmark Subset Selection for Cost-Efficient Software Engineering Agent Regression Testing

Mahmoud Ayyad, Zehao Wang, Jiho Shin and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Evaluating a full regression benchmark for SWE‑agents costs hundreds of millions of LLM tokens, and simple subset selection methods such as random or stratified sampling produce high variance and unrepresentative subsets.

Approach

The method parses raw agent trajectories, sanitizes them to remove explicit outcome signals, and feeds the cleaned trajectories into an embedding model to obtain numerical representations. Instances are first grouped by their recent test outcome to preserve pass/fail rates, then for each outcome group the instance whose embedding is closest to the group centroid is selected. The selected deterministic subset is evaluated on later runs and its resolve rate is used to estimate the full suite resolve rate. This trajectory‑aware selection replaces stochastic sampling with a deterministic, embedding‑driven choice.

Result

The centroid‑based trajectory‑aware method achieved the lowest estimation error, reducing average error by 3--11% and worst‑case error by 4--11% relative to a typical random draw, and by 38--46% relative to the 95th‑percentile draw of the strongest baseline. A 10% subset kept median estimation error below 5% while cutting token cost by roughly 90%. Consist‑Strat baseline median RMSE at 5% subset was 0.0367 versus Random 0.0744.

Why it matters

Researchers and engineers building SWE‑agents should adopt trajectory‑aware subset selection to dramatically lower evaluation cost while maintaining accurate regression estimates.

Method details
  • Embedding model generates numerical representations from sanitized trajectories.
  • Outcome stratification groups test instances by recent pass/fail outcomes before selection.
  • Centroid Pooled selection picks instances nearest to the centroid of each outcome group in embedding space.
  • Evaluated 76 subset selection configurations including random, embedding‑based, clustering‑based, and hybrid methods.
  • Baselines include Random, Diff‑Strat, Consist‑Strat, Repo‑Strat, RepoDiff, and RepoConsist.
  • Three regression scenarios: same‑configuration reruns, model and configuration changes, and agent framework changes.
Numbers
  • average estimation error reduction 3--11% vs typical draw
  • worst‑case error reduction 4--11% vs typical draw
  • worst‑case error reduction 38--46% vs 95th‑percentile draw of strongest baseline
  • 10% subset median estimation error <5% with ~90% token cost reduction
  • Consist‑Strat median RMSE 0.0367 at 5% subset vs Random 0.0744
  • Centroid Pooled RMSE 0.0679 (8.4% relative change) at 5% Multi‑Model
Limitations

The paper does not claim applicability beyond the evaluated regression scenarios and does not provide results for other agent domains.

Our approach reduces the average estimation error by 3--11% and the worst-case error by 4--11% relative to the typical draw and 38--46% relative to the 95th-percentile draw of the strongest baseline.Found in the source text, word for word.

Picked because: Shows how to select cost-effective subsets of regression-testing benchmarks for software-engineering agents, reducing token-usage by orders of magnitude while preserving detection power.

Paper 4 of 5

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

Zhilin Wang, Shaokun Zhang, Yifan Zhang and 9 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Prior evaluation of Computer-Use Agents focused only on final deliverables using functional verifiers, which hides how and why agents fail. Simply adding more end‑state tests does not reveal process errors such as keyboard versus click mistakes.

Approach

The authors introduce OSWorld‑Pro, a benchmark of over 300 tasks broken into more than 2800 subgoals with 67,000 human annotations. They employ human‑aligned LLM‑Judges that score each step on subgoal targeting, feasibility, progression, and completion using Macro‑F1. The judges evaluate an entire trajectory in a single API request, allowing partial credit for early subgoal success. Performance is aggregated at step, subgoal, and task levels, and compared against human labels. This process‑based evaluation exposes failure modes that outcome‑based metrics miss.

Result

LLM‑Judges closely match human judgments (94.1% step‑level agreement) and enable measurement of model performance on OSWorld‑Pro. Claude Opus 5 Max attains only 75.7% task completion, lower than its 83.4% score on OSWorld. Open‑weight models lag further, with the best open‑weight (Qwen 3.8 Flash Next Xhigh) reaching 55.1% overall.

Why it matters

Researchers developing computer‑use agents and evaluation frameworks should adopt OSWorld‑Pro to gain fine‑grained insight into process failures and improve agent robustness.

Method details
  • OSWorld‑Pro contains >300 tasks, >2800 subgoals, and >67,000 human annotations.
  • LLM‑Judges are evaluated with Macro‑F1 against human labels, achieving 94.1% agreement on step‑level fields.
  • Inference uses a single API call with up to 500 MB of screenshots; only OpenAI GPT‑5.6 models handled the payload.
  • Baseline comparison includes the original OSWorld benchmark where top Opus model scores 83.4%.
  • Human annotator agreement measured by Cohen's kappa is 0.8 overall.
Numbers
  • Claude Opus 5 Max task completion 75.7% vs OSWorld 83.4%
  • Human step‑level Macro‑F1 98.4% (Target), 91.6% (Feasible), 94.1% (Progress), 97.6% (Complete)
  • LLM‑Judge GPT‑5.6‑Sol Max Macro‑F1 97.0% (Target), 61.9% (Feasible), 74.7% (Progress), 94.3% (Complete)
  • Claude Opus 4.8 Max overall performance 77.7%
  • Qwen 3.8 Flash Next Xhigh overall 55.1%
  • Minimax M3 Xhigh overall 28.9%
Limitations

The paper does not state any explicit limitations.

OSWorld‑Pro is challenging even for state-of-the-art LLMs, with top performers like Claude Opus 5 achieving only 75.7% vs. 83.4% on OSWorld.Found in the source text, word for word.

Picked because: Provides OSWorld-Pro, a process-level evaluation framework for computer-use agents that surfaces failure modes during execution, offering concrete diagnostics for agent verification.

Paper 5 of 5

SPECTRA: Adaptive Execution of Speculative Decoding on a Runtime-Reconfigurable Tiled Architecture

Gabriele Tombesi, William Baisi, Je Yang and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AR

Problem

Prior edge LLM inference struggled with speculative decoding because verification creates a runtime‑dependent intermediate regime between memory‑bound GEMV and compute‑bound GEMM operations, and static accelerator designs cannot efficiently handle both regimes.

Approach

SPECTRA is a runtime‑reconfigurable tiled architecture where each accelerator tile contains a compute engine that can switch between systolic execution for GEMMs and vector‑lane execution for GEMVs. Tiles are interconnected via a network‑on‑chip and a host processor configures both tile‑level and system‑level parameters per kernel. The system dynamically selects the number of active tiles, kernel partitioning, and communication patterns to match the current phase of speculative decoding. Reconfigurability operates at kernel granularity, allowing the fabric to adapt to the varying arithmetic intensity of draft decode, prefill, and verification phases.

Result

SPECTRA achieves up to 2.09× speedup from tile‑level reconfiguration and an additional 1.25× gain from system‑level adaptability over fixed designs, and delivers up to 3.6× higher throughput and 2.9× higher power efficiency compared to edge GPU platforms.

Why it matters

Edge device designers and FPGA accelerator developers should care because SPECTRA shows how runtime reconfigurability can substantially improve LLM inference efficiency on resource‑constrained hardware.

Method details
  • Pythia‑70M/160M, SmolLM2‑135M/360M, GPT2‑124M/774M model pairs were used as benchmarks
  • Prototype implemented on a 20‑tile FPGA (14 accelerator, 4 memory, 1 processor, 1 I/O) on a UltraScale+ XCVU19P
  • Baselines compared were a vector‑only datapath, a systolic‑only datapath, and fixed static designs
  • Ablations included computation‑fabric flexibility, system‑level sharding flexibility, speculation‑length sweep, output‑length sweep, and prompt‑length sweep
  • Resource utilization of the reconfigurable design was 88,206 LUTs, 103,564 FFs, 112 BRAM, 18 URAM, 308 DSPs
Numbers
  • speedup, 2.09x, over fixed designs (tile‑level reconfiguration)
  • speedup, 1.25x, over fixed designs (system‑level adaptability)
  • throughput, 3.6x, higher than edge GPU platforms
  • power efficiency, 2.9x, higher than edge GPU platforms
Limitations

The paper does not evaluate models larger than GPT‑2 774M or assess performance on non‑FPGA or ASIC platforms, and it focuses only on inference, not training.

SPECTRA achieves up to 2.09 speedup from tile-level reconfiguration and a further 1.25 gain from system-level adaptability over fixed designs.Found in the source text, word for word.

Picked because: Describes SPECTRA, a runtime-reconfigurable tiled architecture that implements speculative decoding with adaptive execution, delivering a released prototype that speeds up LLM inference on edge hardware.