Organizations lacked a unified ledger for AI inference spend across Kubernetes allocations and LLM provider bills, causing large portions of cost to be unowned; simply labeling pods does not fix it because owner labels on only leader pods leave most GPU spend unattributed.
Approach
The paper introduces unalloc, an open‑source tool that merges cost data from OpenCost, LiteLLM, OpenAI and Anthropic into a single exact ledger and reports unowned spend. It defines metering rules that decide cost attribution inside a shared inference server, comparing token‑based and time‑share meters. The tool is evaluated on five case studies ranging from a vLLM‑style simulator to real PyTorch transformer serving and distributed parallel inference. Results are positioned against recent Shapley‑based energy attribution methods. The approach reveals how attribution breaks at system seams and how fallback keys can reassign cost.
Result
When only LeaderWorkerSet leader pods are labeled, 66% of the deployment's GPU bill is unowned; applying a natural fallback key assigns 61% to a Helm chart name and reduces the headline unallocated share to 4%. On an NVIDIA H100, a token meter gives a retrieval‑heavy tenant 12‑14 percentage points more of the bill than an equal time‑share meter at every load, while GPU utilization stays at 97‑99% across loads of 2 to 16 requests per second. Reading a single page of a billing API reports only a quarter of the total spend.
Why it matters
Operators of shared inference clusters and organizations paying for LLM API usage need accurate cross‑system cost attribution to allocate spend correctly and avoid large unowned expenses.
Method details
unalloc integrates OpenCost, LiteLLM, OpenAI and Anthropic cost streams
Case study 1 uses a vLLM‑style serving simulator with paged KV memory and prefix caching
Case study 2 serves a multi‑tenant trace with a real KV cache on PyTorch transformer
Case study 3 runs tensor‑ and pipeline‑parallel inference on torch.distributed
Case study 4 uses the unmodified CLI against mock provider APIs
Four downstream use cases are also evaluated
Numbers
66% unowned GPU bill when only leader pods labeled
61% assigned to Helm chart name by fallback key
4% headline unallocated share after fallback key
12-14 percentage points more cost for retrieval‑heavy tenant with token meter
97-99% GPU utilization across loads of 2 to 16 rps
one quarter of spend reported by reading one page of billing API
Limitations
The meters used are not ground truth and the study relies on synthetic OpenCost allocations rather than observed billing data, so exact true attribution is not established.
owner labels set only on LeaderWorkerSet leader pods leave 66% of that deployment's GPU bill unownedFound in the source text, word for word.
Picked because: Introduces unalloc, an open-source tool that integrates Kubernetes cost data with LLM provider billing to attribute inference spend, giving engineers actionable cost-visibility for self-hosted deployments.
Haoran Ye, Yuxing Lu, Haonan Dong and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Agent harnesses boost performance but their gains are tied to the harness at deployment, forcing a model to use a suboptimal shared harness or manage many specialized ones, and the obvious fix of simply attaching the best harness fails because the harnesses differ in action space and information, preventing direct supervision.
Approach
Harness-Zero introduces an agent-as-harness that, guided by an optimized harness, intervenes on student responses to translate them into the target harness's action space, creating compatible training demonstrations. These corrected trajectories are then used to fine‑tune the model via LoRA supervised fine‑tuning, internalizing the specialized harness behavior while keeping a fixed minimal target harness at deployment.
Result
After fine‑tuning, Harness-Zero raises macro‑average task success from 23.3% to 44.3%, surpassing the 41.7% achieved when the specialized harness remains attached, and agent-as-harness outperforms code-as-harness with an average of 81.1% versus 78.1%. It also recovers harness‑induced behaviors at an average rate of 82.3% across 28 patterns.
Why it matters
Researchers building agent systems can internalize domain‑specific harness benefits into model weights, reducing dependence on external scaffolds and simplifying deployment.
Method details
Base model Qwen3.5-9B fine‑tuned with LoRA rank 32 for two epochs using the Tinker recipe.
Baselines: code-as-harness, base model without distillation, base model with the specialized harness attached (41.7% success).
Ablations show harness‑guided review is the most effective supervision source and that weaker harnessing agents can make review harmful.
Numbers
macro‑average task success, 23.3%, base model without distillation
macro‑average task success, 44.3%, after Harness‑Zero SFT
macro‑average task success, 41.7%, with specialized harness still attached
agent‑as‑harness vs code‑as‑harness, 81.1% vs 78.1% average success
behavior recovery rate, 82.3%, average across 28 patterns
training rollouts, 487 SpreadsheetBench, 282 AppWorld, 500 USPTO
Limitations
The method relies on a capable harnessing model; weaker models make review harmful and increase collection cost, and it does not fully eliminate the need for a harness.
Harness-Zero improves the base model's macro-average task success from 23.3% to 44.3%, even exceeding the 41.7% it reaches with that harness still attached.Found in the source text, word for word.
Picked because: Presents Harness-Zero, a method to distill and compress LLM agent harnesses into lightweight, deployable components, enabling practical tooling and verification of agent pipelines.
Mahmoud Ayyad, Zehao Wang, Jiho Shin and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Evaluating a full regression benchmark for SWE‑agents costs hundreds of millions of LLM tokens, and simple subset selection methods such as random or stratified sampling produce high variance and unrepresentative subsets.
Approach
The method parses raw agent trajectories, sanitizes them to remove explicit outcome signals, and feeds the cleaned trajectories into an embedding model to obtain numerical representations. Instances are first grouped by their recent test outcome to preserve pass/fail rates, then for each outcome group the instance whose embedding is closest to the group centroid is selected. The selected deterministic subset is evaluated on later runs and its resolve rate is used to estimate the full suite resolve rate. This trajectory‑aware selection replaces stochastic sampling with a deterministic, embedding‑driven choice.
Result
The centroid‑based trajectory‑aware method achieved the lowest estimation error, reducing average error by 3--11% and worst‑case error by 4--11% relative to a typical random draw, and by 38--46% relative to the 95th‑percentile draw of the strongest baseline. A 10% subset kept median estimation error below 5% while cutting token cost by roughly 90%. Consist‑Strat baseline median RMSE at 5% subset was 0.0367 versus Random 0.0744.
Why it matters
Researchers and engineers building SWE‑agents should adopt trajectory‑aware subset selection to dramatically lower evaluation cost while maintaining accurate regression estimates.
Method details
Embedding model generates numerical representations from sanitized trajectories.
Outcome stratification groups test instances by recent pass/fail outcomes before selection.
Centroid Pooled selection picks instances nearest to the centroid of each outcome group in embedding space.
Evaluated 76 subset selection configurations including random, embedding‑based, clustering‑based, and hybrid methods.
Baselines include Random, Diff‑Strat, Consist‑Strat, Repo‑Strat, RepoDiff, and RepoConsist.
Three regression scenarios: same‑configuration reruns, model and configuration changes, and agent framework changes.
Numbers
average estimation error reduction 3--11% vs typical draw
worst‑case error reduction 4--11% vs typical draw
worst‑case error reduction 38--46% vs 95th‑percentile draw of strongest baseline
10% subset median estimation error <5% with ~90% token cost reduction
Consist‑Strat median RMSE 0.0367 at 5% subset vs Random 0.0744
Centroid Pooled RMSE 0.0679 (8.4% relative change) at 5% Multi‑Model
Limitations
The paper does not claim applicability beyond the evaluated regression scenarios and does not provide results for other agent domains.
Our approach reduces the average estimation error by 3--11% and the worst-case error by 4--11% relative to the typical draw and 38--46% relative to the 95th-percentile draw of the strongest baseline.Found in the source text, word for word.
Picked because: Shows how to select cost-effective subsets of regression-testing benchmarks for software-engineering agents, reducing token-usage by orders of magnitude while preserving detection power.
Zhilin Wang, Shaokun Zhang, Yifan Zhang and 9 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Prior evaluation of Computer-Use Agents focused only on final deliverables using functional verifiers, which hides how and why agents fail. Simply adding more end‑state tests does not reveal process errors such as keyboard versus click mistakes.
Approach
The authors introduce OSWorld‑Pro, a benchmark of over 300 tasks broken into more than 2800 subgoals with 67,000 human annotations. They employ human‑aligned LLM‑Judges that score each step on subgoal targeting, feasibility, progression, and completion using Macro‑F1. The judges evaluate an entire trajectory in a single API request, allowing partial credit for early subgoal success. Performance is aggregated at step, subgoal, and task levels, and compared against human labels. This process‑based evaluation exposes failure modes that outcome‑based metrics miss.
Result
LLM‑Judges closely match human judgments (94.1% step‑level agreement) and enable measurement of model performance on OSWorld‑Pro. Claude Opus 5 Max attains only 75.7% task completion, lower than its 83.4% score on OSWorld. Open‑weight models lag further, with the best open‑weight (Qwen 3.8 Flash Next Xhigh) reaching 55.1% overall.
Why it matters
Researchers developing computer‑use agents and evaluation frameworks should adopt OSWorld‑Pro to gain fine‑grained insight into process failures and improve agent robustness.
Method details
OSWorld‑Pro contains >300 tasks, >2800 subgoals, and >67,000 human annotations.
LLM‑Judges are evaluated with Macro‑F1 against human labels, achieving 94.1% agreement on step‑level fields.
Inference uses a single API call with up to 500 MB of screenshots; only OpenAI GPT‑5.6 models handled the payload.
Baseline comparison includes the original OSWorld benchmark where top Opus model scores 83.4%.
Human annotator agreement measured by Cohen's kappa is 0.8 overall.
Numbers
Claude Opus 5 Max task completion 75.7% vs OSWorld 83.4%
The paper does not state any explicit limitations.
OSWorld‑Pro is challenging even for state-of-the-art LLMs, with top performers like Claude Opus 5 achieving only 75.7% vs. 83.4% on OSWorld.Found in the source text, word for word.
Picked because: Provides OSWorld-Pro, a process-level evaluation framework for computer-use agents that surfaces failure modes during execution, offering concrete diagnostics for agent verification.
Gabriele Tombesi, William Baisi, Je Yang and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AR
Problem
Prior edge LLM inference struggled with speculative decoding because verification creates a runtime‑dependent intermediate regime between memory‑bound GEMV and compute‑bound GEMM operations, and static accelerator designs cannot efficiently handle both regimes.
Approach
SPECTRA is a runtime‑reconfigurable tiled architecture where each accelerator tile contains a compute engine that can switch between systolic execution for GEMMs and vector‑lane execution for GEMVs. Tiles are interconnected via a network‑on‑chip and a host processor configures both tile‑level and system‑level parameters per kernel. The system dynamically selects the number of active tiles, kernel partitioning, and communication patterns to match the current phase of speculative decoding. Reconfigurability operates at kernel granularity, allowing the fabric to adapt to the varying arithmetic intensity of draft decode, prefill, and verification phases.
Result
SPECTRA achieves up to 2.09× speedup from tile‑level reconfiguration and an additional 1.25× gain from system‑level adaptability over fixed designs, and delivers up to 3.6× higher throughput and 2.9× higher power efficiency compared to edge GPU platforms.
Why it matters
Edge device designers and FPGA accelerator developers should care because SPECTRA shows how runtime reconfigurability can substantially improve LLM inference efficiency on resource‑constrained hardware.
Method details
Pythia‑70M/160M, SmolLM2‑135M/360M, GPT2‑124M/774M model pairs were used as benchmarks
Prototype implemented on a 20‑tile FPGA (14 accelerator, 4 memory, 1 processor, 1 I/O) on a UltraScale+ XCVU19P
Baselines compared were a vector‑only datapath, a systolic‑only datapath, and fixed static designs
Ablations included computation‑fabric flexibility, system‑level sharding flexibility, speculation‑length sweep, output‑length sweep, and prompt‑length sweep
Resource utilization of the reconfigurable design was 88,206 LUTs, 103,564 FFs, 112 BRAM, 18 URAM, 308 DSPs
Numbers
speedup, 2.09x, over fixed designs (tile‑level reconfiguration)
speedup, 1.25x, over fixed designs (system‑level adaptability)
throughput, 3.6x, higher than edge GPU platforms
power efficiency, 2.9x, higher than edge GPU platforms
Limitations
The paper does not evaluate models larger than GPT‑2 774M or assess performance on non‑FPGA or ASIC platforms, and it focuses only on inference, not training.
SPECTRA achieves up to 2.09 speedup from tile-level reconfiguration and a further 1.25 gain from system-level adaptability over fixed designs.Found in the source text, word for word.
Picked because: Describes SPECTRA, a runtime-reconfigurable tiled architecture that implements speculative decoding with adaptive execution, delivering a released prototype that speeds up LLM inference on edge hardware.