arXiv digest

Monday

September 7, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Beyond Scalar Flexibility: From Eligible AI Workloads to Dependable Load Relief

Meiyi Li · abstract · pdf

quote verifiedfigures checkedread: full texteess.SY

Problem

Prior grid studies model data‑center flexibility as a fixed percentage of load, but no public production trace showed how much eligible load persists across event durations or co‑moves across clusters, so the scalar assumption lacks empirical grounding.

Approach

The authors reconstruct hourly power from a 185‑day trace of 155,410 GPUs, building workload‑conditioned GPU power curves and a four‑layer semantic flexibility envelope. They compute immediate eligible curtailment while retaining allocated‑GPU idle power, then evaluate 95%‑available relief for multiple durations using Monte Carlo medians. Scalars are calibrated against the envelope (mean‑calibrated and tail‑calibrated) to assess over‑/under‑statement. Aggregation across 13 clusters is analyzed to measure firmness gains, and the production scheduler’s delay‑based capacity is quantified.

Result

The study finds that 95%‑available relief declines from 2.51 MW for a one‑hour event to 1.95 MW for a 24‑hour event under full eligibility, and that a mean‑calibrated scalar overstates these values by 17% to 47% while a tail‑calibrated scalar understates the one‑hour product by 6% and overstates the 24‑hour product by 17%; aggregating clusters improves firmness but cross‑cluster covariance limits gains.

Why it matters

Grid planners and demand‑response operators should use the duration‑reliability‑portfolio surface instead of a single flexibility percentage to size reliable load‑relief contracts for AI data‑centers.

Method details
  • 4,439 hourly power observations reconstructed from a 185‑day trace of 155,410 GPUs
  • 300 Monte Carlo model draws expose parameter sensitivity
  • Immediate eligible curtailment averages 3.55 MW (12.1% of workload power, 6.35% of median facility power)
  • 95%‑available relief: 2.51 MW for one hour, 2.32 MW for four hours, 1.95 MW for 24 hours
  • Aggregating 13 clusters raises four‑hour firmness from 0.38 to 0.66
  • Newly deferrable arrivals average 0.008 MW and have zero 95%‑available capacity
Numbers
  • 55.8 MW fleet time‑averaged Monte Carlo median facility demand
  • 3.55 MW immediate eligible curtailment average
  • 12.1% of workload power eligible
  • 6.35% of median facility power eligible
  • 2.51 MW 95%‑available relief for one hour
  • 2.32 MW 95%‑available relief for four hours
  • 1.95 MW 95%‑available relief for 24 hours
  • 0.38 four‑hour firmness for single cluster
  • 0.66 four‑hour firmness after aggregating 13 clusters
  • 0.008 MW newly deferrable arrivals average
Limitations

The paper only measures historical eligibility, not actual delivered response, and cannot assess sub‑hour dynamics, checkpointing overhead, or control latency.

The production scheduler exposes almost no additional delay-based capacity: newly deferrable arrivals average 0.008 MW and have zero 95%-available capacity.Found in the source text, word for word.

Picked because: Provides a concrete workload‑semantic flexibility analysis of a large GPU fleet with released trace reconstruction and metrics useful for data‑center DevOps automation.

Paper 2 of 5

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

Dain Kim, Eungi Cho, Kyumin Kim and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Open-source LLM agents underperform on multi-step tool‑calling with live Korean public APIs, and existing benchmarks rely on emulated or hand‑authored simulations that do not capture real execution failures.

Approach

The paper introduces EDGE, a two‑phase framework. Phase A builds an execution‑grounded dynamic graph by proposing candidate edges with an LLM, then pruning edges that fail when actually called against live APIs. Phase B traverses the verified graph to assemble multi‑step trajectories, types each junction by response cardinality, and generates Korean queries and answers for training. The resulting dataset is used to fine‑tune models with the GRPO objective, yielding improved multi‑step tool‑calling performance.

Result

On KOPA‑Bench the GRPO‑fine‑tuned 4B model achieves pass@1 0.3094 and pass@4 0.4690, surpassing the base 4B model (pass@1 0.1758, pass@4 0.3103) and also improving the Action score to 0.3462 from 0.2140. On the out‑of‑distribution BFCL benchmark, multi‑turn performance improves by +4.04pp for the 4B model and +5.87pp for the 9B model.

Why it matters

Organizations that must run on‑premise LLM agents for Korean government services can adopt EDGE to obtain near‑state‑of‑the‑art multi‑step tool‑calling performance without large proprietary models, and researchers gain a realistic benchmark for future tool‑calling work.

Method details
  • Model sizes: Qwen3.5‑4B base and fine‑tuned (GRPO) and Qwen3.5‑9B fine‑tuned; also Qwen3.5‑27B used as untuned reference.
  • Architecture: decoder‑only transformer from the Qwen3.5 family.
  • Datasets: KOPA‑Bench with 145 real‑world tasks across 10 platforms and six domains; EDGE‑synthesized dataset filtered for execution success.
  • Training objective: compare standard supervised fine‑tuning (SFT) with GRPO on the same EDGE tasks.
  • Baselines: Qwen3.5‑27B untuned, Qwen3.5‑4B base, gemma‑4‑26B‑A4B‑it, EXAONE‑4.5‑33B, and other open‑source models.
  • Ablations: training objective (SFT vs GRPO), trajectory composition (pure‑sequential vs parallel+Mixed vs full), data filtering impact, and graph pruning effectiveness.
Numbers
  • pass@1 0.3094 vs base 0.1758 (Qwen3.5‑4B)
  • pass@4 0.4690 vs base 0.3103 (Qwen3.5‑4B)
  • Action 0.3462 vs base 0.2140 (Qwen3.5‑4B)
  • BFCL multi‑turn improvement +4.04pp for 4B model
  • BFCL multi‑turn improvement +5.87pp for 9B model
  • KOPA‑Bench contains 145 tasks across 10 platforms and six domains
Limitations

The study is limited to Korean public APIs and does not evaluate generalization to other languages or non‑governmental services.

Fine-tuned via GRPO on the resulting dataset, our 9B model nearly matches the untuned 27B model from the same familyFound in the source text, word for word.

Picked because: Introduces KOPA‑Bench and a data‑synthesis recipe for multi‑step tool‑calling with open APIs, giving engineers a ready‑to‑use benchmark for LLM agent tooling.

Paper 3 of 5

CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents

Haoting Shi, Wenhao Wang, Weicheng Fang and 6 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Prior computer-use agents act mostly through the GUI, leading to inefficient trajectories, and existing hybrid environments require extensive manual engineering per application, while CLI‑native agents lack visual perception and GUI‑native agents are inefficient for command‑driven operations.

Approach

The method builds CUA-Universe, a pipeline that (1) adapts real desktop applications into reproducible VMs with discovered or generated CLI surfaces (App-Forge), (2) synthesizes compositional hybrid tasks of controllable difficulty from reusable operations (Task-Weave), and (3) steers rollouts along efficient hybrid paths to harvest verified trajectories (Path-Steer). These components generate a large set of hybrid GUI+CLI episodes, which are used to fine‑tune a 9B language model with LoRA. The resulting model learns to coordinate GUI navigation and CLI commands over shared application state.

Result

The 9B model improves success and efficiency on hybrid benchmarks, achieving a CUA‑Verse score increase of +39.3 points with 37% fewer steps and 60% fewer tokens, an OSWorld success‑rate gain of +16.8 points with 57% fewer steps and 44% fewer tokens, and an OSWorld‑MCP score gain of +7.84 points with 27% fewer steps and 30% fewer tokens.

Why it matters

Researchers and developers of computer‑use agents should care because the pipeline demonstrates a scalable way to generate hybrid GUI+CLI supervision that yields measurable efficiency gains without increasing model size.

Method details
  • Base model: Qwen3.5-9B fine‑tuned with LoRA (rank 8, dropout 0.05).
  • Training data: 4,923 verified episodes (235,408 step‑level records).
  • Fine‑tuning: 3 epochs on A100 GPUs, roughly two days, using ms‑swift framework.
  • Generated environments: 16 applications with up to 46 CLI commands each.
  • Baselines compared: Kimi K2.5, Seed2.1 Pro, GPT-5.5, Qwen3.5-9B, EvoCUA-8B.
  • Ablation: out‑of‑domain 8‑app LoRA vs. full 16‑app LoRA on OSWorld.
Numbers
  • CUA‑Verse Score +39.3 pts vs baseline
  • CUA‑Verse steps -37% vs baseline
  • CUA‑Verse tokens -60% vs baseline
  • OSWorld SR +16.8 pts vs baseline
  • OSWorld steps -57% vs baseline
  • OSWorld tokens -44% vs baseline
Limitations

The paper does not state explicit limitations.

Our 9B model improves both success and efficiency on CUA-Verse (Score +39.3 pts; -37% steps, -60% tokens), OSWorld (SR +16.8 pts; -57% steps, -44% tokens), and OSWorld-MCP (Score +7.84 pts; -27% steps, -30% tokens).Found in the source text, word for word.

Picked because: Releases CUA‑Universe, a scalable hybrid GUI+CLI environment enabling self‑hosted agents to coordinate both modalities in real‑world software tasks.

Paper 4 of 5

Testing Interchangeability in LLM Agent Teams

Jianxin Gao, Tianyi Yu, Linna Deng and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing multi-agent deployments assume agents are freely interchangeable, but swapping agents often degrades coordination. The obvious fix of treating a swap as a simple competence replacement fails because partner-specific conventions cause hidden communication overhead.

Approach

The paper introduces a swap test where multiple independent teams are formed from a single base model, each agent keeping a private notebook over ten formation episodes. Role‑matched agents are then exchanged between teams while a placebo condition reproduces the roster change without swapping. Metrics such as task score, communication per unit of progress, and protocol signature divergence are recorded. By comparing swap, placebo, and naive (inexperienced) conditions, the method isolates the cost of partner‑specific conventions. Ablations over model, decoding temperature, and formation length further probe the mechanisms.

Result

Swapping agents leaves task score almost unchanged but raises communication per unit of progress by 16 to 63 percent. Greedy decoding halves the swap penalty relative to higher temperature sampling, while longer formation histories increase both swap cost and team divergence. Partner‑specific effects are modest but measurable, especially in high‑coupling settings like Collab-Overcooked.

Why it matters

Operators of production LLM agent systems should care because agent rotation impacts coordination efficiency more than raw performance, suggesting the need for monitoring communication overhead and possibly clearing partner notes.

Method details
  • Base models: GPT-5.6 Luna, Gemini 3.7 Flash, Claude Sonnet 5.
  • Tasks: Hanabi and Collab-Overcooked.
  • Teams: eight per setting; six per configuration in ablations.
  • Formation episodes: 5, 10 (default), and 20.
  • Decoding temperatures tested: 0.0 (greedy), 0.7, 1.0, 1.5, 2.0.
  • Ablations: model variation, temperature variation, and formation‑length variation.
Numbers
  • communication increase 16 to 63 percent compared to placebo
  • GPT-5.6 Luna task score 63.7
  • Gemini 3.7 Flash task score 54.9
  • Claude Sonnet 5 task score 73.3
  • temperature 0.0 yields task score 66.2
  • formation 5 episodes task score 60.6; 10 episodes 63.7; 20 episodes 68.6
Limitations

The study is limited to dyadic teams, short formation horizons (up to twenty episodes), and relies on a notebook design that may exaggerate partner‑specific effects.

a swap costs little in task score but raises the communication a team spends per unit of progress by 16 to 63 percentFound in the source text, word for word.

Picked because: Empirically evaluates interchangeability of LLM agents across teams, offering practical verification insights for deploying interchangeable agent components.

Paper 5 of 5

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

Konstantin Grotov, Valentin Malykh · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

LLM agents for software engineering act confidently wrong and only detect failures after costly execution, while closed‑API agents expose no logits and sampling‑based uncertainty estimators are too expensive.

Approach

Speculative Uncertainty (SU) inverts speculative decoding by feeding the agent’s completed token trajectory into a small open‑weight draft model to obtain cross‑likelihoods, separates reasoning and action spans to form phase‑aware features, fits a linear calibrator on these features, and applies a pre‑execution veto gate that blocks actions predicted to fail.

Result

Using SU with a veto gate reduced the per‑call execution error rate by 6 to 8 percentage points and lowered token cost by 14 to 19% across both Qwen3‑Coder‑480B and Claude 3.5 Sonnet, with zero‑shot transfer to out‑of‑distribution benchmarks.

Why it matters

Teams deploying closed‑API coding agents can add a cheap draft‑model forward pass and veto gate to substantially cut errors and compute cost without modifying the agent itself.

Method details
  • Agent models: Qwen3-Coder-480B and Claude 3.5 Sonnet
  • Draft model: Qwen3-4B trained with SFT or teacher‑forced distillation
  • Training data: SWE‑rebench OpenHands trajectories and real‑world GitHub issues, plus OOD splits SWE‑Bench Verified and DA‑Code
  • Baselines: Verbalized Confidence, Last‑TP, Global‑TP, HTC, and SU‑Uniform ablation
  • Hyperparameters: learning rate 1e-5 (cosine, 5% warmup), epochs 2, batch size 32 sequences, max seq length 32768 tokens, weight decay 0.1, gradient clipping 1.0, precision bf16
Numbers
  • error rate reduction, 6 to 8 percentage points, compared to baseline without veto
  • token cost reduction, 14 to 19 %, compared to baseline without veto
  • SWE‑Bench Verified task‑success change, 5 percentage points, compared to baseline
  • learning rate, 1e-5, used for draft model training
  • epochs, 2, used for draft model training
Limitations

The study only measures code‑execution success as the objective, reports single‑run point estimates without variance, and assumes the draft model’s overhead is negligible for this workload.

cutting execution error rate by 6 to 8 percentage points and token cost by 14 to 19% in deploymentFound in the source text, word for word.

Picked because: Presents the Speculative Uncertainty method to derive failure signals from black‑box LLM agents, a directly applicable technique for robust agent deployment.