arXiv digest

Thursday

August 27, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Slasher: Power Flexibility for Cloud Datacenters

Liuzixuan Lin, Fiodar Kazhamiaka, Alok Gautam Kumbhare and 10 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.DC

Problem

Existing datacenter power‑modulation systems were built per scenario, lacking coordination across racks, servers, and VMs, which caused missed optimization opportunities and conflicting actions; simply adding separate handlers for each scenario does not work because they cannot share telemetry or jointly reason about impact at scale.

Approach

Slasher introduces a hierarchical control stack with a regional orchestrator that assigns power‑shedding goals to local controllers in each data hall. Each local controller runs an offline mode that continuously ingests telemetry and pre‑computes Decision and Impact tables, and an online mode that selects the minimal set of actions to meet the current limit. A Decision Manager tracks inventory and lever catalog, while an Impact Estimator maps power reductions to expected service impact using CVaR‑based ES metrics. The Consolidator arbitrates multiple active scenarios by enforcing the most restrictive limit, and the Action Performer executes the chosen actions via specialized handlers. Fast‑path hardware controllers handle sub‑second actions such as battery discharge or processor throttling.

Result

The impact model distinguishes services by their tolerance to capacity reduction; Service B can tolerate more capacity loss than Service A while keeping expected shortfall negligible, as shown in the ES impact functions of Figure 12.

Why it matters

Datacenter operators and cloud platform engineers should care because Slasher offers a unified, hierarchical framework to meet diverse power‑reduction targets with minimal workload disruption.

Method details
  • Hierarchical design with regional orchestrator and per‑datacenter local controllers
  • Decision Manager maintains inventory of racks, servers, VMs, batteries and generators
  • Impact Estimator synthesizes Decision Table and Impact Table from telemetry
  • Offline mode pre‑computes action playbooks based on probabilistic load models
  • Online mode uses Consolidator to enforce most restrictive limit and selects actions
  • Hardware controllers provide sub‑second response for energy‑storage actions
Limitations

The paper does not provide quantitative evaluation of the control algorithms or compare against alternative power‑modulation approaches.

Service B can tolerate more capacity loss than Service AFound in the source text, word for word.

Picked because: Presents Slasher, a concrete system for power‑flexibility in cloud datacenters that can be integrated into DevOps pipelines for dynamic power management.

Paper 2 of 5

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

Srimonti Dutta, Akshata Kishore Moharir · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Answer-only evaluation was considered sufficient, but structured-data tasks can produce benchmark-correct answers with invalid computation traces, leading to silent failures that the obvious fix of measuring only answer accuracy does not catch.

Approach

The paper defines Trace Integrity as a set of criteria (explicit, executable, schema‑valid, operator‑faithful, replayable, answer‑consistent, auditable) and operationalizes it with execution contracts that bind user intent, schema elements, operator plans, assumptions, executable queries, verification status, and the final answer. It also introduces the CAIT (Correct Answer / Invalid Trace) Rate to quantify how often correct‑looking answers rely on invalid traces. Three prompting strategies are compared under the same executor and a deterministic trace validator, allowing separation of answer correctness from trace validity. By evaluating both answer accuracy and trace‑level metrics, the method reveals distinct deployment reliability signals.

Result

Across the three prompting conditions, answer accuracy ranged from 20.0% to 24.0% while Trace Integrity Pass Rates ranged from 39.0% to 43.0%, and CAIT Rates remained high (55.0%, 59.1%), demonstrating that correct answers often coexist with invalid traces and that these signals are orthogonal.

Why it matters

Developers and deployment teams of LLM data agents should care because evaluating only answer correctness can miss silent computational failures that affect reliability and compliance.

Method details
  • Model: claude‑haiku‑4‑5 with temperature 0.0
  • Dataset: 100 BIRD Mini‑Dev examples
  • Prompting conditions compared: Direct SQL, Operation Summary + SQL, Contract‑First SQL
  • Metrics reported: answer accuracy, Trace Integrity Pass Rate, CAIT Rate
  • Validator: deterministic operator‑level checks for schema validity, operator matching, and answer‑trace consistency
Numbers
  • Answer accuracy, 20%, Direct SQL
  • Answer accuracy, 22%, Operation Summary + SQL
  • Answer accuracy, 24%, Contract‑First SQL
  • Trace Integrity Pass Rate, 39%, Direct SQL
  • Trace Integrity Pass Rate, 43%, Operation Summary + SQL
  • Trace Integrity Pass Rate, 40%, Contract‑First SQL
Limitations

The study is a scoped proof‑of‑concept using a single model, a single 100‑example dataset, and deterministic validation that does not capture full semantic equivalence, so results are not generalizable across models or datasets.

Answer accuracy is an insufficient reliability signal for LLM data agents.Found in the source text, word for word.

Picked because: Introduces the Trace Integrity framework, offering concrete criteria and tooling to audit and verify LLM data‑agent computations for reliable deployment.

Paper 3 of 5

ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs

Somgyuan Li, Ahmed M. Abdelmoniem, Shiqiang Wang · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing cascade routing methods make one-shot, query-level decisions and cannot adapt to the dynamic, state‑dependent nature of multi‑step workflows, leading to cost‑infeasible single‑model policies.

Approach

ProgRouter routes LLM agents online by scoring task progress with a multi‑view progress scorer, predicting future progress gain with a dual‑path predictor (structured and semantic paths) and an adaptive meta‑gating mechanism, and enforcing long‑term cost limits via a virtual‑queue budget tracker. At each workflow step the system estimates the progress gain for each candidate LLM and selects the one that best balances expected gain against the remaining time and energy budget.

Result

ProgRouter attains the highest quality‑cost tradeoff on all four benchmarks, achieving 93.0% pass on HumanEval Plus within a 4800 J budget, 79.4% pass, 3376 J energy and 10.3 s execution on MBPP, 84.3% pass, 6112 J energy and 19.0 s execution on MATH‑500, and 92.1% citation precision on ASQA while staying under the 19000 J budget, outperforming all baselines.

Why it matters

Researchers and engineers building multi‑agent LLM workflows should care because ProgRouter demonstrates how to maintain high task quality while respecting realistic energy and time budgets.

Method details
  • Coding model zoo includes Qwen 2.5‑Coder (0.5B, 7B, 14B, 32B) and Qwen 3.5 (2B, 4B, 9B, 27B, 35B).
  • Math model zoo includes Granite 4.1 (3B, 8B, 30B) and Gemma 4 (2B, 4B, 26B, 31B).
  • QA model zoo includes Qwen 3.5 (2B, 4B, 9B, 27B, 35B) and Qwen 3.6 (27B, 35B).
  • Benchmarks: HumanEval Plus (164 tasks), MBPP (200 coding tasks), MATH‑500 (200 math problems), ASQA (100 QA tasks).
  • Baselines compared: MasRouter, CASCADIA, Educated Guessing.
  • Ablation: Removing the progress predictor drops pass rate from 93.0% to 89.0% on HumanEval Plus.
Numbers
  • HumanEval Plus pass 93.0% within 4800 J budget
  • MBPP pass 79.4%, energy 3376 J, time 10.3 s
  • MATH‑500 pass 84.3%, energy 6112 J, time 19.0 s
  • ASQA citation precision 92.1% within 19000 J budget
  • Gemma 4 31B consumes up to 26276 J on MATH‑500
  • Educated Guessing 98.8% Granite 4.1 8B on MATH‑500
Limitations

The paper does not state explicit limitations; none are reported.

ProgRouter achieves the highest citation precision (92.1%) among all evaluated methods while remaining within the 19000 J energy budgetFound in the source text, word for word.

Picked because: Describes ProgRouter, an online orchestration layer that routes multi‑agent LLM workflows based on real‑time quality‑cost trade‑offs, with released code for practical use.

Paper 4 of 5

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

Tongyan Hu, Bryan Hooi · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Existing defenses are static: safety behavior is fixed at deployment and cannot accumulate defensive experience or adapt to new jailbreak strategies, and simply updating model parameters at test time is not feasible.

Approach

The paper introduces a self‑evolving multi‑agent framework that maintains a persistent rule memory. When a jailbreak succeeds, the Self‑Reflection Agent abstracts the failure into a method‑level rule and stores it. The Rule Triggering Agent classifies inputs and retrieves relevant rules, the Policy Decision Agent selects a response mode (ALLOW, HARD_REFUSE, SOFT_REFUSE), and the Response Generation Agent produces the final output according to the chosen policy. All components operate via prompting without any parameter updates, enabling test‑time adaptation across both open‑weight and black‑box models.

Result

Across four black‑box jailbreak families and multiple models, the proposed framework substantially reduces attack success rates while preserving benign utility, remains robust under an adaptive composite‑wrapper attack, and does not increase over‑refusal as the memory grows.

Why it matters

LLM developers and safety engineers should care because the method offers a parameter‑free, test‑time defense that can continuously adapt to emerging jailbreak techniques.

Method details
  • Open‑source models: instruction‑tuned Qwen and Llama families.
  • API‑based models: Gemini‑class proprietary models accessed via standard interfaces.
  • Safety evaluation dataset: Advbench with 520 harmful prompts.
  • Benign utility datasets: MMLU and GSM8K.
  • Baseline defenses compared: No Defense, Defense Prompt, Self‑Reminder, AutoDefense.
  • Attack families evaluated: DeepInception, CodeChameleon, ReNeLLM, FlipAttack.
Numbers
  • Advbench prompts, 520, dataset size
  • ASR‑gpt threshold τ, 7, judge threshold
  • Rule Triggering Agent selects up to, 2, rule ids per input
  • Memory initialized empty, 0, initial rule count
Limitations

The paper does not explicitly discuss any limitations.

our method substantially reduces attack success rates while preserving benign utilityFound in the source text, word for word.

Picked because: Provides a self‑evolving multi‑agent defense against LLM jailbreak attacks, delivering actionable mechanisms and open‑source components for securing LLM services.

Paper 5 of 5

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

Jiarui Yan, Weiwei Sun, Sijie Li and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

LLMs can write correct isolated code but are far weaker at autonomous machine‑learning development, collapsing into narrow loops that never pivot or reopen abandoned work; simple instruction prompts only close the part of the gap that reduces to instructions.

Approach

The authors release TraceML, a version‑level corpus of human and agent Kaggle development traces with rich labels, and train open‑weight labelers (teacher gpt‑5.4‑mini, students Qwen3‑1.7B) to annotate them. They then design a compact planning prompt (“skill”) that encodes anti‑loop constraints, human‑prior practices, periodic self‑checks, and task‑specific priors. In the harness experiment the same Codex CLI backend runs under a 12‑hour budget with either the baseline prompt or the skill injected every 30 minutes. Two ablation arms isolate the effect of the prompt content versus the injection schedule. Performance is measured as the best valid held‑out score compared to human percentile bands.

Result

Across the seven paired competitions the harness prompt improves scores in five competitions, stays within noise in two, and never regresses. For example, gquest Spearman rises from 0.371 to 0.429, aes2 QWK from 0.771 to 0.817/0.808, and hms KL from 1.050 to 0.718. Ablation B (schedule only) matches or falls below baseline everywhere, confirming that gains come from the prompt content.

Why it matters

Researchers building autonomous ML agents should care because the TraceML corpus and the planning prompt reveal concrete behavioral gaps and a simple instruction‑based lever that can improve performance on realistic Kaggle tasks.

Method details
  • Codex CLI agent (version 0.146.0) runs on the gpt‑5.4‑mini backend.
  • MLEvolve evolutionary search agent also runs on gpt‑5.4‑mini.
  • Labeler teacher model: gpt‑5.4‑mini; student models: two Qwen3‑1.7B.
  • Dataset: 4,465 human Kaggle trajectories across 134 competitions; paired subset: 7 competitions with 430 human and 207 agent trajectories.
  • Baseline arm uses the standard task prompt; harness arm adds the planning skill and re‑injects it every 30 minutes.
  • Ablation A delivers a single content block; Ablation B keeps the 30‑minute cadence but strips the planning content.
Numbers
  • gquest Spearman 0.429 vs baseline 0.371
  • aes2 QWK 0.817 (or 0.808) vs baseline 0.771
  • hms KL 0.718 vs baseline 1.050
  • commonlit RMSE 0.505/0.517 vs baseline 0.510
  • equity C-index 0.670/0.680 vs baseline 0.675
  • five of seven competitions improve, two are within noise, none regress
Limitations

The paper does not show that the planning prompt can fully close the human‑agent gap; it only yields partial score gains and leaves the effort profile agent‑shaped.

A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores, but the effort profile stays agent-shaped.Found in the source text, word for word.

Picked because: Offers an empirical study of human‑agent interaction in ML development pipelines, yielding actionable best‑practice recommendations for engineers building autonomous ML tooling.