arXiv digest

Thursday

October 1, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Turbo Harness: Instance-Adaptive Harness Optimization

Tunyu Zhang, Hao Wang, Kai Xu and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing harness optimization produces a single global harness applied uniformly, which may be suboptimal for individual task instances, and naïvely searching per instance is infeasible.

Approach

Turbo Harness reuses the exhaust (traces, reflections, evaluations) from a completed global harness search to build a structured playbook. A lightweight LLM harness editor is trained with reinforcement learning to condition on both the task instance and the playbook, generating a patch that adapts the global harness into an instance‑specific harness. At inference the editor produces the patch, which is applied (or falls back to the global harness) before the frozen execution model runs. The editor is trained with GRPO using task performance as reward, while the execution model remains frozen.

Result

On SWE‑smith‑MR with a frozen Claude Haiku 4.5 executor, the globally optimized Meta‑Harness achieves 50.7% pass rate. An untrained 9B editor without the playbook reaches 51.3%, with the playbook 50.0%, and RL‑trained without the playbook 49.3%. Combining RL training with playbook conditioning raises the pass rate to 64.0%, a 14.7‑point improvement over the RL‑trained editor without the playbook. Stronger frozen editors achieve 55.3% (Sonnet‑4.5) and 61.3% (Opus‑4.6), but the RL‑trained Qwen3.5‑9B matches the best.

Why it matters

Researchers building self‑improving agents and harness optimization pipelines should care because Turbo Harness shows that lightweight, RL‑trained editors can extract and apply global search knowledge to boost per‑instance performance without retraining the main model.

Method details
  • Harness editor: Qwen3.5-9B full‑parameter fine‑tuned with FSDP and GRPO.
  • Execution models: frozen LLMs per benchmark (Qwen3.5‑9B for agentic tasks, Claude Haiku 4.5 and Gemini 3.7 Flash for coding, Claude Sonnet 4.5 for TB2.1).
  • Benchmarks: seven tasks (ALFWorld, ScienceWorld, DBBench, WebShop, SWE‑smith‑MR, SWE‑bench Verified, Terminal‑Bench‑2.1).
  • Baselines: Meta‑Harness, Default (mini‑swe‑agent or basic tool‑calling), Terminus‑Kira, Terminus‑2, ReAct, Self‑Refine, Reflection, Harness‑R1.
  • Ablations: compare untrained editor, playbook only, RL only, and combined RL+playbook; also vary editor model (Qwen3.5‑9B, Sonnet‑4.5, Opus‑4.6).
Numbers
  • Pass rate, 64.0%, compared to Meta‑Harness 50.7%
  • Pass rate, 61.3%, compared to Meta‑Harness 50.7%
  • Pass rate, 51.3%, compared to Meta‑Harness 50.7%
Limitations

The paper does not claim to evaluate generalization to unseen repositories or to budgets beyond those tested.

Combining RL training with playbook conditioning raises the pass rate to 64.0%, a 14.7-point improvement over the same RL-trained editor without the playbook.Found in the source text, word for word.

Picked because: Introduces Turbo Harness, a concrete framework for instance-adaptive harness optimization that can be directly applied to improve LLM agent tooling.

Paper 2 of 5

Compression Footprints as Security Signals for Model-Poisoning Defense in Federated Learning

Sachi Shome, William Eiers · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Existing robust aggregation defenses treat lossy compression only as an error source and rely on update geometry, which modern model‑poisoning attacks can evade, and the obvious fix of using client‑reported compression metadata is insecure because Byzantine clients can falsify it.

Approach

CRAFT is a server‑side robust aggregation method that extracts a five‑dimensional compression footprint from each decompressed client update using the SZ2 error‑bounded lossy compressor, robustly scales these footprints, identifies a majority‑supported footprint core, computes a continuous trust weight for each client based on its distance to the core, and finally aggregates the decompressed updates with a weighted coordinate‑wise median.

Result

Across 18 dataset‑attack settings CRAFT attains the highest final test accuracy in 7 settings and is within 1.7 percentage points of the best method in the remaining settings, while matching mean performance on clean IID training.

Why it matters

FL system designers and researchers needing Byzantine‑resilient aggregation can adopt CRAFT to gain robustness without extra communication overhead, leveraging existing compression pipelines as a security signal.

Method details
  • Uses SZ2 EBLC in relative‑error mode to recompress each received update on the server
  • Simulates 25 clients with 9 malicious (36% participation) under IID data partitions
  • Evaluates on three datasets: CIFAR‑10, Fashion‑MNIST, and Purchase
  • Compares against six robust aggregation baselines: Mean, Median, Trimmed Mean, Krum, Bulyan, and DnC
  • Applies the same trust‑weight parameters (footprint‑distance scale and trust‑decay power) across all main experiments
Numbers
  • 7/18 settings best final accuracy
  • within 1.7 percentage points of the best in the others
  • CIFAR‑10 ALIE 58.83 % (best, next DnC 57.31 %)
  • Fashion‑MNIST ALIE 91.97 % (best, next DnC 88.27 %)
  • Purchase ALIE 81.13 % (close to DnC 81.76 %)
Limitations

The paper does not address backdoor attacks, privacy attacks, non‑IID data distributions, or attacks targeting the compression library itself.

CRAFT consistently achieves the best accuracy in 7 out of 18 settings and within 1.7 percentage points of the best in the others.Found in the source text, word for word.

Picked because: Presents a practical security signal using compression footprints to detect model‑poisoning attacks in federated learning, useful for self‑hosted infrastructure security.

Paper 3 of 5

cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Existing CUA benchmarks suffer from a reproducibility crisis because they rely on complex infrastructure with varying machine and container configurations that confound speed evaluation, and simply measuring wall‑time does not isolate the agent's performance.

Approach

cua‑speedrun introduces a uniform virtual‑machine setup, a consistent execution pipeline, and a common agent interface that decouple the four dimensions of evaluation (agent, benchmark, environment infrastructure, and agent loop). It measures speed from the moment the instruction is given until the agent terminates, while counting model responses and desktop actions separately. Model serving is handled via API endpoints or a self‑hosted vLLM server on L40S GPUs. Task‑selection methods reduce the number of evaluation tasks without losing statistical power. The framework records score, wall‑time, cost, and token usage for each configuration.

Result

Across the benchmarks, configurations achieve up to 91.6% score with varying time and cost; for example GPT‑6 Astra / xhigh reaches 91.6% score in 126.8 s at $0.708. Energy‑based task selection yields a mean error of 4.42 percentage points, better than 5.08 for difficulty‑stratified random selection and 8.08 for an IRT‑inspired method.

Why it matters

Researchers and engineers developing computer‑use agents should care because cua‑speedrun provides a reproducible, cost‑aware benchmark for evaluating and comparing agent speed and efficiency across diverse tasks.

Method details
  • Model serving uses API endpoints or a self‑hosted vLLM inference server on L40S GPUs.
  • Evaluation runs on four CUA benchmarks with a fresh VM environment for each task.
  • Task selection reduces OSWorld from 295 tasks to 32 tasks using energy‑based selection.
  • Model responses and desktop actions are counted separately; a step is one request containing keyboard, mouse, or wait actions.
  • Cost is computed from provider‑recorded charges or token usage, excluding environment hosting and verification calls.
Numbers
  • Score 91.6%, Time 126.8 s, Cost $0.708 for GPT‑6 Astra / xhigh
  • Score 91.6%, Time 250.1 s, Cost $0.221 for Gemini 3.8 Flash / medium
  • Score 90.8%, Time 89.5 s, Cost $0.531 for GPT‑6 Astra / low
  • Mean error 4.42 pp for energy‑based selection (vs 5.08 pp random, 8.08 pp IRT)
  • Cost $0.108 for Gemini 3.8 Flash / low with 91.6% score
Limitations

The paper notes that validation gains do not generalize to a held‑out model, indicating limited predictive power beyond the evaluated models.

Benchmarks are based on complex infrastructure with varying machine and container configurations that confound the evaluation of the execution speed of CUAs.Found in the source text, word for word.

Picked because: Provides a standardized benchmark of computer‑use agent speed and cost, giving engineers actionable data for deploying faster, cheaper agents.

Paper 4 of 5

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

Yong Du, Tongbo Chen, Zhengxi Lu and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CV

Problem

Existing computer-use agents rely on sparse outcome rewards that give no supervision for intermediate actions, and applying on-policy self-distillation (OPSD) directly suffers from fixed guidance becoming misaligned with the student and guidance‑induced probability shifts conflicting with step‑level correctness.

Approach

ComputerSD introduces an online self‑distillation loop where a fine‑tuned GUI analyzer generates step‑level guidance and a value score after each action. The value score gates the token‑level OPSD signals, regulating their strength. These gated token‑level signals are combined with trajectory‑level GRPO loss. The whole system is trained asynchronously, sampling multiple trajectories in parallel and updating the policy with both token‑level and trajectory‑level objectives.

Result

ComputerSD raises success rates to 39.8% on Qwen3-VL-8B-Thinking and 47.9% on EvoCUA-8B, surpassing outcome‑only GRPO (37.9% and 43.8%) and the respective baselines (33.8% and 41.3%). It also improves OOD performance, adding 4.5 points on OSWorld‑Verified OOD for Qwen3-VL-8B-Thinking and 3.8 points for EvoCUA-8B, and gains 5.3 points on the cross‑platform WindowsAgentArena benchmark for Qwen3-VL-8B-Thinking.

Why it matters

Researchers and engineers building computer‑use agents should consider ComputerSD because it provides token‑level supervision from real‑time feedback, leading to higher success rates and better generalization without extra inference cost.

Method details
  • GUI analyzer initialized from Qwen3-VL-8B-Thinking and fine‑tuned with expert annotations.
  • Expert model for annotation is Kimi‑K3.
  • Applied to two 8B‑scale backbones: Qwen3-VL-8B-Thinking (general) and EvoCUA-8B (specialized).
  • Training data: OSWorld‑Verified with 222 in‑domain and 139 OOD tasks.
  • Training runs for 180 policy updates with batches of 32 trajectories (8 trajectories per 4 tasks).
  • Ablations include removing GUI analyzer SFT, removing real‑time feedback, and removing the value gate.
Numbers
  • Success rate 39.8% vs GRPO 37.9% on Qwen3-VL-8B-Thinking (general)
  • Success rate 47.9% vs GRPO 43.8% on EvoCUA-8B (specialized)
  • In‑domain OSWorld improvement +7.0 points for Qwen3-VL-8B-Thinking
  • OOD OSWorld improvement +4.5 points for Qwen3-VL-8B-Thinking
  • In‑domain OSWorld improvement +8.4 points for EvoCUA-8B
  • OOD OSWorld improvement +3.8 points for EvoCUA-8B
Limitations

The paper does not explicitly discuss any limitations of the proposed method.

For each backbone, we compare ComputerSD with outcome-only GRPO under the same online training and evaluation settings.Found in the source text, word for word.

Picked because: Describes ComputerSD, an online self‑distillation method with real‑time feedback for computer‑use agents, offering a ready‑to‑use technique for continuous improvement.

Paper 5 of 5

PhantomEnvironments: Training LLM Agents in Fictional Worlds

Anmol Kabra, Swathi Saravana Selvam, Albert Gong and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Existing RL training for LLM agents relies on costly human-curated data or LLM-generated environments that risk hallucinations and benchmark contamination, making scalable, verifiable, long-horizon environments unavailable.

Approach

The paper builds rule‑generated synthetic worlds (PhantomEnvironments) from the PhantomWiki dataset, creating templated articles and context‑free‑grammar multi‑hop questions. These are indexed for retrieval, and agents are trained via RL using the GRPO algorithm with XML‑tagged search and answer actions. Dense retrieval during training uses an e5‑base‑v2 encoder and a flat FAISS index fetching the top‑3 documents. Four LLM families are fine‑tuned for one epoch with two seeds, and performance is evaluated on real‑world multi‑hop benchmarks.

Result

Synthetic training consistently improves F1 scores on all six real‑world multi‑hop search benchmarks for all four LLMs, with the largest gains on newer and harder datasets such as SynthWorlds‑RM, SynthWorlds‑SM, and FRAMES, and the improvements persist across training checkpoints without overfitting.

Why it matters

Researchers and engineers building RL‑fine‑tuned LLM agents should consider rule‑generated fictional environments as a cheap, scalable source of training data that yields transferable search abilities.

Method details
  • Trains Qwen2.5-3B-Instruct, Qwen2.5-7B-Instruct, Llama-3.2-3B-Instruct, and Phi-4-mini-instruct.
  • Uses GRPO RL algorithm with F1 reward computed after SQuAD‑style normalization.
  • Trains on 55K synthetic questions and 55K real NQ+HotpotQA questions for comparison.
  • Retrieves top‑3 documents per query with intfloat/e5-base-v2 dense encoder and flat FAISS index.
  • Evaluates with Qwen3-Embedding-4B dense retriever, up to 20 turns and 32768 token context limit.
  • Runs each setting for 1 epoch with 2 independent training seeds on 2 B200s GPUs over 2 days.
Numbers
  • training questions, 55K, synthetic vs real
  • maximum hops per question, 7, rule‑generated chains
  • documents retrieved per query, top‑3, training retriever
  • evaluation questions per benchmark, 500, standard sets
  • FRAMES evaluation questions, 824, benchmark size
  • training seeds, 2, independent runs
Limitations

The paper does not establish performance beyond the evaluated benchmarks or on truly novel real‑world domains.

Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply.Found in the source text, word for word.

Picked because: Shows how to train LLM agents in synthetic fictional worlds, eliminating the need for costly real environments and enabling scalable agent development.