arXiv digest

Wednesday

September 9, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Distributed Linear Programming on GPU Clusters at Extreme Scale

Arnaud Deza, Santanu Dey, Pascal Van Hentenryck · abstract · pdf

quote unverifiedfigures checkedread: full textmath.OC

Problem

Previous GPU LP solvers kept the matrix or primal‑dual state on a single node, so phases other than matrix‑vector multiplication re‑introduced a single‑node memory limit; simply moving the matrix‑vector product to GPUs does not solve the overall memory bottleneck.

Approach

ShardLP distributes both the constraint matrix and the primal‑dual vectors across GPUs from input sharding through output, maintaining persistent ownership. It implements the PDHG (Chambolle‑Pock) iteration with diagonal scaling, presolve, adaptive step sizes, restart, and feasibility polishing on the distributed data. Support‑aware communication skips GPUs that store no coefficients for a given row in column‑partitioned solves. The system runs one MPI rank per GPU and uses collective reductions only for scalars or partition‑sized data, never gathering the full matrix or solution.

Result

ShardLP meets the published PDLP acceptance criterion on nine of eleven Google benchmark instances, exceeding the CPU baseline which succeeded on eight. On the largest instance it solves in 593 s of solver time (933 s end‑to‑end), compared with 21.06 h on the CPU. Scaled runs solve a 13.604 billion‑variable, 40.807 billion‑nonzero LP in 1,905 s on 76 GPUs, and support‑aware communication cuts communication by 92.97% while improving runtime by up to 1.52×.

Why it matters

Researchers and practitioners needing to solve extremely large sparse LPs that exceed single‑node memory will benefit from ShardLP's distributed memory ownership and communication optimizations.

Method details
  • Largest benchmark: 1.185 billion variables and 6.338 billion nonzeros.
  • Hardware: 8 H200 GPUs on a single node for the Google PDLP benchmark; up to 76 GPUs across 29 nodes for larger solves.
  • Datasets: Google PDLP eleven instances, synthetic multicommodity‑flow (MCF13.60B, MCF8.35B, MCF8B, MCF5B), KDD12 L1‑SVM, QAPLIB tai256c AJ, MS1 matching.
  • Baselines: published CPU PDLP study (Applegate et al., 2026) and prior GPU systems D‑PDLP and MPAX.
  • Ablation: support‑aware communication reduces modeled data movement by 92.97% and yields solver‑time speedups of 1.27×‑1.52×.
Numbers
  • 1.185 billion variables, 6.338 billion nonzeros, 593 s solver time vs 21.06 h CPU
  • 9 of 11 instances solved vs 8 of 11 for CPU baseline
  • 13.604 billion variables, 40.807 billion nonzeros solved in 1,905 s on 76 GPUs
  • Support‑aware communication cuts modeled communication by 92.97%
  • Solver‑time speedup from support‑aware communication: 1.27×‑1.52×
Limitations

ShardLP requires the LP to be pre‑sharded offline and currently supports support‑aware communication only for column partitions with fixed matrix support.

The model could not quote the source for its claim, so nothing above has been checked against the paper. Read the abstract before trusting it.Citation check failed.

Picked because: Introduces SHARDLP, a released distributed GPU linear programming solver that demonstrates practical techniques for scaling compute workloads across clusters.

Paper 2 of 5

ExecCritic: Learn to Test, Test to Improve for Coding Agents

Leitian Tao, Baolin Peng, Haorui Wang and 7 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Execution feedback only helps coding agents when the tests fully capture the requested behavior, but agent‑generated tests are often incomplete or incorrect, and when the same trajectory creates both the patch and the test their errors can align, giving false confidence.

Approach

ExecCritic separates testing from repair by introducing two distinct agents that share the same Qwen‑3.5‑35B‑A3B backbone. The Test agent independently generates a repository‑native test bundle (code, execution command, JSON contract) which the harness validates and then freezes. The Repair agent then revises the source code using the fixed test’s execution feedback without altering the test. Training occurs in two stages: Learn to Test (SFT then RL to produce behaviorally valid tests) and Test to Improve (feedback‑conditioned repair training). This role‑specific reinforcement learning allows each agent to be optimized for its own objective while remaining compatible at inference time.

Result

Post‑training, the Qwen Test agent’s Base‑to‑Gold success rises from 22.2% to 62.2%, and when paired with the post‑trained Qwen Repair agent the overall resolved rate reaches 72.6%, an 11.4‑point improvement over the original no‑test baseline of 61.2%. GPT‑5.6‑sol generated tests raise the baseline to 65.3% while Qwen‑generated tests drop it to 57.3%.

Why it matters

Researchers and engineers building coding assistants should care because separating test generation from repair and training them with role‑specific RL markedly improves feedback‑driven patch resolution without requiring larger models.

Method details
  • Backbone model for both agents: Qwen‑3.5‑35B‑A3B.
  • Training data: SWE‑ReBench for both Test and Repair agents.
  • Test‑agent training: SFT on DeepSeek‑V4‑Flash‑0731 trajectories then 200 GRPO RL steps with dynamic sampling (16 issue groups, 8 rollouts per group, up to 5 test attempts).
  • Repair‑agent training: 100 RL steps, same dynamic sampling, up to 5 feedback‑guided revision rounds of 40 turns each.
  • Baselines compared: untrained Base‑35B, its SFT checkpoint, Codex‑5.3, DeepSeek‑V4‑Flash‑0731, and GPT‑5.6‑sol.
  • Ablations include: standard repair vs. Test‑to‑Improve with/without direct‑solve bonus (Table 3) and SFT teacher comparison (Table 4).
Numbers
  • Base‑to‑Gold success after RL, 62.2%, compared to SFT baseline 39.6%
  • Composed Qwen agents resolved rate, 72.6%, an 11.4‑point gain over no‑test baseline 61.2%
  • GPT‑5.6‑sol generated tests raise no‑test baseline, 65.3%, compared to baseline 61.2%
  • Qwen‑generated tests reduce no‑test baseline, 57.3%, compared to baseline 61.2%
  • Full Test‑to‑Improve final resolved rate, 72.6%, versus Standard training final 70.3%
  • Cross‑language Rust test‑generation success, 66.7%, compared to Python 62.2%
Limitations

The paper does not demonstrate that generated tests can fully replace privileged Oracle fail‑to‑pass feedback, and it shows limited improvement on the SWE‑bench Pro subset where test coverage is insufficient.

The trained Test agent therefore reaches a level comparable to Codex-5.3 at , while remaining points below DeepSeek-V4-Flash-0731 and points below GPT-5.6-sol.Found in the source text, word for word.

Picked because: Presents ExecCritic, a concrete test‑verify‑revise pipeline for coding agents that engineers can adopt to improve automated code repair reliability.

Paper 3 of 5

ToolLoop: Closed-Loop Tool-Use Data Synthesis via Decomposed Generation and Dynamic Self-Feedback

Min Zeng, Yuzhou Liu, Zhenyu Cao and 5 others · abstract · pdf

quote verified6 figures not in sourceread: full textcs.CL

Problem

Existing synthetic tool-use data pipelines use a generate‑then‑filter paradigm with static post‑hoc verification, which leads to inefficient data and imbalanced feature distributions; simply adding more filtering does not fix the issue because static filters only discard bad samples without correcting underlying generation errors.

Approach

ToolLoop introduces a closed‑loop framework that decomposes data synthesis into three progressive stages-ground truth sampling, backward derivation of user queries, and forward derivation of tool calls. Each stage incorporates dynamic self‑feedback that verifies and refines the model's output before moving to the next stage. This transforms the pipeline from generate‑then‑filter into a generate‑verify‑refine loop. The feedback mechanism iteratively guides the model toward high‑quality generation, ensuring coherence, parallelizability, and correct argument grounding. The overall system is applied to a 4B‑parameter base model and trained on a small synthetic dataset.

Result

ToolLoop‑4B achieves 86.40% overall accuracy on BFCL in non‑reasoning mode, with 91.29% on non‑live scenarios and 81.50% on live scenarios, while the isolate variant reaches 86.07% overall. On ACEBench, ToolLoop‑4B attains 72.1% overall accuracy using only 18.3% of the training data required by APIGen‑4B.

Why it matters

Researchers and engineers building tool‑use language models should care because ToolLoop delivers higher accuracy with far fewer synthetic examples, enabling more data‑efficient fine‑tuning for tool‑calling tasks.

Method details
  • Base model: Qwen3-4B-Instruct-2507 (4 B parameters)
  • Synthetic training set: 11 K examples for ToolLoop‑4B, 10 K for the isolate variant
  • Benchmarks: Berkeley Function Calling Leaderboard (BFCL‑v4, 2 501 test instances) and ACEBench
  • Baselines compared: APIGen‑4B (60 K examples), ToolMind‑4B (55 K examples), plus several commercial and open‑source models
  • Ablation variants: w/o Feedback (79.97% overall), w/ Final Filtering (82.56% overall), full ToolLoop (86.40% overall)
Numbers
  • 86.40% overall accuracy on BFCL (ToolLoop‑4B) vs 83.11% for APIGen‑4B (3.29‑point gain)
  • 86.07% overall accuracy for ToolLoop‑4B‑Isolate (0.33 points below full model)
  • 91.29% non‑live BFCL accuracy (ToolLoop‑4B) vs 89.90% for APIGen‑4B (1.39‑point gain)
  • 81.50% live BFCL accuracy (ToolLoop‑4B) vs 76.31% for APIGen‑4B (5.19‑point gain)
  • 72.1% overall ACEBench accuracy (ToolLoop‑4B) vs 67.0% for APIGen‑4B (5.1‑point gain) using 18.3% of the training data
  • 79.97% overall accuracy without feedback (ablation) vs 86.40% with full ToolLoop
Limitations

The paper shows limited improvement on the Profile category of ACEBench, where all fine‑tuned models underperform the base model, and it still trails the best single‑turn performance (ToolMind‑4B).

a 4B parameter model trained with our 11K synthetic examples achieves 86.40% accuracyFound in the source text, word for word.
These figures do not appear in the source text: 10 K, 11 K, 4 B, 5.19, 55 K, 60 K. Treat them as unverified.Number check failed.

Picked because: Describes ToolLoop, a closed‑loop framework for synthesizing high‑quality tool‑use data, enabling engineers to build and evaluate LLM‑driven tool agents.

Paper 4 of 5

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

Zhou Yu, Bin Bi, Shiva Kumar Pentyala and 8 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Imitating expert trajectories under the evolved harness caused the weaker model to regress on all seven tasks, because the weaker model adopted the expert's planning strategy without the competence to execute it, breaking the fit with the harness.

Approach

The paper introduces an on-policy expert-correction pipeline driven by a meta-level MLE agent. The agent first generates rollouts of the weaker model under the evolved harness, then localizes the single failing turn in each rollout. The expert model rewrites only that turn, preserving the weaker model's original planning distribution. The corrected trajectories are used to LoRA‑SFT the weaker model. This method keeps the model‑harness fit while injecting expert knowledge.

Result

Evolving the harness raised mean test success from 29.2% to 78.0% (+48.8 points). The expert model improved from 84.4% to 93.6% (+9.2) on the evolved harness. Full imitation caused an average regression of 14.9 points, while on‑policy correction increased mean success to 79.7% (+1.7) and yielded gains on five of seven tasks.

Why it matters

Practitioners aiming to improve smaller LLMs on domain‑specific enterprise tasks can gain modest performance boosts without costly full model training by combining harness evolution with on‑policy expert correction.

Method details
  • Weaker model: Qwen3‑Coder‑30b‑a3b‑instruct (30B total / 3B active).
  • Stronger expert model: gemini‑3.1‑pro‑preview (used for correction and as teaching signal).
  • Additional model family tested: Gemma‑4‑26b‑a4b‑it (26B/4B active).
  • Dataset: seven enterprise agentic benchmarks (payroll auditing, budget approval, stock alerting, IoT anomaly detection, browser automation, website management, code refactoring).
  • Fine‑tuning: LoRA‑SFT on H100/H200 GPUs, converting expert trajectories to student chat and tool‑call format.
  • Baselines: base harness, evolved harness, SFT‑imit (full imitation), SFT‑corr (on‑policy correction).
Numbers
  • base harness mean success 29.2%
  • evolved harness mean success 78.0% (+48.8 points)
  • expert model success under base harness 84.4%
  • expert model success under evolved harness 93.6% (+9.2)
  • imitation regression average 14.9 points
  • on‑policy correction mean success 79.7% (+1.7)
Limitations

The approach does not fully close the gap to the expert model, likely due to limits of lightweight LoRA fine‑tuning.

On-policy correction matches or beats the evolved-harness base on every taskFound in the source text, word for word.

Picked because: Shows how automated harness evolution can boost weaker models, offering actionable methods for constructing robust LLM agent pipelines.

Paper 5 of 5

PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games

Ryan Truong, Lance Ying, Samuel J. Gershman and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Prior to this work, creating new video-game environments or adapting existing ones required extensive hand‑coding, making the process laborious. The obvious fix of manually editing code does not scale because each game must be individually programmed and integrated into the RL pipeline.

Approach

PlayTrain first uses an LLM‑prompting pipeline to generate or modify a JavaScript game from a short human description. The generated JS file is then compiled ahead of time into a shared library and executed inside a QuickJS interpreter embedded in the trainer process. QuickJS is coupled with a custom C++ p5 layer that forwards drawing commands to a Rust rasterizer which writes directly into the observation buffer. The environment runs as a standard OpenAI‑gym instance, allowing any RL algorithm to interact with it. Double‑buffered and single‑buffered execution modes are supported to compare throughput against baseline C++ environments.

Result

PlayTrain achieves over 1M agent‑steps per second on a single GPU node, with double‑buffered IMPALA‑Nature‑CNN reaching 1.07M steps/s across all 24 games. Per‑core, PlayTrain is 12.62× faster than ALE and 2.19× faster than ProcGen, while QuickJS runs 13.4× faster than Node/V8 and 117× faster than a headless browser. In multi‑threaded training, PlayTrain clones are on average 2.25× faster than ProcGen and 5.80× faster than ALE.

Why it matters

Researchers and engineers building RL benchmarks will benefit from rapid, LLM‑driven generation of high‑performance JavaScript environments without hand‑coding, enabling faster experimentation.

Method details
  • Uses Nature‑CNN and IMPALA‑CNN vision encoders
  • Trains PPO and IMPALA algorithms
  • Clones 8 Atari (ALE) and 16 ProcGen games for a total of 24 environments
  • Baseline comparison against original C++ EnvPool implementations
  • Evaluates double‑buffered versus single‑buffered execution
  • Measures per‑core and multi‑thread throughput on a node with four H100 GPUs and 92 CPU cores
Numbers
  • agent‑steps/s, 1.07M, double‑buffered IMPALA with Nature‑CNN across all 24 games
  • agent‑steps/s, 0.35M, double‑buffered IMPALA‑CNN across all 24 games
  • speedup, 12.62×, PlayTrain vs ALE per‑core geometric average
  • speedup, 2.19×, PlayTrain vs ProcGen per‑core geometric average
  • backend speedup, 13.4×, QuickJS vs Node/V8
  • backend speedup, 117×, QuickJS vs browser (Playwright)
Limitations

The paper does not demonstrate performance on games whose logic runs per frame rather than drawing commands, where PlayTrain can be slower than the original.

QuickJS steps them 13.4 faster than Node/V8 and 117 faster than the browserFound in the source text, word for word.

Picked because: Provides PlayTrain, an open‑source RL framework for generating adaptable JavaScript games with LLMs, useful for full‑stack prototyping and self‑hosted deployment.