Arnaud Deza, Santanu Dey, Pascal Van Hentenryck · abstract · pdf
quote unverifiedfigures checkedread: full textmath.OC
Problem
Previous GPU LP solvers kept the matrix or primal‑dual state on a single node, so phases other than matrix‑vector multiplication re‑introduced a single‑node memory limit; simply moving the matrix‑vector product to GPUs does not solve the overall memory bottleneck.
Approach
ShardLP distributes both the constraint matrix and the primal‑dual vectors across GPUs from input sharding through output, maintaining persistent ownership. It implements the PDHG (Chambolle‑Pock) iteration with diagonal scaling, presolve, adaptive step sizes, restart, and feasibility polishing on the distributed data. Support‑aware communication skips GPUs that store no coefficients for a given row in column‑partitioned solves. The system runs one MPI rank per GPU and uses collective reductions only for scalars or partition‑sized data, never gathering the full matrix or solution.
Result
ShardLP meets the published PDLP acceptance criterion on nine of eleven Google benchmark instances, exceeding the CPU baseline which succeeded on eight. On the largest instance it solves in 593 s of solver time (933 s end‑to‑end), compared with 21.06 h on the CPU. Scaled runs solve a 13.604 billion‑variable, 40.807 billion‑nonzero LP in 1,905 s on 76 GPUs, and support‑aware communication cuts communication by 92.97% while improving runtime by up to 1.52×.
Why it matters
Researchers and practitioners needing to solve extremely large sparse LPs that exceed single‑node memory will benefit from ShardLP's distributed memory ownership and communication optimizations.
Method details
Largest benchmark: 1.185 billion variables and 6.338 billion nonzeros.
Hardware: 8 H200 GPUs on a single node for the Google PDLP benchmark; up to 76 GPUs across 29 nodes for larger solves.
Baselines: published CPU PDLP study (Applegate et al., 2026) and prior GPU systems D‑PDLP and MPAX.
Ablation: support‑aware communication reduces modeled data movement by 92.97% and yields solver‑time speedups of 1.27×‑1.52×.
Numbers
1.185 billion variables, 6.338 billion nonzeros, 593 s solver time vs 21.06 h CPU
9 of 11 instances solved vs 8 of 11 for CPU baseline
13.604 billion variables, 40.807 billion nonzeros solved in 1,905 s on 76 GPUs
Support‑aware communication cuts modeled communication by 92.97%
Solver‑time speedup from support‑aware communication: 1.27×‑1.52×
Limitations
ShardLP requires the LP to be pre‑sharded offline and currently supports support‑aware communication only for column partitions with fixed matrix support.
The model could not quote the source for its claim, so nothing above has been checked against the paper. Read the abstract before trusting it.Citation check failed.
Picked because: Introduces SHARDLP, a released distributed GPU linear programming solver that demonstrates practical techniques for scaling compute workloads across clusters.
Leitian Tao, Baolin Peng, Haorui Wang and 7 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Execution feedback only helps coding agents when the tests fully capture the requested behavior, but agent‑generated tests are often incomplete or incorrect, and when the same trajectory creates both the patch and the test their errors can align, giving false confidence.
Approach
ExecCritic separates testing from repair by introducing two distinct agents that share the same Qwen‑3.5‑35B‑A3B backbone. The Test agent independently generates a repository‑native test bundle (code, execution command, JSON contract) which the harness validates and then freezes. The Repair agent then revises the source code using the fixed test’s execution feedback without altering the test. Training occurs in two stages: Learn to Test (SFT then RL to produce behaviorally valid tests) and Test to Improve (feedback‑conditioned repair training). This role‑specific reinforcement learning allows each agent to be optimized for its own objective while remaining compatible at inference time.
Result
Post‑training, the Qwen Test agent’s Base‑to‑Gold success rises from 22.2% to 62.2%, and when paired with the post‑trained Qwen Repair agent the overall resolved rate reaches 72.6%, an 11.4‑point improvement over the original no‑test baseline of 61.2%. GPT‑5.6‑sol generated tests raise the baseline to 65.3% while Qwen‑generated tests drop it to 57.3%.
Why it matters
Researchers and engineers building coding assistants should care because separating test generation from repair and training them with role‑specific RL markedly improves feedback‑driven patch resolution without requiring larger models.
Method details
Backbone model for both agents: Qwen‑3.5‑35B‑A3B.
Training data: SWE‑ReBench for both Test and Repair agents.
Test‑agent training: SFT on DeepSeek‑V4‑Flash‑0731 trajectories then 200 GRPO RL steps with dynamic sampling (16 issue groups, 8 rollouts per group, up to 5 test attempts).
Repair‑agent training: 100 RL steps, same dynamic sampling, up to 5 feedback‑guided revision rounds of 40 turns each.
Baselines compared: untrained Base‑35B, its SFT checkpoint, Codex‑5.3, DeepSeek‑V4‑Flash‑0731, and GPT‑5.6‑sol.
Ablations include: standard repair vs. Test‑to‑Improve with/without direct‑solve bonus (Table 3) and SFT teacher comparison (Table 4).
Numbers
Base‑to‑Gold success after RL, 62.2%, compared to SFT baseline 39.6%
Composed Qwen agents resolved rate, 72.6%, an 11.4‑point gain over no‑test baseline 61.2%
Qwen‑generated tests reduce no‑test baseline, 57.3%, compared to baseline 61.2%
Full Test‑to‑Improve final resolved rate, 72.6%, versus Standard training final 70.3%
Cross‑language Rust test‑generation success, 66.7%, compared to Python 62.2%
Limitations
The paper does not demonstrate that generated tests can fully replace privileged Oracle fail‑to‑pass feedback, and it shows limited improvement on the SWE‑bench Pro subset where test coverage is insufficient.
The trained Test agent therefore reaches a level comparable to Codex-5.3 at , while remaining points below DeepSeek-V4-Flash-0731 and points below GPT-5.6-sol.Found in the source text, word for word.
Picked because: Presents ExecCritic, a concrete test‑verify‑revise pipeline for coding agents that engineers can adopt to improve automated code repair reliability.
Min Zeng, Yuzhou Liu, Zhenyu Cao and 5 others · abstract · pdf
quote verified6 figures not in sourceread: full textcs.CL
Problem
Existing synthetic tool-use data pipelines use a generate‑then‑filter paradigm with static post‑hoc verification, which leads to inefficient data and imbalanced feature distributions; simply adding more filtering does not fix the issue because static filters only discard bad samples without correcting underlying generation errors.
Approach
ToolLoop introduces a closed‑loop framework that decomposes data synthesis into three progressive stages-ground truth sampling, backward derivation of user queries, and forward derivation of tool calls. Each stage incorporates dynamic self‑feedback that verifies and refines the model's output before moving to the next stage. This transforms the pipeline from generate‑then‑filter into a generate‑verify‑refine loop. The feedback mechanism iteratively guides the model toward high‑quality generation, ensuring coherence, parallelizability, and correct argument grounding. The overall system is applied to a 4B‑parameter base model and trained on a small synthetic dataset.
Result
ToolLoop‑4B achieves 86.40% overall accuracy on BFCL in non‑reasoning mode, with 91.29% on non‑live scenarios and 81.50% on live scenarios, while the isolate variant reaches 86.07% overall. On ACEBench, ToolLoop‑4B attains 72.1% overall accuracy using only 18.3% of the training data required by APIGen‑4B.
Why it matters
Researchers and engineers building tool‑use language models should care because ToolLoop delivers higher accuracy with far fewer synthetic examples, enabling more data‑efficient fine‑tuning for tool‑calling tasks.
Method details
Base model: Qwen3-4B-Instruct-2507 (4 B parameters)
Synthetic training set: 11 K examples for ToolLoop‑4B, 10 K for the isolate variant
Benchmarks: Berkeley Function Calling Leaderboard (BFCL‑v4, 2 501 test instances) and ACEBench
Baselines compared: APIGen‑4B (60 K examples), ToolMind‑4B (55 K examples), plus several commercial and open‑source models
Ablation variants: w/o Feedback (79.97% overall), w/ Final Filtering (82.56% overall), full ToolLoop (86.40% overall)
Numbers
86.40% overall accuracy on BFCL (ToolLoop‑4B) vs 83.11% for APIGen‑4B (3.29‑point gain)
86.07% overall accuracy for ToolLoop‑4B‑Isolate (0.33 points below full model)
91.29% non‑live BFCL accuracy (ToolLoop‑4B) vs 89.90% for APIGen‑4B (1.39‑point gain)
81.50% live BFCL accuracy (ToolLoop‑4B) vs 76.31% for APIGen‑4B (5.19‑point gain)
72.1% overall ACEBench accuracy (ToolLoop‑4B) vs 67.0% for APIGen‑4B (5.1‑point gain) using 18.3% of the training data
79.97% overall accuracy without feedback (ablation) vs 86.40% with full ToolLoop
Limitations
The paper shows limited improvement on the Profile category of ACEBench, where all fine‑tuned models underperform the base model, and it still trails the best single‑turn performance (ToolMind‑4B).
a 4B parameter model trained with our 11K synthetic examples achieves 86.40% accuracyFound in the source text, word for word.
These figures do not appear in the source text: 10 K, 11 K, 4 B, 5.19, 55 K, 60 K. Treat them as unverified.Number check failed.
Picked because: Describes ToolLoop, a closed‑loop framework for synthesizing high‑quality tool‑use data, enabling engineers to build and evaluate LLM‑driven tool agents.
Zhou Yu, Bin Bi, Shiva Kumar Pentyala and 8 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Imitating expert trajectories under the evolved harness caused the weaker model to regress on all seven tasks, because the weaker model adopted the expert's planning strategy without the competence to execute it, breaking the fit with the harness.
Approach
The paper introduces an on-policy expert-correction pipeline driven by a meta-level MLE agent. The agent first generates rollouts of the weaker model under the evolved harness, then localizes the single failing turn in each rollout. The expert model rewrites only that turn, preserving the weaker model's original planning distribution. The corrected trajectories are used to LoRA‑SFT the weaker model. This method keeps the model‑harness fit while injecting expert knowledge.
Result
Evolving the harness raised mean test success from 29.2% to 78.0% (+48.8 points). The expert model improved from 84.4% to 93.6% (+9.2) on the evolved harness. Full imitation caused an average regression of 14.9 points, while on‑policy correction increased mean success to 79.7% (+1.7) and yielded gains on five of seven tasks.
Why it matters
Practitioners aiming to improve smaller LLMs on domain‑specific enterprise tasks can gain modest performance boosts without costly full model training by combining harness evolution with on‑policy expert correction.
Method details
Weaker model: Qwen3‑Coder‑30b‑a3b‑instruct (30B total / 3B active).
Stronger expert model: gemini‑3.1‑pro‑preview (used for correction and as teaching signal).
Additional model family tested: Gemma‑4‑26b‑a4b‑it (26B/4B active).
Ryan Truong, Lance Ying, Samuel J. Gershman and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
Prior to this work, creating new video-game environments or adapting existing ones required extensive hand‑coding, making the process laborious. The obvious fix of manually editing code does not scale because each game must be individually programmed and integrated into the RL pipeline.
Approach
PlayTrain first uses an LLM‑prompting pipeline to generate or modify a JavaScript game from a short human description. The generated JS file is then compiled ahead of time into a shared library and executed inside a QuickJS interpreter embedded in the trainer process. QuickJS is coupled with a custom C++ p5 layer that forwards drawing commands to a Rust rasterizer which writes directly into the observation buffer. The environment runs as a standard OpenAI‑gym instance, allowing any RL algorithm to interact with it. Double‑buffered and single‑buffered execution modes are supported to compare throughput against baseline C++ environments.
Result
PlayTrain achieves over 1M agent‑steps per second on a single GPU node, with double‑buffered IMPALA‑Nature‑CNN reaching 1.07M steps/s across all 24 games. Per‑core, PlayTrain is 12.62× faster than ALE and 2.19× faster than ProcGen, while QuickJS runs 13.4× faster than Node/V8 and 117× faster than a headless browser. In multi‑threaded training, PlayTrain clones are on average 2.25× faster than ProcGen and 5.80× faster than ALE.
Why it matters
Researchers and engineers building RL benchmarks will benefit from rapid, LLM‑driven generation of high‑performance JavaScript environments without hand‑coding, enabling faster experimentation.
Method details
Uses Nature‑CNN and IMPALA‑CNN vision encoders
Trains PPO and IMPALA algorithms
Clones 8 Atari (ALE) and 16 ProcGen games for a total of 24 environments
Baseline comparison against original C++ EnvPool implementations
Evaluates double‑buffered versus single‑buffered execution
Measures per‑core and multi‑thread throughput on a node with four H100 GPUs and 92 CPU cores
Numbers
agent‑steps/s, 1.07M, double‑buffered IMPALA with Nature‑CNN across all 24 games
agent‑steps/s, 0.35M, double‑buffered IMPALA‑CNN across all 24 games
speedup, 12.62×, PlayTrain vs ALE per‑core geometric average
speedup, 2.19×, PlayTrain vs ProcGen per‑core geometric average
backend speedup, 13.4×, QuickJS vs Node/V8
backend speedup, 117×, QuickJS vs browser (Playwright)
Limitations
The paper does not demonstrate performance on games whose logic runs per frame rather than drawing commands, where PlayTrain can be slower than the original.
QuickJS steps them 13.4 faster than Node/V8 and 117 faster than the browserFound in the source text, word for word.
Picked because: Provides PlayTrain, an open‑source RL framework for generating adaptable JavaScript games with LLMs, useful for full‑stack prototyping and self‑hosted deployment.