Bo Liu, Simon Yu, Yiding Jiang and 15 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Existing training environment pools are hand‑curated, statically synthesized, or frozen‑verifier, keeping the goal distribution fixed as the learner scales, which limits diversity and continual improvement.
Approach
SPADE frames self‑play as a single LLM assuming two roles: an Environment Designer that writes complete Gym‑style executable environments and a Reasoning Agent that learns to act in them. The Designer is conditioned on freshly sampled corpus documents and an accumulated memory of past environments with regret scores. Regret is estimated as the reward gap between runs with and without privileged hints, guiding the Designer to generate environments at the edge of the agent's capabilities. The two roles are trained jointly so that the curriculum co‑evolves with the agent. Hint‑based regret is blended with a difficulty anchor to form the Designer's reward.
Result
Across eight held‑out benchmarks SPADE gains an average of +5.3 points over the strongest fixed‑environment baseline, lifts tool‑use performance by +5.7 on BFCL‑v4 multi‑turn and +13.9 on ACEBench‑Agent, and shows larger margins as model scale increases (e.g., +5.2, +5.7, +8.1 average points for 4B, 8B, and 30B backbones respectively).
Why it matters
Researchers building self‑improving language agents and RL‑based curriculum generators should care because SPADE demonstrates that jointly learning environment design and reasoning can substantially boost performance across diverse tasks.
Method details
Model sizes up to 30B parameters (Qwen3‑30B‑A3B‑Instruct‑2507)
Environment Designer and Reasoning Agent are LLMs using an OpenAI‑Gym‑style reset()/step() interface
Corpus grounding uses k documents from DCLM, MegaScience for games and Nemotron code corpus for tool‑use
Baseline comparisons include Fixed‑env GRPO and Fixed‑env RLVE trained on the same base models
Ablations test freezing the Designer, removing memory, and using EMA‑based learning‑potential reward
Numbers
Avg across eight benchmarks, +5.3, vs strongest fixed‑environment baseline
BFCL‑v4 multi‑turn, +5.7, vs baseline
ACEBench‑Agent, +13.9, vs baseline
Qwen3‑4B avg, +5.2, vs Fixed‑env RLVE
Qwen3‑8B avg, +5.7, vs Fixed‑env RLVE
Qwen3‑30B avg, +8.1, vs Fixed‑env RLVE
Limitations
The paper does not establish theoretical guarantees of open‑ended self‑improvement beyond the reported empirical gains.
SPADE improves over the strongest fixed‑environment baseline by +5.3 on average across eight held‑out math, science, code, and reasoning benchmarksFound in the source text, word for word.
Picked because: Introduces SPADE, a self‑play RL framework that lets a single LLM generate and adapt its own training goals, providing concrete tooling for building and verifying autonomous language agents.
Ramneet Kaur, Pradyumna Chari, Ramesh Raskar and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Supervised probes that classify activations as collusive require labeled attack examples and fail to detect future unseen latent codes, making monitoring of private latent communication brittle.
Approach
The paper introduces Verifiable Latent Alignments (VLA), which links each private latent handoff to the resulting public action via a shared event identifier, enabling causal analysis. VLA employs a three‑layer unsupervised monitor: Layer 1 detects anomalies in a neutral latent manifold, Layer 2 measures counterfactual influence on the action distribution, and Layer 3 uses a sparse auto‑encoder for interpretation. A steerability framework adds black‑box prompt steering and white‑box matched‑neutral counterfactual steering. The monitor is trained only on neutral interactions and applied sequentially to detect covert coordination.
Result
The sequential monitor attains a mean AUROC of 0.993 for homogeneous agents and 0.854 for heterogeneous pairs when pooling text‑ and latent‑collusion rows as positives. Monitoring incurs only a small normalized load relative to all possible directed pairs, and full white‑box steering recovers the bid distribution perfectly while cutting collusive low‑bid behavior by 47.3 percentage points.
Why it matters
Platform operators and governance bodies should care because VLA offers a way to audit and intervene on hidden latent coordination without needing labeled attack data, improving safety of multi‑agent LLM systems.
Method details
Qwen3-0.6B, Mistral-7B, Vicuna-7B, and Phi-4-mini model families are used for homogeneous and heterogeneous experiments
The NandaTown auction benchmark with 25 to 100 bidders provides the evaluation dataset
Layers 1 and 3 of the monitor are fit on neutral latent rows; Layer 2 uses matched counterfactual queries at test time
Baseline comparison includes homogeneous vs heterogeneous AUROC without training on attack examples
Ablation reports AUROC of 0.993 for homogeneous agents and 0.854 for heterogeneous pairs
Numbers
AUROC, 0.993, homogeneous agents
AUROC, 0.854, heterogeneous pairs
Bid‑distribution recovery, 100%, full white‑box steering
Collusive low‑bid reduction, 47.3 percentage points, full white‑box steering
Bidder count, 25 to 100, Qwen3-0.6B auctions
Limitations
The paper does not claim to detect covert coordination without access to matched neutral counterfactuals or to scale to arbitrary activation capture across all agent pairs.
full whitebox steering achieves 100% bid-distribution recovery and reduces collusive low-bid behavior by 47.3 percentage pointsFound in the source text, word for word.
Picked because: Presents Verifiable Latent Alignments, a concrete system for detecting and steering hidden latent‑state communication in LLM agents, directly addressing covert coordination verification.
Tate Berenbaum, Muthaiah Venkatachalam · abstract · pdf
quote verifiedfigures checkedread: full textcs.DC
Problem
Standard export pipelines (torch.export, torch.onnx.export) fail on reshape and transpose ops in rotary position embeddings, and a naive per‑stage export runs well below monolithic inference because it misses an OpenVINO GPU optimization.
Approach
The method splits a LLM by layers into per‑stage shards, each compiled with torch.jit.trace and pre‑computed rotary embeddings, producing independent OpenVINO graphs. A beam_idx Gather operation is injected into each shard to trigger the IndirectKVCache fusion GPU optimization, restoring parity with the unsplit model. Speculative decoding is applied on stateful OpenVINO models using a small draft model. Multiple user requests are interleaved across stages via micro‑batching, giving each request its own KV cache. The system runs across two or more Intel AI PCs connected by a standard network.
Result
A two‑node Llama 3.1 8B INT4 pipeline serves two concurrent users at 1.79× the single‑user throughput of the unsplit model. Speculative decoding yields a mean speedup of 1.33× (29.98 vs 22.66 tok/s). Micro‑batching with two streams raises aggregate throughput from 16.33 to 29.34 tok/s, a 1.80× gain. The Gemma 4 v2 single‑node export reaches 13.98 tok/s.
Why it matters
Engineers deploying large language models on Intel AI PCs with limited memory can use this pipeline to run models that exceed a single device’s capacity while maintaining or improving throughput.
Method details
Llama 3.1 8B INT4 model split across two nodes (Panther Lake Arc B390 iGPU).
Speculative decoding uses Llama 3.2 1B INT4 as draft model.
Gemma 4 E2B 5.1B FP32 model evaluated on a four‑node Lunar Lake deployment.
Baseline A: monolithic openvino_genai C++ decode loop (24.54 tok/s).
Ablation B (pre‑beam_idx) runs at 21.26 tok/s, B (_beam) runs at 24.45 tok/s.
Numbers
Tok/s 24.54 vs 1.00 baseline (mono openvino_genai).
Tok/s 24.45 vs 1.00 baseline (B _beam).
Mean speedup 1.33× (29.98 tok/s vs 22.66 tok/s).
Two‑node pipeline 1.79× single‑user throughput.
Gemma 4 v2 single‑node 13.98 tok/s.
Micro‑batching gain 1.80× (16.33 to 29.34 tok/s).
Limitations
The paper does not provide measurements for the v2_beam configuration in multi‑node or micro‑batch settings, and it does not evaluate continuous batching or models larger than 70 B parameters.
The beam_idx Gather injection recovers of throughput (from to tok/s) without touching the weights, tokenization, or decode loop.Found in the source text, word for word.
Picked because: Describes a released pipeline‑parallel inference system for 70B‑parameter LLMs on Intel AI PC fleets, offering practical artifacts for self‑hosted, cloud‑scale deployment.
Huan-ang Gao, Haohan Chi, Yong Yan and 7 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
Standard multi-teacher on-policy distillation (M-OPD) captures only 35.6% of the available headroom and suffers severe degradation on concise tasks because the token-level optimization budget is severely misallocated, not because of gradient conflict.
Approach
Open-MOPD keeps the teacher routing and on-policy distillation objective but introduces three orthogonal mechanisms: token-share balancing equalises domain token shares in the loss, gap-following allocation scales domain loss weights by a clipped factor derived from running mean reward magnitudes, and reward refresh recomputes the student-dependent reward at each inner PPO update while reusing cached teacher log‑probabilities. These mechanisms operate on batch, training‑progress, and inner‑update time scales respectively and are combined into a single training loop.
Result
Open-MOPD raises the integration‑gap recovery from 35.6% to 83.4%, increasing the total score to 31.24 compared with 28.05 for Naive M-OPD, and improves per‑domain averages (Math 22.42 vs 21.26, Code 21.73 vs 19.26, IF 49.58 vs 43.64).
Why it matters
Researchers building multi‑teacher RL distillation pipelines should adopt Open-MOPD to eliminate token‑budget imbalance and achieve near‑oracle performance without excessive compute.
Method details
Base model: SmolLM3-3B-Base, a 3B decoder‑only model with 64K context length.
Datasets: AIME24 and AIME25 (math), LiveCodeBench v5/v6 (code), IFEval and IFBenchtest (instruction).
Training: four epochs of mixed‑domain SFT followed by PPO with inner‑step count as in Algorithm 1, run on a single 8 A100‑80GB node.
Ablations: individual evaluation of token‑share balancing, gap‑following allocation, and reward refresh.
Hardware constraint: recipe must run on a single 8 A100‑80GB node.
Numbers
headroom recovery, 35.6%, Naive M-OPD
headroom recovery, 83.4%, Open-MOPD
total score, 28.05, Naive M-OPD
total score, 31.24, Open-MOPD
Math avg, 21.26, Naive M-OPD
Math avg, 22.42, Open-MOPD
Code avg, 19.26, Naive M-OPD
Code avg, 21.73, Open-MOPD
IF avg, 43.64, Naive M-OPD
IF avg, 49.58, Open-MOPD
truncation rate, 15.2%, SmolLM3-3B-Base
Limitations
The paper does not establish how the method scales beyond the 3B model or to domains beyond the three evaluated.
elevating headroom recovery from 35.6% to 83.4% in a single deployable student.Found in the source text, word for word.
Picked because: Offers an open‑source diagnostic and remediation recipe for multi‑teacher on‑policy distillation, enabling engineers to balance capabilities when consolidating RL experts.
Joy Jia Yin Lim, Xin Huang, Hao Peng and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Frontier LLM agents could execute post‑training pipelines but their training strategy was locked in at the start, so they could not revise high‑level plans as evidence accumulated; simply adding more experience, guidance, or compute did not unlock strategy‑level capability.
Approach
The authors first analyze a large corpus of public post‑training trajectories to confirm that agents lock into a default strategy early. They then intervene with three escalating mechanisms: an experience‑driven scaffold that supplies a skill library and evaluator agent, controlled human guidance that revises the initial plan, and additional inference compute during execution. Each intervention builds on the Claude Code agent with Opus 4.6 and is evaluated on GSM8K, HumanEval, and AIME 2025 under a fixed compute budget. The study measures whether these fixes improve execution performance and whether they induce strategy changes.
Result
The experience‑driven scaffold raised execution scores by +12.6 points on GSM8K and +40.8 points on HumanEval but did not change the locked‑in strategy; human guidance successfully shifted the initial strategy but the agent reverted to local adjustments once training began; extra inference compute helped on easier tasks but yielded almost no gain on the hardest benchmark.
Why it matters
Researchers building automated AI‑for‑AI systems should note that improving execution resources alone will not overcome the strategic bottleneck in post‑training pipelines.
Method details
Base model Qwen3-1.7B-Base is used for the main experiments.
Benchmarks: GSM8K, HumanEval, and AIME 2025.
Each trajectory runs under a 10‑hour compute budget on one NVIDIA H100 80GB GPU.
Intervention experiments run three times on four NVIDIA A800 GPUs.
Agents compared: Claude Code (Opus 4.6), Claude Code (GLM‑5.2), and Codex CLI (GPT‑5.2).
Dataset includes 20 distinct agent configurations across seven benchmarks and four base models ranging from 1.7B to 4B parameters.
Numbers
+12.6 points on GSM8K
+40.8 on HumanEval
10‑hour compute budget
four NVIDIA A800 GPUs
three runs
seven benchmarks
Limitations
The paper does not provide a mechanism that enables agents to spontaneously revise their strategy during execution.
experience-driven scaffold improves execution across the board (+12.6 points on GSM8K and +40.8 on HumanEval) but leaves the strategy staticFound in the source text, word for word.
Picked because: Provides an empirical study of post‑training LLM agents that separates execution‑level from strategy‑level capabilities, yielding actionable insights for building reliable agent toolchains.