Benhao Huang, Chufan Shi, Junlin Chen and 4 others · abstract · pdf
quote verified2 figures not in sourceread: full textcs.LG
Problem
Fixed-depth training breaks KV sharing and Huginn’s broad depth prior dilutes supervision at the target depth. Existing injection schemes let the state’s component along the input amplify or cancel the injection.
Approach
The paper learns a depth prior from prediction feedback and adds an entropy term to keep it broad. It introduces orthogonal injection that projects the carryover off the input direction. Training uses truncated backpropagation through the last recurrences, which shapes fixed points without warm‑up. At inference, rollout states are saved and reused for RL updates, avoiding a second forward pass. Together these components enable KV sharing, faster distillation, and faster RL updates.
Result
The learned prior and orthogonal injection lower perplexity at every scale relative to Huginn’s prior and existing injection schemes. A distilled student achieves 1.79× faster prefilling, and RL updates are 2× faster than full backpropagation. The learned prior with a 3× smaller KV cache matches the downstream average of fixed‑depth training with the full cache.
Why it matters
Researchers building looped language models and RL pipelines should care because the methods reduce training and inference cost while preserving accuracy.
Method details
Model scales from 100M to 1.6B parameters using a looped Huginn architecture with prelude, recurrent core, and coda blocks.
Ablations: learned depth prior vs Huginn prior; orthogonal injection vs other injection recipes; KV cache size reduced by 3×.
Numbers
prefill speedup, 1.79x faster compared to baseline
RL update speedup, 2x faster compared to full BPTT
KV cache reduction, 3x smaller matches full cache performance
parameter count range, 100M to 1.6B
non‑emb parameters S scale, 0.104 B
total parameters L scale, 2.451 B
Limitations
The paper does not establish results beyond perplexity or for models larger than 1.6B parameters.
a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the replayed trajectory.Found in the source text, word for word.
These figures do not appear in the source text: 0.104 B, 2.451 B. Treat them as unverified.Number check failed.
Picked because: Shows concrete engineering tricks (truncated backprop, KV sharing, fast prefilling) that cut training and inference cost for looped language models, with released code.
Haozhen Zhang, Haodong Yue, Quanyu Long and 6 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Existing agent memory systems build query-agnostic memory, incurring unnecessary preprocessing and discarding details needed later; recent runtime‑adaptive approaches specialize in fixed operations or schemes, leaving flexible control over quality, cost, and latency unexplored.
Approach
MemPilot introduces an orchestrator that iteratively decides between retrieving compressed, query‑agnostic memory and delegating query‑specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The orchestrator controls evidence amount, curation instructions, model selection, and visual access. Policy optimization uses reinforcement learning with Group Relative Policy Optimization, objective‑wise advantage decoupling, and prefix‑based marginal utility estimation. This enables fine‑grained allocation of runtime computation to meet different performance‑cost‑latency preferences.
Result
Across five multimodal benchmarks MemPilot variants achieve the highest F1 and L‑J scores while keeping inference cost low. For the Qwen3‑VL‑4B‑Instruct model, MemPilot‑Bal reaches 57.26 F1 and 64.27 L‑J on Mem‑Gallery with a cost of 1.4e‑2, outperforming all baselines. Similar quality‑cost advantages are observed for the Qwen3.5‑9B model, where MemPilot‑Bal attains 53.55 F1 and 67.27 L‑J with 2.1e‑2 cost.
Why it matters
Developers of LLM agents who need to balance answer quality with inference cost and latency should consider MemPilot for flexible, on‑demand memory curation.
Method details
Answer models: Qwen3-VL-4B-Instruct and Qwen3.5-9B.
Training: policy optimized with Group Relative Policy Optimization (GRPO) using objective‑wise advantage decoupling and prefix‑based marginal utility.
Heterogeneous model pool includes Llama‑3.2‑3B‑Instruct, Llama‑3.1‑8B‑Instruct, Llama‑3.3‑70B‑Instruct, Qwen3‑30B‑A3B‑Instruct‑2507, Qwen3‑Next‑80B‑A3B‑Instruct, Qwen3‑235B‑A22B‑2507, Qwen3‑VL‑8B‑Instruct, Qwen3‑VL‑30B‑A3B‑Instruct, Qwen3‑VL‑235B‑A22B‑Instruct, Gemini‑2.5‑Flash‑Lite, Claude‑Haiku‑4.5.
Numbers
MemPilot‑Bal F1 57.26 vs next best 54.24 (A‑Mem) on Mem‑Gallery
MemPilot‑Bal L‑J 64.27 vs next best 51.45 (A‑Mem) on Mem‑Gallery
MemPilot‑Bal Cost 1.4e‑2 vs next best 2.8e‑2 (A‑Mem) on Mem‑Gallery
MemPilot‑Bal (Qwen3.5‑9B) F1 53.55 vs next best 48.19 (MemPilot‑Perf) on Mem‑Gallery
Limitations
The paper does not discuss any limitations.
MemPilot spans distinct quality-cost regimes across both models and benchmarks.Found in the source text, word for word.
Picked because: Introduces MemPilot, a practical toolkit for on‑demand multimodal memory management in LLM agents, enabling engineers to integrate adaptive memory without heavy preprocessing.
Yifan Zhang, Yutong Dai, Viraj Prabhu and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Training web agents with reinforcement learning relies on binary task success, which is too sparse for credit assignment, and calling frontier language model judges at every step is too expensive and unavailable at deployment.
Approach
CLIFT introduces a conformal self‑verification loop. During training the agent answers natural‑language verification questions about its rollouts; a Compositional Conformal Certifier keeps only those question signals whose URL‑conditional evidence agrees with a training‑time judge, assigns signed trust weights via polarity‑aware lift, and blends the verifier score (clipped at zero) into per‑step rewards. At test time the certified question bank is frozen and used for Conformal Trajectory Selection (CTS): the agent generates a greedy rollout and diverse retries, the self‑verifier summarizes each URL trace, and a conservative majority‑vote rule decides whether to swap to a better trajectory without any external judge.
Result
CLIFT attains state‑of‑the‑art performance on WebArena Infinity among open‑source agents, transfers a certified bank to GPT‑5.5 for VisualWebArena achieving state‑of‑the‑art under the canonical harness, and improves a live‑web agent in zero‑shot evaluation on Online Mind2Web. In the WAI ablation, CLIFT+CTS raises success from 61.8 to 74.6 while reducing average steps to 15.5.
Why it matters
Researchers and engineers building open‑source web agents should care because CLIFT provides a reusable verification signal that improves training credit assignment and enables judge‑free test‑time scaling.
Method details
Trains Gemma‑4 31B with LoRA on the WAI benchmark using the full URL‑path / host / global CCC hierarchy.
Uses GPT‑5.5 as a closed‑model backbone at test time; the model weights are not updated.
Benchmarks: WebArena Infinity (WAI), VisualWebArena (VWA), and Online Mind2Web (OM2W).
Baselines include vanilla GRPO (scalar comparative‑judge reward) and a no‑RL greedy Gemma‑4 reference.
CTS improves 5 apps, ties 4, and never regresses; average swap rate is 17.8%.
Numbers
Success rate 74.6 (CLIFT+CTS) vs 61.8 (Gemma‑4 base) on WAI
Success rate 72.6 (CLIFT) vs 64.7 (vanilla GRPO) on WAI
Certified question bank size 132 questions, 100 (75.8%) certified
CTS improves 5 apps, ties 4, regresses 0 apps
Average CTS swap rate 17.8%
Limitations
The paper does not state any limitations.
CLIFT achieves state-of-the-art performance among open-source web agents.Found in the source text, word for word.
Picked because: Presents CLIFT, a self‑verification framework for web‑agent training and test‑time scaling that provides reproducible safety checks and open‑source artifacts.
Olga Tsymboi, Ramil Latypov, Aleksandr Medvedev and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Hard multi-step search fails because existing retrievers cannot plan bounded multi-round searches and often exploit lexical shortcuts; simply increasing depth or model size does not fix the evidence recall issue.
Approach
T-Search builds on Qwen3.6-35B-A3B and runs a bounded multi-round search harness that returns ranked evidence with short justifications. It is trained in two stages: round-sliced supervised fine-tuning on teacher trajectories followed by GSPO reinforcement learning on a recall reward. The system can fuse multiple rollouts to improve recall and supports interchangeable retrievers such as dense, BM25, or LLM rerankers. The downstream answer generation is left to any separate model, allowing the backend to be swapped without retraining.
Result
A single rollout of T-Search achieves an average Recall@10 of 56.0 across seven English and Russian benchmarks, 14.4 points above its base model, and three fused rollouts raise the average to 61.3, outperforming all listed baselines. Retriever robustness experiments show Recall@10 ranging from 49.38 with a small dense retriever to 62.87 with an LLM reranker.
Why it matters
Teams that need controllable evidence retrieval can adopt T-Search to boost recall while keeping their existing answer generators, and they can swap the underlying index without retraining the model.
Method details
Base model Qwen3.6-35B-A3B (35B mixture-of-experts).
Synthetic search tasks generated natively in English and Russian, with about 8k SFT questions and 2k RL questions per language.
Two-stage training: supervised fine-tuning on teacher trajectories then reinforcement learning (GSPO) with a recall reward.
Evaluation uses five benchmarks (BrowseComp-Plus, SealQA-Hard, Russian BrowseComp-Plus, TRuST, SynthComp) with a five-round cap.
Baselines include GLM-5.1, Kimi-K2.62, Qwen3.6-27B3, DeepSeek-V4-Flash4, among others.
Ablations compare single rollout vs three fused rollouts and test retriever variants (BM25, dense, LLM reranker).
Numbers
Recall@10 56.0, single rollout, average over seven benchmarks
14.4 points above base model
Recall@10 61.3, three fused rollouts, average
Retriever Recall@10 49.38 with Qwen3-Embedding-0.6B
Retriever Recall@10 62.87 with LLM reranker
TRuST contains 324 Russian questions
SynthComp contains 395 synthetic questions per language
Limitations
T-Search only returns evidence with justifications and relies on a downstream generator, so it cannot be deployed or evaluated as a standalone question‑answering system.
T-Search improves over its base model on every benchmark and achieves the highest average Recall@10 among single-rollout configurations in the T-Search harness (55.96).Found in the source text, word for word.
Picked because: Offers T‑Search, an open‑weight, plug‑and‑play agentic retriever with a playground, allowing engineers to deploy multi‑step search pipelines without retraining downstream models.
Before this work distributed beam search could not keep a single global beam when the retained set exceeded a single GPU's memory, and simply adding more GPUs caused duplicate‑state handling and race conditions that broke exact top‑B selection.
Approach
The method implements an abstract distributed pipeline that keeps states and candidates on the GPUs while the CPU only handles control and ancestry records. It performs one global reduction per key with race‑free routing, using integer scores, Hash128 equivalence, and a fully specified physical‑layout tie order to guarantee the same reduced‑key top‑B as a monolithic run. The pipeline relies on a static audit to identify cross‑buffer uniqueness and multi‑rank scatter risks. By ensuring a single global reduction and deterministic routing, the system achieves exact beam selection across many GPUs.
Result
On eight H200 GPUs the Cube4 depth‑8 run completed in 931.266 s with an effective beam width of 2,900,361,216 and a derived throughput of 74.746 million nominal pairs per second. Two T4 GPUs measuring Megaminx achieved 30.274 million logical children per second at an effective width of 82,837,504. On a separate RTX 3060 host eight‑GPU speedups were 7.487× (common profile) and 5.605× (selected stable), with weak actual‑work gains of 6.113× and 7.295× respectively.
Why it matters
Researchers building high‑throughput distributed beam search for combinatorial puzzles should care because the paper shows how to maintain exact top‑B selection across many GPUs without fitting the entire beam on a single device.
Megaminx scalar MLP model: 34,766,849 parameters, scalar output, used with MultiGPUBeamSearch.
Megaminx QMLP2RB model: 23,978,008 parameters, 24‑action output, used with MultiGPUBeamSearch.
Workload Cube4 puzzle 1000 with 24 generators, depth 8, effective beam width 2,900,361,216 on eight H200 GPUs.
Workload Megaminx puzzle depth 8, effective beam width 82,837,504 measured on two T4 GPUs.
Baseline single‑GPU RTX 3070 E2E median 3.031 s per depth 20.
Numbers
E2E time 931.266 s for Cube4 depth 8 on eight H200 GPUs
Effective beam width 2,900,361,216 records on H200 run
Derived nominal pairs rate 74.746 million pairs/s on H200
Megaminx logical children rate 30.274 million children/s on two T4 GPUs
Eight‑GPU RTX 3060 speedup 7.487× over single‑GPU baseline (common profile)
Eight‑GPU RTX 3060 weak actual‑work gain 6.113× (common profile)
Limitations
The cross‑platform records do not establish strong scaling and the static audit leaves cross‑buffer de‑duplication unresolved.
Eight H200s completed one saturated Cube4 depth at 2,900,361,216 retained records and a derived 74.746 million nominal pairs/s;Found in the source text, word for word.
Picked because: Describes a distributed beam‑search system that runs across many GPUs as a single logical beam, delivering high‑throughput inference for large‑scale generation tasks.