Jiale Chen, Vage Egiazarian, Eldar Kurtić and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
KV cache memory and bandwidth grow with context length and batch size, limiting efficient long-context inference, and simple low‑bit quantization or generic transforms (identity, random orthogonal, Hadamard, OSCAR) cause large reconstruction error especially at aggressive bitwidths.
Approach
WUSH‑KV builds separate data‑aware transforms for keys and values from second‑order statistics of calibration data, folds the value transform into model weights and applies the key transform after RoPE, and then quantizes the transformed caches with clipped quantizers such as QuEST INT; the transforms are near‑optimal under mild assumptions and reduce quantization error.
Result
WUSH‑KV consistently yields the lowest perplexity among all quantized transforms at 4‑, 3‑, and 2‑bit, and improves downstream reasoning accuracy over OSCAR, especially at 2‑bit where it narrows the gap to full‑precision.
Why it matters
Researchers and engineers building long‑context LLM inference pipelines can adopt WUSH‑KV to achieve low‑bit KV cache storage with minimal quality loss, enabling cheaper memory and bandwidth usage.
Method details
Evaluated on Qwen3‑8B, Qwen3‑4B‑Thinking‑2507, and Qwen3‑32B models.
Calibration uses 128 FineWeb‑Edu sequences of length 32,768.
Perplexity measured on WikiText‑2; downstream tasks include AIME 2025, MATH‑500, GPQA Diamond, LiveCodeBench v6.
Quantizers: QuEST INT (theoretical analysis) and OSCAR‑style percentile‑clipped affine quantizer.
Calibration runs take ~12 minutes on a single NVIDIA L40S GPU and peak at 39 GiB memory.
Numbers
Perplexity 2‑bit WUSH 10.51 vs OSCAR 13.74 on Qwen3‑8B WikiText‑2
Perplexity 4‑bit WUSH 9.73 vs OSCAR 9.77 on Qwen3‑8B WikiText‑2
AIME 2025 score Qwen3‑8B WUSH 46.7% vs OSCAR 34.4%
MATH‑500 score Qwen3‑8B WUSH 92.9% vs OSCAR 89.3%
Calibration time ~12 minutes on NVIDIA L40S GPU
Peak GPU memory 39 GiB during calibration
Limitations
The paper does not provide exact theoretical guarantees for the OSCAR‑style quantizer and does not retune clipping ratios for WUSH‑KV, leaving potential headroom unexamined.
WUSH achieves the lowest perplexity among the quantized transforms at every bitwidth.Found in the source text, word for word.
Picked because: Introduces WUSH-KV, a data‑adaptive low‑bit quantization for KV‑cache memory that cuts inference memory and bandwidth, with released code useful for long‑context LLM deployment.
Prior MoE inference systems use reactive offloading and caching, waiting for router outputs before moving experts, which leads to inefficient cache use and cannot overlap data transfers with compute on tight VRAM budgets.
Approach
Mira introduces a proactive runtime that couples a lightweight per-layer predictor with a two‑level GPU cache. The predictor consumes hidden states and routing statistics to forecast which experts will be needed two layers ahead. Predicted experts are asynchronously prefetched into the STAGE region while frequently used experts reside in the HOT region. Expert FFN weights are stored offline in a custom INT8 format and dequantized on‑the‑fly during fused compute. A telemetry‑driven rebalancer continuously adjusts the HOT set based on token‑level assignment statistics. The system uses separate CUDA streams for H2D transfers and compute to overlap communication and execution.
Result
Mira reduces expert‑induced stalls and achieves a 5.71x speedup in average throughput over state‑of‑the‑art baselines, an 11.71x acceleration in time‑to‑first‑token, and a 3.84x average speedup in beam‑search inference. Accuracy remains comparable to FP16, with Arc‑E 0.871 vs 0.870 FP16. Throughput gains increase as VRAM budget shrinks, e.g., 1.49x speedup at 16 GB VRAM.
Why it matters
Researchers and engineers deploying large MoE models on memory‑constrained single GPUs should consider Mira to achieve higher throughput without sacrificing accuracy.
Method details
Model evaluated: Mixtral‑8x7B and DeepSeek‑V2‑Lite‑Chat
Dataset for inference performance: ShareGPT conversations
Quantization: custom INT8 layout for expert feed‑forward blocks
Predictor horizon: forecasts expert usage two layers ahead
Numbers
throughput speedup, 5.71x, compared against state‑of‑the‑art baselines
TTFT acceleration, 11.71x, compared against baselines
beam‑search speedup, 3.84x, compared against baselines
Arc‑E accuracy, 0.871, vs FP16 0.870
VRAM 16 GB speedup, 1.49x, Mira+P over Mira
Limitations
The paper does not evaluate multi‑GPU scaling or applicability to non‑MoE transformer architectures.
Mira reduces expert‑induced stalls. Compared against state‑of‑the‑art baselines, Mira achieves a 5.71x speedup in average throughput on a memory‑constrained GPU.Found in the source text, word for word.
Picked because: Presents Mira, a memory‑efficient MoE inference system using adaptive caching and predictive expert staging, enabling self‑hosted MoE models on single‑GPU hardware.
Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Prior LLM agents used a generic Plan+ReAct loop that generated a plan and then executed it, but this approach often failed to preserve the declared planning structure, especially for longer plans. The obvious fix of simply running the same planner without enforcing its pattern does not work because the executor does not respect the plan’s hierarchical or sequential constraints.
Approach
The paper introduces Planning-as-Routing, where an LLM first declares one of four planning modes (Predefined, Sequential, Hierarchical, Search). A deterministic router then dispatches the task to a pattern‑specific executor that is hard‑wired to follow the declared structure. Each executor implements the control flow required by its mode, ensuring structural fidelity during execution. The same backbone LLM is used for both declaration and execution unless otherwise noted. Few‑shot examples can be provided to the declaration module to improve mode selection. The system is evaluated across four benchmarks with three LLM backbones.
Result
Pattern‑specific executors raise task success on ALFWorld from 0.48 to 0.92 and on SWE‑bench Verified from 0.36 to 0.44 compared with generic Plan+ReAct. Only 22‑45% of trajectories preserve the declared structure under generic Plan+ReAct, while pattern‑specific executors enforce it by design. The forced‑pattern oracle exceeds the strongest fixed pattern, but most of the remaining gap (‑0.025 to +0.027) is explained by extra execution attempts rather than per‑task pattern complementarity.
Why it matters
Researchers building LLM agents for long‑horizon tasks should adopt mode‑specific executors to ensure structural fidelity, and developers should focus on improving mode‑selection mechanisms.
Planning modes: Predefined, Sequential, Hierarchical, Search with deterministic routing to mode‑specific executors
Evaluation uses three random seeds (7, 13, 42) and reports mean and standard deviation
Baselines: generic Plan+ReAct and Flat ReAct; comparisons include Best Fixed pattern and Oracle ceiling
Plan quality judged by a Gemma-4-26B LLM‑as‑judge using a 0 to 3 rubric
Numbers
Task success on ALFWorld, 0.48, Plan+ReAct baseline
Task success on ALFWorld, 0.92, pattern‑specific executor
Task success on SWE‑bench Verified, 0.36, Plan+ReAct baseline
Task success on SWE‑bench Verified, 0.44, pattern‑specific executor
Structure preservation under generic Plan+ReAct, 22‑45%, overall trajectories
Oracle gap, -0.025 to +0.027, after three retries of strongest fixed pattern
Limitations
The study only evaluates four fixed planning patterns and does not address dynamic mode switching or broader pattern sets.
pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench VerifiedFound in the source text, word for word.
Picked because: Provides a verification framework to assess whether LLM agents actually execute the plans they generate, offering concrete metrics and tooling for reliable agent deployment.
Paras Dahal, Anton Bakhtin, Taco Cohen and 9 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Prior agents interleaved control decisions with object‑level work, making it hard to compose useful work and causing performance to plateau as compute grew; simply adding more model calls (direct control) does not solve the scaling issue.
Approach
The paper introduces agentic meta‑reasoning, separating a controller from task workers. The controller maintains a compact account of the run, persistent memory, and decides next actions via a four‑step cycle: assess, propose, evaluate, and dispatch. Workers execute the object‑level computation using the context supplied by the controller. Between decisions the controller does not replay the full history, reducing overhead. This staged control allows the system to allocate budget to promising options rather than following a fixed chain of thought.
Result
Meta‑reasoning consistently outperforms direct control and external baselines across all four benchmarks and three frontier models, achieving the highest scores at the main budget settings and showing larger gains as the budget increases, while also producing more reusable intermediate artifacts.
Why it matters
Researchers and engineers building long‑horizon language model agents should consider meta‑reasoning to better allocate inference‑time compute and improve scalability of complex tasks.
Budgets: 100 model calls for reasoning benchmarks, 1200 calls for ProgramBench, counting both controller and worker calls
Baselines compared: Direct Control Agent, Recursive Language Model, mini‑SWE Agent, Codex, Claude Code
Workers: coding agents on ProgramBench can inspect files, run commands, and test reconstructions
Numbers
ProgramBench test‑pass rate 71.5% meta‑reasoning with GPT‑5.5 vs 58.0% Codex
ProgramBench test‑pass rate 67.2% meta‑reasoning with Opus 4.8 vs 65.5% Claude Code
IMO ProofBench‑Advanced score 94.6% meta‑reasoning with GPT‑5.5 vs 93.3% Direct Control
ARC‑AGI‑2 accuracy 79.2% meta‑reasoning with GPT‑5.5 vs 75.8% Direct Control
Limitations
The staged control adds compute overhead, can underperform at low budgets, and relies on a compact state that may lose important information; performance also varies with the underlying model.
On ProgramBench with GPT-5.5, it reaches 71.5%, compared with 58.0% for Codex.Found in the source text, word for word.
Picked because: Proposes agentic meta‑reasoning, an inference‑time control layer that makes planning and execution decisions explicit, a practical method for managing LLM‑based agents in production.
Rishabh Agrawal, Hejie Cui, Shasha Li and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Using all corrections for self‑distillation can limit learning because plausible corrections that do not change execution provide weak targets that bias the advisor away from useful advice, and the obvious fix of using every correction does not work.
Approach
Advisor Self‑Distillation (AdviSD) combines outcome‑based reinforcement learning with selective self‑distillation from a feedback‑conditioned copy of the advisor. A reflection module proposes corrections; the advisor scores the same recorded executor response with and without its issued advice. The magnitude of the score difference is used as a gate to select which decisions receive distillation supervision. This selection rule avoids needing executor likelihoods or additional rollouts. The method retains only those advice decisions whose impact exceeds a calibrated threshold.
Result
AdviSD achieves the highest BFCL‑v3 accuracy and EnvScaler score for both executors, surpassing GRPO by 4.2 to 6.4 percentage points on BFCL‑v3 and by 3.9 to 5.1 score points on EnvScaler, and it also generalizes to out‑of‑domain tasks and transfers across executor versions and model families.
Why it matters
Researchers building advisory layers for frontier LLMs should care because AdviSD provides a simple, rollout‑free way to improve advice quality and transfer across models.
Method details
Trainable advisor: Qwen3-8B (8B parameters).
Frozen executors: Gemini 3.7 Flash and Claude Sonnet 4.6.
Training datasets: BFCL‑v3 and EnvScaler, with test sets of 320 BFCL tasks (80 per category) and 200 EnvScaler tasks.
Evaluation protocol: three independent training runs per method, checkpoint selected on validation, four test evaluations per run, reporting mean and sample standard deviation.
Numbers
BFCL‑v3 improvement: 4.2 to 6.4 percentage points over GRPO
EnvScaler improvement: 3.9 to 5.1 score points over GRPO
Limitations
The paper does not explicitly state any limitations.
AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler.Found in the source text, word for word.
Picked because: Describes AdviSD, a trainable advisor that steers frozen LLM executors via natural‑language advice and self‑distillation, delivering a concrete approach to improve LLM agent behavior.