Previous approaches either compress prompts after seeing the downstream query or store full earlier text, but the handover writer must decide what to retain before the query is known, and naive compression fails because it lacks the query and can discard needed information.
Approach
The method builds a three‑part handover record: (1) a section that copies decisions and constraints that must remain exact, (2) a statistical part that stores task‑justified sufficient statistics when the task provides an explicit error bound, and (3) a remainder that keeps original observations whose effect cannot be captured by the statistics. The writer constructs this record within a fixed size limit, then the continuation procedure builds a prompt from the record and feeds it to the language model. A deterministic parser extracts the answer or action, and the expected loss is evaluated under the fixed model and decoding rule. This construction isolates the effects of memory, the writer, and the continuation procedure and yields explicit finite‑bit perturbation bounds for Gaussian regression and memory‑risk trade‑offs for non‑parametric regression.
Result
Theoretical analysis shows that predictive equivalence characterizes the coarsest deterministic sufficient handover and gives a fixed‑length bit requirement; Gaussian regression yields an exact finite‑dimensional handover with bounded perturbation, and non‑parametric regression provides an achievable memory‑risk relation and a distinct memory floor.
Why it matters
Researchers building LLM‑based agents that must continue tasks across session boundaries should care, because the framework specifies exactly what information must be retained and how memory limits affect achievable risk.
Method details
Uses Gaussian linear regression to obtain an exact finite‑dimensional handover and finite‑bit predictive bounds
Derives upper and lower bounds for non‑parametric regression that relate memory size to squared prediction error
Assumes a fixed tokenizer, vocabulary, and maximum response length when constructing prompts
Models the decoder as a deterministic parsing rule that extracts the task answer from the generated sequence
Limitations
The paper provides only theoretical results and does not present empirical experiments on actual language models or datasets.
Predictive equivalence characterizes the coarsest deterministic sufficient state under the exogeneity condition of Proposition 3.1Found in the source text, word for word.
Picked because: Studies session handover for large language models, providing practical methods and artifacts for preserving context across limits.
Systems concatenate many sources into one prompt, conflating source interpretation with decision aggregation, which leads to count‑scale drift; simply thresholding the sum of unnormalized weights fails because the operating point slides with the number of consulted sources and reader reliability.
Approach
The method defines a four‑field evidence tuple (hypothesis, reliability bucket, rationale, provenance) to separate interpretation from aggregation. A language model reads each source independently to produce the tuple, and a calibrated log‑likelihood‑ratio pooling step combines the tuples arithmetically. Two instantiations are presented: a small sequence encoder trained on an auxiliary task and a tree‑ensemble model with a censored survival loss. The pooling of calibrated LLRs fixes the drift without changing model architecture.
Result
The hybrid configuration that pools calibrated log‑likelihood ratios achieves 0.921 AUPRC, outperforming the hand‑crafted baseline which scores 0.805 AUPRC, and the GRU‑only encoder reaches 0.833 AUPRC, showing the benefit of the two‑stage design.
Why it matters
Practitioners building multi‑source language‑model systems or triage engines should adopt the split‑labor design to improve precision and obtain calibrated aggregation without architectural changes.
The capacity‑partition hypothesis remains unproven because the end‑to‑end neural comparison is not feature‑matched and the reader was selected on the audit set, so the results are confounded.
a small sequence encoder on an easy auxiliary objective plus a tree ensemble carrying the censored survival loss reaches 0.921 AUPRC against 0.805 for a hand‑crafted baselineFound in the source text, word for word.
Picked because: Proposes a split‑pipeline separating source interpretation from decision aggregation, offering concrete implementations for LLM‑based tool use.
Alexy Skoutnev, Kirill Acharya, Gaston Longhitano and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Prior frontier models and base agents could not reliably infer the hidden rules and goals of ARC-AGI-3 games, leading to low scores; simply adding an off‑the‑shelf harness without a world model only raised performance modestly and did not solve the underlying inference issue.
Approach
Twin consists of four parts: a problem statement, three harness routines (Validate, Explore, Plan), a checked executor, and a goal‑discovery module. A coding agent writes an executable Python world model (the twin) and validates it by replaying every logged transition. Mismatches become counterexamples that trigger model repair in the Explore phase. Planning searches for the shortest route inside the validated twin, and ExecuteChecked runs the plan step‑by‑step, halting on any mismatch. Goal discovery infers the win condition before any reward or falls back to search when needed.
Result
Twin clears 179 of 183 levels (97.8%) and is more action‑efficient than humans on 158 of those levels (88.3%); it infers the goal before any reward on 156 cleared levels (87.2%). The benchmark action‑efficiency score reaches 93.3, clearing 23 of 25 games, far above the base model alone (7.8%) and the off‑the‑shelf harness (61.1%).
Why it matters
Researchers building agents that must learn and act in unknown environments will benefit from Twin's test‑time world‑model inference, which dramatically improves performance on the ARC‑AGI‑3 benchmark.
Method details
Base model is OpenAI Codex (GPT‑5.6 Sol) driving the Twin loop
Dataset is the full public ARC‑AGI‑3 benchmark: 25 games, 183 levels
Ablation disabling the Twin harness drops score from 93.3 to 61.1 and cleared games from 23 to 13
Numbers
179/183 levels cleared (97.8%)
158/179 levels more efficient than humans (88.3%)
156/179 levels goal inferred before reward (87.2%)
Score 93.3 out of 100 (human reference 100)
2.60 billion processed tokens
91.4 hours wall‑clock inference
Limitations
Replay validation assumes deterministic, exactly representable dynamics and only certifies logged transitions; planning and goal discovery run under fixed search budgets and cannot handle latent state or continuous observations.
Twin clears 179 out of 183 levels (97.8%), and does so more efficiently than humans in 158 out of 179 levels (88.3%).Found in the source text, word for word.
Picked because: Presents Twin, a test‑time digital‑twin system that automatically builds executable world models for unknown games, with released benchmarks.
Ziyang Luo, Zhongyao Chu, Xinjie He and 4 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Frozen language models under‑use evidence in their residual stream and cannot detect insufficient input, leading to confabulation; simply adding a second clean pass fixes the detection but doubles inference cost.
Approach
YOPO writes a conditional steering probe into mid‑stack residual layers to boost answer generation, then reads a zero‑shot sufficiency direction from the same pass. Because the steering write perturbs the residual read, a small reconstruction network is trained to map the steered residual back to its pre‑steering state using mean‑squared error on paired (steered, clean) residuals. The reconstructed residual is fed to the fixed sufficiency direction, and a percentile threshold decides abstention. An optional supervised BCE boost can be stacked on the reconstruction for higher in‑domain accuracy.
Result
YOPO more than doubles three‑way accuracy from 0.375 to 0.798 on 1.5B alphaNLI and outperforms the two‑pass reference at every scale, achieving 0.798/0.830/0.893 versus 0.753/0.790/0.863. The one‑pass gate loses up to 8 AUROC points on small models but the reconstruction restores most of the gap.
Why it matters
Researchers and engineers using frozen LLMs for reasoning should adopt YOPO to obtain answer and abstention decisions in a single pass, saving inference cost while improving accuracy.
Method details
Backbone: frozen Qwen2.5 models of 1.5B, 3B, and 7B parameters.
Steering probe writes at injection layers (e.g., layers at 1.5B, read layer 24 for 3B).
three‑way accuracy 0.375->0.798 on 1.5B alphaNLI compared to frozen baseline
one‑pass AUROC 0.798/0.830/0.893 vs two‑pass 0.753/0.790/0.863 across 1.5B/3B/7B
AUROC reference 1.5B 0.962 (in‑domain) vs 0.918 (transfer)
AUROC reference 3B 0.967 (in‑domain) vs 0.922 (transfer)
AUROC reference 7B 0.985 (in‑domain) vs 0.968 (transfer)
up to 8 AUROC points loss on small models when steering write interferes
Limitations
The paper does not establish performance on tasks beyond the evaluated reasoning and QA benchmarks, and the multi‑layer reconstruction fails to converge at 7B.
three‑way accuracy more than doubles the frozen baseline (0.375->0.798 on 1.5B alphaNLI)Found in the source text, word for word.
Picked because: Shows how to answer and abstain in a single forward pass of a frozen LLM, delivering an efficient inference technique and code.
Panjing He, Mingyue Cheng, Yucong Luo and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Existing methods flatten spreadsheets into sequential strings, losing intra-sheet boundaries and inter-sheet semantics, so LLMs cannot exploit the global spatial context that human experts use. The obvious fix of using larger prompts fails because token overhead prevents modeling fine‑grained cell relations and cross‑sheet dependencies.
Approach
SheetCompass first reconstructs the raw grid into a unified hierarchical graph that encodes columns and cross‑sheet relationships as topological nodes. It then feeds this graph into a dual‑level memory system that combines static expert knowledge with dynamic reasoning experience. A collaborative multi‑agent workflow (including an Explorer and a Reflector) navigates the graph, anchors targets, generates execution scripts, and performs closed‑loop self‑reflection. The agents iteratively reason over the graph while consulting memory, allowing precise script generation and error correction. This integration preserves both local layout geometry and global semantic connections for robust spreadsheet automation.
Result
SheetCompass consistently outperforms all baselines across the three benchmarks, achieving 71.3% pass@1 on SCB, 22.0% hard restriction on SB, and 52.3% pass@1 on SheetRM with the full model. Under the GPT‑4 backbone it improves pass@1 on SCB to 63.2%, hard restriction on SB to 18.3% and pass@1 on SheetRM to 43.5%, and scaling to GPT‑5 yields additional absolute gains of 6.4% soft restriction and 6.9% hard restriction on SB. Ablation studies show that removing the hierarchical graph drops SCB pass@1 to 56.4% and SB hard restriction to 14.5%, confirming each component’s contribution.
Why it matters
Researchers building LLM‑based agents for spreadsheet automation should consider SheetCompass for its graph‑guided reasoning and memory‑driven multi‑agent design, and practitioners needing reliable, error‑corrected spreadsheet scripts can benefit from its higher success rates.
Method details
Datasets: SCB (39), SB (43), SheetRM (7)
Backbones: GPT‑5 as primary closed‑source LLM and GPT‑4o‑mini as cost‑efficient backbone
Metrics: exec@1, pass@1 for SCB and SheetRM; soft restriction and hard restriction for SB
Ablations evaluate removal of hierarchical graph, dual‑level memory, and multi‑agent workflow
Reasoning cycles hyperparameter set to 2 for default operation
Numbers
SCB Pass@1 71.3% compared to baselines
SB Hard restriction 22.0% compared to baselines
SheetRM Pass@1 52.3% compared to baselines
Hierarchical Graph removal SCB Pass@1 56.4% (drop of 14.9%) compared to full model
Dual-level Memory removal SB Hard restriction 18.8% compared to full model
Multi-agent removal SCB Pass@1 59.1% compared to full model
Limitations
The paper does not evaluate scalability to very large workbooks or generalization to spreadsheet domains beyond the three benchmarks used.
SheetCompass explicitly models structural relationships within and across worksheets while maintaining task-relevant information in memory, enabling agents to reason more effectively over complex workbooks.Found in the source text, word for word.
Picked because: Introduces SheetCompass, hierarchical relation graphs that enable LLMs to reason over complex spreadsheets, accompanied by datasets and models.
Beatrice Alessandra Motetti, Emilien Guandalino, Daniele Jahier Pagliari and 4 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Existing generative models synthesize content but often lack information grounding, leading to hallucinations, and simple retrieval does not solve the grounding issue.
Approach
Wyvern is a multi‑agent framework that sequentially runs three modules: a search module that retrieves and processes relevant resources, a report generation module that writes the text and integrates the most informative figures and an overview table, and a grounding module that revises atomic claims against the collected evidence. Each module is implemented as an ensemble of LLM agents orchestrated with LangChain. The search agents use Serper.dev and document parsers, the generation agents use DeepSeek‑V3, and the grounding agents use DeepSeek‑R1 with a claims auto‑revision stage. The final step improves citation recall and precision before the report is output.
Result
Human evaluators preferred Wyvern's figures as more informative in 87% of cases and rated its reports as more useful than the three baselines in 63% to 100% of instances. Automatic metrics show citation recall up to 73.79% (a 2.3× gain over baselines) and citation precision up to 75.45% (a 1.6× gain). The claims revision module raised citation recall from 60.92% to 73.79% and precision from 67.15% to 75.45%.
Why it matters
Researchers needing automatically generated, citation‑grounded multimodal reports should consider Wyvern for its improved factuality and figure integration.
Method details
Implemented with LangChain
Search uses Serper.dev API for top‑relevant resources
Reasoning agents use DeepSeek‑R1 and other agents use DeepSeek‑V3, both at temperature 0
Evaluator models are DeepSeek‑V3 (temp 1) and Qwen3‑32B (temp 0.6)
Baselines compared are STORM, WebThinker, and WikiAutoGen
Human study involved 27 volunteers, 23 completed questionnaires with nine pairwise comparisons per baseline
Numbers
Figures informativeness preference, 87%, over recent baseline
Usefulness preference over STORM, 100%, over WebThinker, 62.50%, over WikiAutoGen, 87.50%
Citation recall, 73.79%, versus STORM (+21.50 pp) and WikiAutoGen (+42.34 pp)
Citation precision, 75.45%, gain 1.12× over baseline
Citation recall after revision, 73.79%, up from 60.92% (gain 1.21×)
Limitations
The automatic LLM‑based evaluation still fails to capture human judgments for some baselines and exhibits ordering bias in up to 36.11% of comparisons.
Wyvern’s reports are rated as more useful than those produced by three alternative methods in 63% to 100% of instances.Found in the source text, word for word.
Picked because: Describes Wyvern, a multi‑agent framework for generating grounded multimodal reports, with open‑source components.
Yuhao Zhan, Bingxiang He, Zecong Tang and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Existing evaluations optimize agents under fixed execution conditions and never test recovery after those conditions change. Simply re‑using the source solution in the mutated target fails because agents must infer hidden physics changes and redesign the code.
Approach
The paper introduces PACE‑Bench, a simulator‑grounded benchmark of 144 source‑to‑target adaptation pairs across six physics domains. Agents receive a working source design that fails in the target and must iteratively adapt it using diagnostic sandbox feedback within a 20‑attempt budget. The Uniform Suffix lists variables that may differ without revealing the exact change. The evaluation measures physical inference and mechanism redesign. Ten self‑evolving methods are compared against a Vanilla baseline.
Result
Reflexion combined with Qwen3‑14B succeeds on only 35.9% of full‑benchmark pairs, while GPT‑5.5 solves 66.7% of the Statics subset under the full 20‑attempt budget. Vanilla performance on the Statics subset shows Pass@2 ranging from 8.3% (Qwen3‑4B) to 66.7% (GPT‑5.5) and Score@2 up to 78.1 for GPT‑5.5.
Why it matters
Researchers developing self‑evolving LLM agents and physics‑based code generation should care because PACE‑Bench reveals that current methods struggle with dynamic physics adaptation and that mechanism redesign, not just parameter inference, is the key bottleneck.
Dataset: 144 source‑to‑target pairs across six physics categories; Statics subset contains 24 pairs; Kinematics subset used for cost‑normalized results.
Interaction budget: 20 attempts per pair (five‑attempt ablation also evaluated).
Baseline: Vanilla iterative submission without additional self‑evolving mechanisms.
Ablations: five‑attempt budget comparison; revealing exact physical changes does not improve performance.
Numbers
Pass@2 35.9% for Reflexion + Qwen3‑14B on full benchmark
Pass@2 66.7% for GPT‑5.5 on Statics subset
Pass@2 37.5% for Qwen3‑14B (Vanilla) on Statics subset
Pass@2 45.8% for DeepSeek‑V4‑Pro (Vanilla) on Statics subset
Score@2 78.1 for GPT‑5.5 (Vanilla) on Statics subset
Score@2 37.5 for Qwen3‑14B (Vanilla) on Statics subset
Limitations
The benchmark remains far from saturated; even the best configurations fail on a substantial fraction of pairs and the paper does not demonstrate a method that reliably solves all adaptation cases.
simulator‑grounded reflection is more reliable than unverified self‑revisionFound in the source text, word for word.
Picked because: Provides PACE‑Bench, a benchmark suite for physics adaptation via code evolution, useful for evaluating adaptive agents.
LLM evaluations traditionally use fixed sampling budgets, testing every item the same number of times even after estimates become precise. This uniform repetition wastes compute because uncertainty varies across items and model-task combinations. Simply increasing the overall budget does not address the uneven variance and leads to unnecessary sampling.
Approach
The paper introduces optstop, a precision‑based adaptive stopping framework that treats evaluation as a sequential measurement problem. It builds on hierarchical Bayesian inference to monitor the width of posterior credible intervals for each item or group. When the interval falls below a user‑specified delta, sampling stops for that unit. A safeguard samples more cautiously as measured performance approaches zero to protect rare‑success cases. optstop integrates with the inspect_ai evaluation framework for live early stopping and can also be applied retrospectively.
Result
In an illustrative 200‑item, 10‑epoch evaluation optstop removed 57%, 97% of planned trials across nine validation settings while yielding overall conclusions equivalent to the full run. Model comparison showed the logit‑normal hierarchy had far fewer divergences than the Beta‑Binomial (120 fewer in the mid‑range, 192 fewer near boundaries, and 9,206 fewer at exact boundaries).
Why it matters
Researchers and engineers conducting large‑scale LLM evaluations should adopt optstop to allocate compute based on uncertainty rather than fixed repetition counts, reducing wasted resources while preserving statistical conclusions.
Method details
Hierarchical logit‑normal models are used for binary, ordinal, and continuous outcome pathways.
MCMC inference runs with 4 chains and 1,000 post‑warmup draws by default.
Downstream analyses use 4 chains with 2,000 posterior draws per chain (8,000 total draws) and 94% HDIs.
Experiments evaluate 200 items over 10 epochs per cell.
Benchmarks include MATH Level 5 (GPT‑3.5 Turbo), GPQA Diamond (GPT‑4o), MMLU (GPT‑4o), and WritingBench with Claude Sonnet 4.5.
Numbers
trial removal, 57%, 97%, across nine validation settings
absolute bias, 0.022, logit‑normal vs Beta‑Binomial mid‑range
CI width, 0.070, logit‑normal vs Beta‑Binomial mid‑range
divergence ratio, 120, mid‑range (Beta‑Binomial divergences divided by logit‑normal)
effective sample size, 4,600, logit‑normal at boundaries (vs 7 for Beta‑Binomial)
Limitations
The paper does not establish how optstop performs on evaluation designs that do not fit a hierarchical Bayesian framework.
it removes 57%, 97% of planned trials across nine validation settings, with overall conclusions equivalent to the full run.Found in the source text, word for word.
Picked because: Offers optstop, a Bayesian optimal‑stopping method for LLM evaluation that adaptively allocates samples, with released implementation.
Zewen Jin, Shen Fu, Zeping Duan and 8 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
MoE inference in small‑batch decoding is memory‑bound and bottlenecked by expert weight loading, and existing fixes such as post‑training weight compression or fine‑grained expert design either degrade accuracy or add extra computation and communication overhead.
Approach
DeaMoE groups experts into several departments that share most parameters, while each expert retains a few private parameters. A two‑stage routing strategy avoids redundant weight loading. The shared department backbone and expert‑specific transforms are connected by SiLU non‑linear operators at gate, up, and down interfaces. Expert matrices are initialized to the identity to start from a well‑conditioned state. The design matches the baseline in total parameters and per‑token FLOPs but reduces the amount of weight movement during decoding.
Result
DeaMoE cuts per‑step loaded weights by up to 50.9% and yields up to 1.33× end‑to‑end TPOT speedup for the 7B model on an A40 GPU, while microbenchmarks show peak speedups of 2.00× on A40 and 1.97× on H100 for DeepSeek‑V3; downstream accuracy on BoolQ improves to 62.39 versus 61.47 for the baseline.
Why it matters
Teams deploying interactive LLM services with large‑scale MoE models and small‑batch decoding should adopt DeaMoE to lower latency and increase throughput, especially on bandwidth‑limited GPUs.
Method details
7.3B total parameters for both DeaMoE‑7B and Baseline‑7B
Pre‑training on RedPajama‑v1 dataset for 110B tokens
Training uses PyTorch FSDP with SiLU activation and identity initialization for expert matrices
Inference integrated into vLLM v0.13.0 with Triton operators and CUDA Graph on NVIDIA A40 and H100 GPUs
Baselines include a vanilla MoE model (Baseline‑7B) and larger models DeepSeek‑V3 and Qwen3‑235B‑A22B for microbenchmarks
Ablations study non‑linear operator choices (SiLU vs GeLU/RMSNorm) and expert‑initialization strategies
Numbers
per‑step loaded weights reduction, 50.9%, vs vanilla MoE
end‑to‑end TPOT speedup, 1.33×, vs Baseline‑7B on A40
peak speedup DeepSeek‑V3, 2.00×, vs baseline on A40
peak speedup DeepSeek‑V3, 1.97×, vs baseline on H100
BoolQ accuracy, 62.39, DeaMoE‑7B vs 61.47 baseline
throughput improvement under 20 ms TPOT budget, 1.83×, vs baseline
Limitations
The paper shows limited or even negative gains for small‑expert configurations on high‑bandwidth hardware such as H100, indicating the method is less effective when expert weights fit in cache.
DeaMoE reduces per‑step loaded weights by up to 50.9% and achieves up to 1.33 end‑to‑end TPOT speedup for the pre‑trained 7B model on A40Found in the source text, word for word.
Picked because: Develops DeaMoE, an efficient Mixture‑of‑Experts architecture optimized for small‑batch decoding, including code and performance results.