arXiv digest

Monday

August 17, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 9

Handover of In-Context Learning State Across Session Boundaries

Masahiro Kato, Taka Kato · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Previous approaches either compress prompts after seeing the downstream query or store full earlier text, but the handover writer must decide what to retain before the query is known, and naive compression fails because it lacks the query and can discard needed information.

Approach

The method builds a three‑part handover record: (1) a section that copies decisions and constraints that must remain exact, (2) a statistical part that stores task‑justified sufficient statistics when the task provides an explicit error bound, and (3) a remainder that keeps original observations whose effect cannot be captured by the statistics. The writer constructs this record within a fixed size limit, then the continuation procedure builds a prompt from the record and feeds it to the language model. A deterministic parser extracts the answer or action, and the expected loss is evaluated under the fixed model and decoding rule. This construction isolates the effects of memory, the writer, and the continuation procedure and yields explicit finite‑bit perturbation bounds for Gaussian regression and memory‑risk trade‑offs for non‑parametric regression.

Result

Theoretical analysis shows that predictive equivalence characterizes the coarsest deterministic sufficient handover and gives a fixed‑length bit requirement; Gaussian regression yields an exact finite‑dimensional handover with bounded perturbation, and non‑parametric regression provides an achievable memory‑risk relation and a distinct memory floor.

Why it matters

Researchers building LLM‑based agents that must continue tasks across session boundaries should care, because the framework specifies exactly what information must be retained and how memory limits affect achievable risk.

Method details
  • Uses Gaussian linear regression to obtain an exact finite‑dimensional handover and finite‑bit predictive bounds
  • Derives upper and lower bounds for non‑parametric regression that relate memory size to squared prediction error
  • Assumes a fixed tokenizer, vocabulary, and maximum response length when constructing prompts
  • Models the decoder as a deterministic parsing rule that extracts the task answer from the generated sequence
Limitations

The paper provides only theoretical results and does not present empirical experiments on actual language models or datasets.

Predictive equivalence characterizes the coarsest deterministic sufficient state under the exogeneity condition of Proposition 3.1Found in the source text, word for word.

Picked because: Studies session handover for large language models, providing practical methods and artifacts for preserving context across limits.

Paper 2 of 9

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

Zhelun Wu · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Systems concatenate many sources into one prompt, conflating source interpretation with decision aggregation, which leads to count‑scale drift; simply thresholding the sum of unnormalized weights fails because the operating point slides with the number of consulted sources and reader reliability.

Approach

The method defines a four‑field evidence tuple (hypothesis, reliability bucket, rationale, provenance) to separate interpretation from aggregation. A language model reads each source independently to produce the tuple, and a calibrated log‑likelihood‑ratio pooling step combines the tuples arithmetically. Two instantiations are presented: a small sequence encoder trained on an auxiliary task and a tree‑ensemble model with a censored survival loss. The pooling of calibrated LLRs fixes the drift without changing model architecture.

Result

The hybrid configuration that pools calibrated log‑likelihood ratios achieves 0.921 AUPRC, outperforming the hand‑crafted baseline which scores 0.805 AUPRC, and the GRU‑only encoder reaches 0.833 AUPRC, showing the benefit of the two‑stage design.

Why it matters

Practitioners building multi‑source language‑model systems or triage engines should adopt the split‑labor design to improve precision and obtain calibrated aggregation without architectural changes.

Method details
  • Evidence tuple fields: hypothesis, reliability bucket, rationale, provenance
  • GRU cadence encoder uses 69 features and achieves 0.833 AUPRC
  • Hybrid model uses 81 features (tree ensemble) and reaches 0.921 AUPRC
  • Dataset: longitudinal corpus with 73,638 sources across 23,553 instances; train 33,840 snapshots / 10,541 instances, test 69,070 snapshots / 11,928 instances
  • Baseline hand‑crafted configuration uses 18 features and attains 0.805 AUPRC
  • Ablation includes random prevalence baseline with 0.185 AUPRC
Numbers
  • AUPRC, 0.921, hybrid model vs hand‑crafted baseline 0.805
  • AUPRC, 0.833, GRU cadence encoder vs hand‑crafted baseline 0.805
  • Precision, 0.824, independent causal reading (vs one‑shot 0.645)
  • Precision, 0.697, prior only configuration
  • Recall, 0.607, one‑shot configuration
  • Macro, 0.466, prior fallback configuration
Limitations

The capacity‑partition hypothesis remains unproven because the end‑to‑end neural comparison is not feature‑matched and the reader was selected on the audit set, so the results are confounded.

a small sequence encoder on an easy auxiliary objective plus a tree ensemble carrying the censored survival loss reaches 0.921 AUPRC against 0.805 for a hand‑crafted baselineFound in the source text, word for word.

Picked because: Proposes a split‑pipeline separating source interpretation from decision aggregation, offering concrete implementations for LLM‑based tool use.

Paper 3 of 9

Twin: Playing an Unknown Game with a Test-Time Digital Twin

Alexy Skoutnev, Kirill Acharya, Gaston Longhitano and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Prior frontier models and base agents could not reliably infer the hidden rules and goals of ARC-AGI-3 games, leading to low scores; simply adding an off‑the‑shelf harness without a world model only raised performance modestly and did not solve the underlying inference issue.

Approach

Twin consists of four parts: a problem statement, three harness routines (Validate, Explore, Plan), a checked executor, and a goal‑discovery module. A coding agent writes an executable Python world model (the twin) and validates it by replaying every logged transition. Mismatches become counterexamples that trigger model repair in the Explore phase. Planning searches for the shortest route inside the validated twin, and ExecuteChecked runs the plan step‑by‑step, halting on any mismatch. Goal discovery infers the win condition before any reward or falls back to search when needed.

Result

Twin clears 179 of 183 levels (97.8%) and is more action‑efficient than humans on 158 of those levels (88.3%); it infers the goal before any reward on 156 cleared levels (87.2%). The benchmark action‑efficiency score reaches 93.3, clearing 23 of 25 games, far above the base model alone (7.8%) and the off‑the‑shelf harness (61.1%).

Why it matters

Researchers building agents that must learn and act in unknown environments will benefit from Twin's test‑time world‑model inference, which dramatically improves performance on the ARC‑AGI‑3 benchmark.

Method details
  • Base model is OpenAI Codex (GPT‑5.6 Sol) driving the Twin loop
  • Dataset is the full public ARC‑AGI‑3 benchmark: 25 games, 183 levels
  • Compute usage: 2.60 billion processed tokens and 91.4 hours wall‑clock inference
  • Baseline scores: OPINE‑World 78.4, Prime Agent 78.3, EWM 63.8, Codex 61.1
  • Ablation disabling the Twin harness drops score from 93.3 to 61.1 and cleared games from 23 to 13
Numbers
  • 179/183 levels cleared (97.8%)
  • 158/179 levels more efficient than humans (88.3%)
  • 156/179 levels goal inferred before reward (87.2%)
  • Score 93.3 out of 100 (human reference 100)
  • 2.60 billion processed tokens
  • 91.4 hours wall‑clock inference
Limitations

Replay validation assumes deterministic, exactly representable dynamics and only certifies logged transitions; planning and goal discovery run under fixed search budgets and cannot handle latent state or continuous observations.

Twin clears 179 out of 183 levels (97.8%), and does so more efficiently than humans in 158 out of 179 levels (88.3%).Found in the source text, word for word.

Picked because: Presents Twin, a test‑time digital‑twin system that automatically builds executable world models for unknown games, with released benchmarks.

Paper 4 of 9

You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model

Ziyang Luo, Zhongyao Chu, Xinjie He and 4 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Frozen language models under‑use evidence in their residual stream and cannot detect insufficient input, leading to confabulation; simply adding a second clean pass fixes the detection but doubles inference cost.

Approach

YOPO writes a conditional steering probe into mid‑stack residual layers to boost answer generation, then reads a zero‑shot sufficiency direction from the same pass. Because the steering write perturbs the residual read, a small reconstruction network is trained to map the steered residual back to its pre‑steering state using mean‑squared error on paired (steered, clean) residuals. The reconstructed residual is fed to the fixed sufficiency direction, and a percentile threshold decides abstention. An optional supervised BCE boost can be stacked on the reconstruction for higher in‑domain accuracy.

Result

YOPO more than doubles three‑way accuracy from 0.375 to 0.798 on 1.5B alphaNLI and outperforms the two‑pass reference at every scale, achieving 0.798/0.830/0.893 versus 0.753/0.790/0.863. The one‑pass gate loses up to 8 AUROC points on small models but the reconstruction restores most of the gap.

Why it matters

Researchers and engineers using frozen LLMs for reasoning should adopt YOPO to obtain answer and abstention decisions in a single pass, saving inference cost while improving accuracy.

Method details
  • Backbone: frozen Qwen2.5 models of 1.5B, 3B, and 7B parameters.
  • Steering probe writes at injection layers (e.g., layers at 1.5B, read layer 24 for 3B).
  • Reconstruction map: single‑layer low‑rank (r=64) corrector or multi‑layer MLP reading concatenated steered residuals.
  • Training objective: mean‑squared error on (steered, clean) residual pairs, no sufficiency labels.
  • Baselines compared: frozen baseline, steering‑only, gate‑only (clean read), two‑pass clean‑read reference.
  • Evaluation datasets: alphaNLI, HellaSwag, SQuAD2, RepLiQA, MuSiQue.
Numbers
  • three‑way accuracy 0.375->0.798 on 1.5B alphaNLI compared to frozen baseline
  • one‑pass AUROC 0.798/0.830/0.893 vs two‑pass 0.753/0.790/0.863 across 1.5B/3B/7B
  • AUROC reference 1.5B 0.962 (in‑domain) vs 0.918 (transfer)
  • AUROC reference 3B 0.967 (in‑domain) vs 0.922 (transfer)
  • AUROC reference 7B 0.985 (in‑domain) vs 0.968 (transfer)
  • up to 8 AUROC points loss on small models when steering write interferes
Limitations

The paper does not establish performance on tasks beyond the evaluated reasoning and QA benchmarks, and the multi‑layer reconstruction fails to converge at 7B.

three‑way accuracy more than doubles the frozen baseline (0.375->0.798 on 1.5B alphaNLI)Found in the source text, word for word.

Picked because: Shows how to answer and abstain in a single forward pass of a frozen LLM, delivering an efficient inference technique and code.

Paper 5 of 9

SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning

Panjing He, Mingyue Cheng, Yucong Luo and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing methods flatten spreadsheets into sequential strings, losing intra-sheet boundaries and inter-sheet semantics, so LLMs cannot exploit the global spatial context that human experts use. The obvious fix of using larger prompts fails because token overhead prevents modeling fine‑grained cell relations and cross‑sheet dependencies.

Approach

SheetCompass first reconstructs the raw grid into a unified hierarchical graph that encodes columns and cross‑sheet relationships as topological nodes. It then feeds this graph into a dual‑level memory system that combines static expert knowledge with dynamic reasoning experience. A collaborative multi‑agent workflow (including an Explorer and a Reflector) navigates the graph, anchors targets, generates execution scripts, and performs closed‑loop self‑reflection. The agents iteratively reason over the graph while consulting memory, allowing precise script generation and error correction. This integration preserves both local layout geometry and global semantic connections for robust spreadsheet automation.

Result

SheetCompass consistently outperforms all baselines across the three benchmarks, achieving 71.3% pass@1 on SCB, 22.0% hard restriction on SB, and 52.3% pass@1 on SheetRM with the full model. Under the GPT‑4 backbone it improves pass@1 on SCB to 63.2%, hard restriction on SB to 18.3% and pass@1 on SheetRM to 43.5%, and scaling to GPT‑5 yields additional absolute gains of 6.4% soft restriction and 6.9% hard restriction on SB. Ablation studies show that removing the hierarchical graph drops SCB pass@1 to 56.4% and SB hard restriction to 14.5%, confirming each component’s contribution.

Why it matters

Researchers building LLM‑based agents for spreadsheet automation should consider SheetCompass for its graph‑guided reasoning and memory‑driven multi‑agent design, and practitioners needing reliable, error‑corrected spreadsheet scripts can benefit from its higher success rates.

Method details
  • Datasets: SCB (39), SB (43), SheetRM (7)
  • Backbones: GPT‑5 as primary closed‑source LLM and GPT‑4o‑mini as cost‑efficient backbone
  • Baselines: static generation methods Binder and VBA script generation; interactive LLM agents SheetCopilot, SheetAgent, OS‑Copilot
  • Metrics: exec@1, pass@1 for SCB and SheetRM; soft restriction and hard restriction for SB
  • Ablations evaluate removal of hierarchical graph, dual‑level memory, and multi‑agent workflow
  • Reasoning cycles hyperparameter set to 2 for default operation
Numbers
  • SCB Pass@1 71.3% compared to baselines
  • SB Hard restriction 22.0% compared to baselines
  • SheetRM Pass@1 52.3% compared to baselines
  • Hierarchical Graph removal SCB Pass@1 56.4% (drop of 14.9%) compared to full model
  • Dual-level Memory removal SB Hard restriction 18.8% compared to full model
  • Multi-agent removal SCB Pass@1 59.1% compared to full model
Limitations

The paper does not evaluate scalability to very large workbooks or generalization to spreadsheet domains beyond the three benchmarks used.

SheetCompass explicitly models structural relationships within and across worksheets while maintaining task-relevant information in memory, enabling agents to reason more effectively over complex workbooks.Found in the source text, word for word.

Picked because: Introduces SheetCompass, hierarchical relation graphs that enable LLMs to reason over complex spreadsheets, accompanied by datasets and models.

Paper 6 of 9

Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports

Beatrice Alessandra Motetti, Emilien Guandalino, Daniele Jahier Pagliari and 4 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing generative models synthesize content but often lack information grounding, leading to hallucinations, and simple retrieval does not solve the grounding issue.

Approach

Wyvern is a multi‑agent framework that sequentially runs three modules: a search module that retrieves and processes relevant resources, a report generation module that writes the text and integrates the most informative figures and an overview table, and a grounding module that revises atomic claims against the collected evidence. Each module is implemented as an ensemble of LLM agents orchestrated with LangChain. The search agents use Serper.dev and document parsers, the generation agents use DeepSeek‑V3, and the grounding agents use DeepSeek‑R1 with a claims auto‑revision stage. The final step improves citation recall and precision before the report is output.

Result

Human evaluators preferred Wyvern's figures as more informative in 87% of cases and rated its reports as more useful than the three baselines in 63% to 100% of instances. Automatic metrics show citation recall up to 73.79% (a 2.3× gain over baselines) and citation precision up to 75.45% (a 1.6× gain). The claims revision module raised citation recall from 60.92% to 73.79% and precision from 67.15% to 75.45%.

Why it matters

Researchers needing automatically generated, citation‑grounded multimodal reports should consider Wyvern for its improved factuality and figure integration.

Method details
  • Implemented with LangChain
  • Search uses Serper.dev API for top‑relevant resources
  • Reasoning agents use DeepSeek‑R1 and other agents use DeepSeek‑V3, both at temperature 0
  • Evaluator models are DeepSeek‑V3 (temp 1) and Qwen3‑32B (temp 0.6)
  • Baselines compared are STORM, WebThinker, and WikiAutoGen
  • Human study involved 27 volunteers, 23 completed questionnaires with nine pairwise comparisons per baseline
Numbers
  • Figures informativeness preference, 87%, over recent baseline
  • Usefulness preference over STORM, 100%, over WebThinker, 62.50%, over WikiAutoGen, 87.50%
  • Citation recall, 73.79%, versus STORM (+21.50 pp) and WikiAutoGen (+42.34 pp)
  • Citation precision, 75.45%, gain 1.12× over baseline
  • Citation recall after revision, 73.79%, up from 60.92% (gain 1.21×)
Limitations

The automatic LLM‑based evaluation still fails to capture human judgments for some baselines and exhibits ordering bias in up to 36.11% of comparisons.

Wyvern’s reports are rated as more useful than those produced by three alternative methods in 63% to 100% of instances.Found in the source text, word for word.

Picked because: Describes Wyvern, a multi‑agent framework for generating grounded multimodal reports, with open‑source components.

Paper 7 of 9

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Yuhao Zhan, Bingxiang He, Zecong Tang and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing evaluations optimize agents under fixed execution conditions and never test recovery after those conditions change. Simply re‑using the source solution in the mutated target fails because agents must infer hidden physics changes and redesign the code.

Approach

The paper introduces PACE‑Bench, a simulator‑grounded benchmark of 144 source‑to‑target adaptation pairs across six physics domains. Agents receive a working source design that fails in the target and must iteratively adapt it using diagnostic sandbox feedback within a 20‑attempt budget. The Uniform Suffix lists variables that may differ without revealing the exact change. The evaluation measures physical inference and mechanism redesign. Ten self‑evolving methods are compared against a Vanilla baseline.

Result

Reflexion combined with Qwen3‑14B succeeds on only 35.9% of full‑benchmark pairs, while GPT‑5.5 solves 66.7% of the Statics subset under the full 20‑attempt budget. Vanilla performance on the Statics subset shows Pass@2 ranging from 8.3% (Qwen3‑4B) to 66.7% (GPT‑5.5) and Score@2 up to 78.1 for GPT‑5.5.

Why it matters

Researchers developing self‑evolving LLM agents and physics‑based code generation should care because PACE‑Bench reveals that current methods struggle with dynamic physics adaptation and that mechanism redesign, not just parameter inference, is the key bottleneck.

Method details
  • Models: Qwen3‑4B, Qwen3‑8B, Qwen3‑14B, Qwen3‑32B, DeepSeek‑V4‑Pro, GPT‑5.5, Gemini‑3.1‑Pro, Claude‑Opus‑4.7, Kimi‑K2.6, MiniMax‑M2.7.
  • Self‑evolving methods: Reflexion, Self‑Refine, ACE, ExpeL, ReasoningBank, Tree‑of‑Thoughts (ToT), CodeEvolve, SEAL, RAGEN, TTT‑Discover.
  • Dataset: 144 source‑to‑target pairs across six physics categories; Statics subset contains 24 pairs; Kinematics subset used for cost‑normalized results.
  • Interaction budget: 20 attempts per pair (five‑attempt ablation also evaluated).
  • Baseline: Vanilla iterative submission without additional self‑evolving mechanisms.
  • Ablations: five‑attempt budget comparison; revealing exact physical changes does not improve performance.
Numbers
  • Pass@2 35.9% for Reflexion + Qwen3‑14B on full benchmark
  • Pass@2 66.7% for GPT‑5.5 on Statics subset
  • Pass@2 37.5% for Qwen3‑14B (Vanilla) on Statics subset
  • Pass@2 45.8% for DeepSeek‑V4‑Pro (Vanilla) on Statics subset
  • Score@2 78.1 for GPT‑5.5 (Vanilla) on Statics subset
  • Score@2 37.5 for Qwen3‑14B (Vanilla) on Statics subset
Limitations

The benchmark remains far from saturated; even the best configurations fail on a substantial fraction of pairs and the paper does not demonstrate a method that reliably solves all adaptation cases.

simulator‑grounded reflection is more reliable than unverified self‑revisionFound in the source text, word for word.

Picked because: Provides PACE‑Bench, a benchmark suite for physics adaptation via code evolution, useful for evaluating adaptive agents.

Paper 8 of 9

Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations

Toby D. Pilditch · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

LLM evaluations traditionally use fixed sampling budgets, testing every item the same number of times even after estimates become precise. This uniform repetition wastes compute because uncertainty varies across items and model-task combinations. Simply increasing the overall budget does not address the uneven variance and leads to unnecessary sampling.

Approach

The paper introduces optstop, a precision‑based adaptive stopping framework that treats evaluation as a sequential measurement problem. It builds on hierarchical Bayesian inference to monitor the width of posterior credible intervals for each item or group. When the interval falls below a user‑specified delta, sampling stops for that unit. A safeguard samples more cautiously as measured performance approaches zero to protect rare‑success cases. optstop integrates with the inspect_ai evaluation framework for live early stopping and can also be applied retrospectively.

Result

In an illustrative 200‑item, 10‑epoch evaluation optstop removed 57%, 97% of planned trials across nine validation settings while yielding overall conclusions equivalent to the full run. Model comparison showed the logit‑normal hierarchy had far fewer divergences than the Beta‑Binomial (120 fewer in the mid‑range, 192 fewer near boundaries, and 9,206 fewer at exact boundaries).

Why it matters

Researchers and engineers conducting large‑scale LLM evaluations should adopt optstop to allocate compute based on uncertainty rather than fixed repetition counts, reducing wasted resources while preserving statistical conclusions.

Method details
  • Hierarchical logit‑normal models are used for binary, ordinal, and continuous outcome pathways.
  • MCMC inference runs with 4 chains and 1,000 post‑warmup draws by default.
  • Downstream analyses use 4 chains with 2,000 posterior draws per chain (8,000 total draws) and 94% HDIs.
  • Experiments evaluate 200 items over 10 epochs per cell.
  • Benchmarks include MATH Level 5 (GPT‑3.5 Turbo), GPQA Diamond (GPT‑4o), MMLU (GPT‑4o), and WritingBench with Claude Sonnet 4.5.
Numbers
  • trial removal, 57%, 97%, across nine validation settings
  • absolute bias, 0.022, logit‑normal vs Beta‑Binomial mid‑range
  • CI width, 0.070, logit‑normal vs Beta‑Binomial mid‑range
  • coverage, 0.80, logit‑normal mid‑range (nominal 94%)
  • divergence ratio, 120, mid‑range (Beta‑Binomial divergences divided by logit‑normal)
  • effective sample size, 4,600, logit‑normal at boundaries (vs 7 for Beta‑Binomial)
Limitations

The paper does not establish how optstop performs on evaluation designs that do not fit a hierarchical Bayesian framework.

it removes 57%, 97% of planned trials across nine validation settings, with overall conclusions equivalent to the full run.Found in the source text, word for word.

Picked because: Offers optstop, a Bayesian optimal‑stopping method for LLM evaluation that adaptively allocates samples, with released implementation.

Paper 9 of 9

DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding

Zewen Jin, Shen Fu, Zeping Duan and 8 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

MoE inference in small‑batch decoding is memory‑bound and bottlenecked by expert weight loading, and existing fixes such as post‑training weight compression or fine‑grained expert design either degrade accuracy or add extra computation and communication overhead.

Approach

DeaMoE groups experts into several departments that share most parameters, while each expert retains a few private parameters. A two‑stage routing strategy avoids redundant weight loading. The shared department backbone and expert‑specific transforms are connected by SiLU non‑linear operators at gate, up, and down interfaces. Expert matrices are initialized to the identity to start from a well‑conditioned state. The design matches the baseline in total parameters and per‑token FLOPs but reduces the amount of weight movement during decoding.

Result

DeaMoE cuts per‑step loaded weights by up to 50.9% and yields up to 1.33× end‑to‑end TPOT speedup for the 7B model on an A40 GPU, while microbenchmarks show peak speedups of 2.00× on A40 and 1.97× on H100 for DeepSeek‑V3; downstream accuracy on BoolQ improves to 62.39 versus 61.47 for the baseline.

Why it matters

Teams deploying interactive LLM services with large‑scale MoE models and small‑batch decoding should adopt DeaMoE to lower latency and increase throughput, especially on bandwidth‑limited GPUs.

Method details
  • 7.3B total parameters for both DeaMoE‑7B and Baseline‑7B
  • Pre‑training on RedPajama‑v1 dataset for 110B tokens
  • Training uses PyTorch FSDP with SiLU activation and identity initialization for expert matrices
  • Inference integrated into vLLM v0.13.0 with Triton operators and CUDA Graph on NVIDIA A40 and H100 GPUs
  • Baselines include a vanilla MoE model (Baseline‑7B) and larger models DeepSeek‑V3 and Qwen3‑235B‑A22B for microbenchmarks
  • Ablations study non‑linear operator choices (SiLU vs GeLU/RMSNorm) and expert‑initialization strategies
Numbers
  • per‑step loaded weights reduction, 50.9%, vs vanilla MoE
  • end‑to‑end TPOT speedup, 1.33×, vs Baseline‑7B on A40
  • peak speedup DeepSeek‑V3, 2.00×, vs baseline on A40
  • peak speedup DeepSeek‑V3, 1.97×, vs baseline on H100
  • BoolQ accuracy, 62.39, DeaMoE‑7B vs 61.47 baseline
  • throughput improvement under 20 ms TPOT budget, 1.83×, vs baseline
Limitations

The paper shows limited or even negative gains for small‑expert configurations on high‑bandwidth hardware such as H100, indicating the method is less effective when expert weights fit in cache.

DeaMoE reduces per‑step loaded weights by up to 50.9% and achieves up to 1.33 end‑to‑end TPOT speedup for the pre‑trained 7B model on A40Found in the source text, word for word.

Picked because: Develops DeaMoE, an efficient Mixture‑of‑Experts architecture optimized for small‑batch decoding, including code and performance results.