arXiv digest

Wednesday

September 16, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management

Yuhua Chen · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Local inference with open-weight 27B models was limited by the memory required for KV attention state, causing the mlx-vlm baseline to hit a 21,000 MiB guard after only 30,720 token positions, and simply compressing KV bits does not solve the large reconstruction workspace overhead.

Approach

JustFit introduces three runtime components that operate together: KVExec performs compressed KV execution with on‑the‑fly reconstruction; PhaseSwap manages which model components stay resident in memory; and StateTrans preserves and transfers state across requests. The components share page tables and ownership leases so that KV tensors are materialized just in time, released when no longer needed, and reused for subsequent requests, all independent of weight quantization.

Result

JustFit raised the single‑request completed context to 212,992 token positions (6.93× the baseline) and a two‑request configuration reached 229,376 positions in aggregate. Throughput on a 32K‑input, 64‑output probe was 19.11 tokens/s, while a repeated 32K+6K workload showed a median peak memory footprint of 16,374 MiB. The system answered 29 of 30 AIME 2026 problems correctly, demonstrating effective reasoning with the compressed state.

Why it matters

Researchers and engineers building local LLM services on memory‑constrained devices should care because JustFit shows how just‑in‑time state management can enable >200K token contexts without extra hardware.

Method details
  • Model: Qwen3.8‑27B with MXFP4 quantization (27 billion parameters)
  • Runtime: MLX on an Apple M4 Pro MacBook with 24 GiB unified memory
  • Baseline: mlx‑vlm runtime without JustFit mechanisms
  • Evaluation datasets: AIME 2026 problem set (30 problems) and synthetic long‑context workloads
  • Ablations: configurations 0‑7 adding KVExec, PhaseSwap, StateTrans as listed in Table 5
Numbers
  • 212,992 positions, 6.93× baseline
  • 30,720 baseline positions
  • 19.11 tokens/s on 32K‑input, 64‑output probe
  • 16,374 MiB median peak process footprint
  • 29/30 AIME problems correct
  • 196,608 input tokens and 16,384 output tokens completed in three independent runs
Limitations

The study is limited to a single Apple‑silicon laptop, does not provide a matched PhaseSwap‑off/on ablation, and does not evaluate broader workloads or hardware platforms.

increasing completed single-request context from the mlx-vlm baseline's 30,720 positions to 212,992 (6.93)Found in the source text, word for word.

Picked because: It delivers a practical runtime for serving large language models on modest hardware with just‑in‑time state management, enabling self‑hosted deployment.

Paper 2 of 5

FlashVector: Agent for Hierarchical Model Serving Stack Optimization

Qi Wu, Lohan Lemire, Kai Meng and 8 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Model serving required expertise across GPU kernels, ML framework graphs, model server code, and feature processing, and the lack of cross‑layer expertise left large inefficiencies; simply optimizing one layer (e.g., rewriting Python deserialization) does not resolve bottlenecks that shift across the stack.

Approach

FlashVector implements a closed loop consisting of Profile, Diagnose, Optimize, and Verify stages that is shared across all stack layers via a layer‑agent abstraction; each layer (GPU kernel, ML framework, model server, feature service) provides its own tooling to plug into the loop. The Knowledge Base stores historical optimization attempts, and the Optimization Loop continuously proposes candidates, validates them globally against replayed traffic, and retains only those that clear measurement noise. Optimizations are applied locally (e.g., C++ rewrite of Triton input conversion) but accepted only after system‑wide verification. The loop runs continuously, feeding back results to guide future searches.

Result

FlashVector delivered up to 2× throughput increase and up to 1.98× latency speedup on the model server and up to 1.6× throughput increase on the feature store; Table 1 shows latency speedups ranging from 1.30× to 1.98× and throughput gains from 1.10× to 2.00× across six production models.

Why it matters

Large‑scale recommender‑system teams and serving‑infrastructure engineers should care because FlashVector shows that a unified LLM‑driven agent can automatically close cross‑layer performance gaps without hand‑crafted per‑layer solutions.

Method details
  • Early‑stage two‑tower retrieval model served by NVIDIA Triton Inference Server
  • Multi‑task recommendation model using PyTorch AOTInductor, FP16 quantized, batch size 2000 per request on RTX PRO 6000 Blackwell GPU
  • Baseline serialization of Triton inputs performed in Python deserialization module
  • Embedding bag kernel originally consumed 28.1% of GPU time
  • FlashVector fused six attention kernels into one, reducing per‑module memory traffic to roughly one‑third
  • Optimization reduced dispatcher overhead from 44.8% to 28.5%
Numbers
  • throughput increase, 2x, model server vs baseline
  • latency speedup, 1.98x, model server vs baseline
  • throughput increase, 1.6x, feature store vs baseline
  • latency speedup, 1.98x, Ranking model 2 vs baseline
  • throughput increase, 2.00x, Ranking model 2 vs baseline
  • embedding bag kernel GPU time, 28.1%, before optimization
Limitations

The paper does not evaluate FlashVector on serving stacks outside Unity’s Vector platform or provide detailed ablations of each agent component.

FlashVector achieved up to 2 throughput increase and up to 1.98 latency speedup on model serverFound in the source text, word for word.

Picked because: It presents FlashVector, an agent that automates hierarchical model‑serving stack optimization across GPU kernels, frameworks, and feature processing, a concrete DevOps tool.

Paper 3 of 5

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

Fengshuo Liu, Ying Liu, Ruize Sun and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Prior to this work, coding‑agent leaderboards treated small aggregate score gaps as meaningful rankings, but many instances are degenerate-solved by all or none-so they cannot differentiate systems. The obvious fix of retiring those degenerate instances does not work because paired significance tests already ignore them and retirement adds no statistical power.

Approach

The authors audit all 254 SWE‑bench submissions using per‑instance verdicts to identify degenerate versus discriminating instances and compute nesting metrics. They apply cell‑mean interaction tests with Holm correction and exact paired McNemar tests to assess separability of adjacent systems. A five‑step audit protocol is introduced: profile shared outcomes, test paired differences, report grouping sensitivity, estimate the instance budget needed for resolution, and publish provenance. The analysis distinguishes model and scaffold contributions by normalising metadata and fitting an additive least‑squares model. Recommendations include reporting comparison‑set‑specific resolution, tiered rankings, and machine‑readable provenance rather than raw scores.

Result

The leading two Verified submissions each solve 396 of 500 instances, while the top ten collectively share 285 successes and 51 failures, leaving only 164 discriminating instances. Frontier solution sets exhibit a median nesting of 0.935 against a baseline of 0.774. Within‑model scaffold variation has a median spread of 29.8 pp, far larger than the 8.8‑point spread of the top thirty overall. Exact paired McNemar tests find no significant separation among any of the 29 adjacent top‑thirty Verified pairs at the 0.05 level.

Why it matters

Benchmark designers and organizations selecting models should stop treating small score differences as rankings and instead report tiered performance with explicit provenance, because current leaderboards lack resolution to order top systems.

Method details
  • Verified split: 500 instances, 134 submissions
  • Lite split: 299 instances, 84 submissions
  • Top‑two entries each resolve 396 of 500 Verified instances
  • Within‑model scaffold range median 29.8 percentage points
  • Within‑scaffold model range median 8.8 percentage points
  • Paired McNemar tests separate 0 of 29 adjacent Verified top‑thirty pairs at alpha=0.05
Numbers
  • Verified top‑two resolved instances, 396/500, compared to each other
  • Top‑ten shared successes, 285, compared to total 500 instances
  • Median nesting, 0.935, compared to baseline 0.774
  • Within‑model scaffold range median, 29.8 percentage points, compared to top‑thirty spread 8.8 points
  • KR‑20 for top‑ten Verified, 0.475
  • Paired McNemar separation, 0 of 29 adjacent Verified pairs, alpha=0.05
Limitations

The observational design cannot identify causal scaffold effects and non‑rejection in tests does not establish equivalence of systems.

The top strip counts how many of all submissions resolve each instance, a difficulty gradient the top ten themselves cannot see.Found in the source text, word for word.

Picked because: It audits the SWE‑bench coding‑agent leaderboard and proposes concrete metrics for evaluating and verifying coding agents, guiding practitioners.

Paper 4 of 5

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

Shuhan Xue, Jianyuan Zhong, Ziyuan Nan and 10 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing scientific agents lacked a mechanism to continuously improve both their procedural harness and underlying model, leading to stagnant performance on evolving tasks. Simply retraining the model without updating the harness does not work because the environment difficulty changes and the fixed harness limits the usefulness of new model capabilities.

Approach

ScienceBuddy introduces recursive-in-recursive self-improvement, coupling inner recursion that refines the agent harness while keeping the model fixed with outer recursion that trains the model via reinforcement learning under the updated harness. The inner loop uses researcher feedback to revise harness procedures and skills. The outer loop augments the environment, generates rubric‑based rewards, and applies the GRPO algorithm to update the model. After each outer update the harness is re‑evaluated with the new model before the next inner adaptation. This nested schedule creates a feedback loop between procedural and model improvements.

Result

Harness adaptation raised validation accuracy from 31.1% to 51.1%, a 20‑point gain, while the best adaptation batch achieved 75.0% first‑response accuracy. Model reinforcement learning increased problem coverage from 48.3% to 67.8%, a 19.5‑point gain, and training accuracy rose steadily over the two‑hour RL period.

Why it matters

Researchers building interactive scientific AI systems should care because ScienceBuddy demonstrates how coordinated procedural and model updates can substantially improve task success and coverage.

Method details
  • Datasets: LAB‑Bench and Biomni‑Eval1 tasks covering literature reading, database judgments, protocol troubleshooting, and gene/variant assessment.
  • Task collection: 895 tasks (96 LitQA2, 511 DbQA, 108 ProtocolQA, 180 GWAS) across 16 subtopics.
  • Inner recursion: harness adaptation on 24 adaptation batches, selecting the harness with highest first‑response accuracy.
  • Outer recursion: model reinforcement learning using GRPO for approximately two hours of RL rollouts.
  • Evaluation: first‑response accuracy, validation accuracy, and problem coverage measured with pass@4 under a fixed attempt budget.
Numbers
  • validation accuracy, 31.1% vs 51.1%, compared initial vs selected harness
  • first‑response accuracy, 75.0%, best batch across 24 adaptation batches
  • problem coverage, 48.3% vs 67.8%, before vs after RL under fixed harness
  • task count, 895, total tasks across four scientific families
  • adaptation batches, 24, batches used for harness refinement
  • training time, ~2 hours, duration of model reinforcement learning
Limitations

The paper does not establish performance beyond the presented benchmark families or long‑term scalability of the recursive‑in‑recursive framework.

Validation accuracy increases from 31.1% to 51.1%, a gain of 20 percentage points with model weights fixed.Found in the source text, word for word.

Picked because: ScienceBuddy is released as an interactive workspace that integrates continuously improving scientific agents into researcher workflows, offering usable tooling.

Paper 5 of 5

Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems

Sara Vera Marjanović, Jiacheng Xu, Aleksandr Laptev and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.MA

Problem

Prior multi-agent systems assumed that adding more open-source models would improve reasoning performance, but in practice larger heterogeneous pools often hurt results, and simply adding any extra model does not yield gains.

Approach

The authors systematically evaluate eight model‑selection strategies using both before‑generation routing and after‑generation majority‑voting or LLM‑as‑a‑judge architectures. They extract pre‑evaluation signals (size, release date, family, domain specialization) and post‑evaluation signals (accuracy, correct‑answer diversity, error diversity) to rank models. Subsets are formed by ordering models according to each signal and selecting the top‑k candidates. Oracle MAS performance is measured as the maximum possible pass@1 score, and real MAS performance is compared against it. The study focuses on scientific reasoning benchmarks to assess how selection impacts overall system stability.

Result

The study finds a large gap between the theoretical oracle MAS potential and actual MAS performance; expanding the candidate pool frequently reduces accuracy below that of the single best base model. Selecting candidates within a single model family yields the best relative improvement over a standalone model, while heterogeneous pools introduce instability.

Why it matters

Researchers and engineers designing multi‑agent systems should prioritize careful model‑pool selection, especially favoring homogeneous families, to avoid performance degradation.

Method details
  • 23 LM agents from 6 architecture families, released 2024‑2026, ranging 2B‑1.6T parameters.
  • 5 generations per model are collected for each question.
  • Benchmarks: Humanity’s Last Exam (HLE), GPQA‑Diamond, Frontier Science‑Olympiad.
  • Training set of 15.5k questions (24.9% multiple‑choice) used for router training and performance calibration.
  • GPT‑OSS‑120b serves as the judge for open‑form answer equivalence, agreeing with human annotators on 93% of decisions (Cohen’s kappa 0.63).
  • Selection methods include size‑sorted, family‑grouped, and LLM‑chosen rankings.
Numbers
  • 93% agreement with human annotators, Cohen’s kappa 0.63, compared against human judgment
  • 15.5k training questions, compared against test benchmarks
  • 24.9% multiple‑choice questions, compared against total training set
  • 5 generations per model, compared against single generation baselines
  • 23 models evaluated, compared against individual model baselines
  • 6 architecture families, compared against cross‑family selection strategies
Limitations

The paper does not establish how its findings generalize beyond the three scientific reasoning benchmarks studied.

Expanding candidate pool sizes often degrades performance below that of the top performing base-model.Found in the source text, word for word.

Picked because: It systematically evaluates model‑selection strategies for multi‑agent systems and provides actionable guidance for building effective model pools.