Local inference with open-weight 27B models was limited by the memory required for KV attention state, causing the mlx-vlm baseline to hit a 21,000 MiB guard after only 30,720 token positions, and simply compressing KV bits does not solve the large reconstruction workspace overhead.
Approach
JustFit introduces three runtime components that operate together: KVExec performs compressed KV execution with on‑the‑fly reconstruction; PhaseSwap manages which model components stay resident in memory; and StateTrans preserves and transfers state across requests. The components share page tables and ownership leases so that KV tensors are materialized just in time, released when no longer needed, and reused for subsequent requests, all independent of weight quantization.
Result
JustFit raised the single‑request completed context to 212,992 token positions (6.93× the baseline) and a two‑request configuration reached 229,376 positions in aggregate. Throughput on a 32K‑input, 64‑output probe was 19.11 tokens/s, while a repeated 32K+6K workload showed a median peak memory footprint of 16,374 MiB. The system answered 29 of 30 AIME 2026 problems correctly, demonstrating effective reasoning with the compressed state.
Why it matters
Researchers and engineers building local LLM services on memory‑constrained devices should care because JustFit shows how just‑in‑time state management can enable >200K token contexts without extra hardware.
Method details
Model: Qwen3.8‑27B with MXFP4 quantization (27 billion parameters)
Runtime: MLX on an Apple M4 Pro MacBook with 24 GiB unified memory
Baseline: mlx‑vlm runtime without JustFit mechanisms
Evaluation datasets: AIME 2026 problem set (30 problems) and synthetic long‑context workloads
Ablations: configurations 0‑7 adding KVExec, PhaseSwap, StateTrans as listed in Table 5
Numbers
212,992 positions, 6.93× baseline
30,720 baseline positions
19.11 tokens/s on 32K‑input, 64‑output probe
16,374 MiB median peak process footprint
29/30 AIME problems correct
196,608 input tokens and 16,384 output tokens completed in three independent runs
Limitations
The study is limited to a single Apple‑silicon laptop, does not provide a matched PhaseSwap‑off/on ablation, and does not evaluate broader workloads or hardware platforms.
increasing completed single-request context from the mlx-vlm baseline's 30,720 positions to 212,992 (6.93)Found in the source text, word for word.
Picked because: It delivers a practical runtime for serving large language models on modest hardware with just‑in‑time state management, enabling self‑hosted deployment.
Qi Wu, Lohan Lemire, Kai Meng and 8 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Model serving required expertise across GPU kernels, ML framework graphs, model server code, and feature processing, and the lack of cross‑layer expertise left large inefficiencies; simply optimizing one layer (e.g., rewriting Python deserialization) does not resolve bottlenecks that shift across the stack.
Approach
FlashVector implements a closed loop consisting of Profile, Diagnose, Optimize, and Verify stages that is shared across all stack layers via a layer‑agent abstraction; each layer (GPU kernel, ML framework, model server, feature service) provides its own tooling to plug into the loop. The Knowledge Base stores historical optimization attempts, and the Optimization Loop continuously proposes candidates, validates them globally against replayed traffic, and retains only those that clear measurement noise. Optimizations are applied locally (e.g., C++ rewrite of Triton input conversion) but accepted only after system‑wide verification. The loop runs continuously, feeding back results to guide future searches.
Result
FlashVector delivered up to 2× throughput increase and up to 1.98× latency speedup on the model server and up to 1.6× throughput increase on the feature store; Table 1 shows latency speedups ranging from 1.30× to 1.98× and throughput gains from 1.10× to 2.00× across six production models.
Why it matters
Large‑scale recommender‑system teams and serving‑infrastructure engineers should care because FlashVector shows that a unified LLM‑driven agent can automatically close cross‑layer performance gaps without hand‑crafted per‑layer solutions.
Method details
Early‑stage two‑tower retrieval model served by NVIDIA Triton Inference Server
Multi‑task recommendation model using PyTorch AOTInductor, FP16 quantized, batch size 2000 per request on RTX PRO 6000 Blackwell GPU
Baseline serialization of Triton inputs performed in Python deserialization module
Embedding bag kernel originally consumed 28.1% of GPU time
FlashVector fused six attention kernels into one, reducing per‑module memory traffic to roughly one‑third
Optimization reduced dispatcher overhead from 44.8% to 28.5%
Numbers
throughput increase, 2x, model server vs baseline
latency speedup, 1.98x, model server vs baseline
throughput increase, 1.6x, feature store vs baseline
latency speedup, 1.98x, Ranking model 2 vs baseline
throughput increase, 2.00x, Ranking model 2 vs baseline
embedding bag kernel GPU time, 28.1%, before optimization
Limitations
The paper does not evaluate FlashVector on serving stacks outside Unity’s Vector platform or provide detailed ablations of each agent component.
FlashVector achieved up to 2 throughput increase and up to 1.98 latency speedup on model serverFound in the source text, word for word.
Picked because: It presents FlashVector, an agent that automates hierarchical model‑serving stack optimization across GPU kernels, frameworks, and feature processing, a concrete DevOps tool.
Fengshuo Liu, Ying Liu, Ruize Sun and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Prior to this work, coding‑agent leaderboards treated small aggregate score gaps as meaningful rankings, but many instances are degenerate-solved by all or none-so they cannot differentiate systems. The obvious fix of retiring those degenerate instances does not work because paired significance tests already ignore them and retirement adds no statistical power.
Approach
The authors audit all 254 SWE‑bench submissions using per‑instance verdicts to identify degenerate versus discriminating instances and compute nesting metrics. They apply cell‑mean interaction tests with Holm correction and exact paired McNemar tests to assess separability of adjacent systems. A five‑step audit protocol is introduced: profile shared outcomes, test paired differences, report grouping sensitivity, estimate the instance budget needed for resolution, and publish provenance. The analysis distinguishes model and scaffold contributions by normalising metadata and fitting an additive least‑squares model. Recommendations include reporting comparison‑set‑specific resolution, tiered rankings, and machine‑readable provenance rather than raw scores.
Result
The leading two Verified submissions each solve 396 of 500 instances, while the top ten collectively share 285 successes and 51 failures, leaving only 164 discriminating instances. Frontier solution sets exhibit a median nesting of 0.935 against a baseline of 0.774. Within‑model scaffold variation has a median spread of 29.8 pp, far larger than the 8.8‑point spread of the top thirty overall. Exact paired McNemar tests find no significant separation among any of the 29 adjacent top‑thirty Verified pairs at the 0.05 level.
Why it matters
Benchmark designers and organizations selecting models should stop treating small score differences as rankings and instead report tiered performance with explicit provenance, because current leaderboards lack resolution to order top systems.
Method details
Verified split: 500 instances, 134 submissions
Lite split: 299 instances, 84 submissions
Top‑two entries each resolve 396 of 500 Verified instances
Within‑model scaffold range median 29.8 percentage points
Within‑scaffold model range median 8.8 percentage points
Paired McNemar tests separate 0 of 29 adjacent Verified top‑thirty pairs at alpha=0.05
Numbers
Verified top‑two resolved instances, 396/500, compared to each other
Top‑ten shared successes, 285, compared to total 500 instances
Median nesting, 0.935, compared to baseline 0.774
Within‑model scaffold range median, 29.8 percentage points, compared to top‑thirty spread 8.8 points
KR‑20 for top‑ten Verified, 0.475
Paired McNemar separation, 0 of 29 adjacent Verified pairs, alpha=0.05
Limitations
The observational design cannot identify causal scaffold effects and non‑rejection in tests does not establish equivalence of systems.
The top strip counts how many of all submissions resolve each instance, a difficulty gradient the top ten themselves cannot see.Found in the source text, word for word.
Picked because: It audits the SWE‑bench coding‑agent leaderboard and proposes concrete metrics for evaluating and verifying coding agents, guiding practitioners.
Shuhan Xue, Jianyuan Zhong, Ziyuan Nan and 10 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Existing scientific agents lacked a mechanism to continuously improve both their procedural harness and underlying model, leading to stagnant performance on evolving tasks. Simply retraining the model without updating the harness does not work because the environment difficulty changes and the fixed harness limits the usefulness of new model capabilities.
Approach
ScienceBuddy introduces recursive-in-recursive self-improvement, coupling inner recursion that refines the agent harness while keeping the model fixed with outer recursion that trains the model via reinforcement learning under the updated harness. The inner loop uses researcher feedback to revise harness procedures and skills. The outer loop augments the environment, generates rubric‑based rewards, and applies the GRPO algorithm to update the model. After each outer update the harness is re‑evaluated with the new model before the next inner adaptation. This nested schedule creates a feedback loop between procedural and model improvements.
Result
Harness adaptation raised validation accuracy from 31.1% to 51.1%, a 20‑point gain, while the best adaptation batch achieved 75.0% first‑response accuracy. Model reinforcement learning increased problem coverage from 48.3% to 67.8%, a 19.5‑point gain, and training accuracy rose steadily over the two‑hour RL period.
Why it matters
Researchers building interactive scientific AI systems should care because ScienceBuddy demonstrates how coordinated procedural and model updates can substantially improve task success and coverage.
Method details
Datasets: LAB‑Bench and Biomni‑Eval1 tasks covering literature reading, database judgments, protocol troubleshooting, and gene/variant assessment.
Inner recursion: harness adaptation on 24 adaptation batches, selecting the harness with highest first‑response accuracy.
Outer recursion: model reinforcement learning using GRPO for approximately two hours of RL rollouts.
Evaluation: first‑response accuracy, validation accuracy, and problem coverage measured with pass@4 under a fixed attempt budget.
Numbers
validation accuracy, 31.1% vs 51.1%, compared initial vs selected harness
first‑response accuracy, 75.0%, best batch across 24 adaptation batches
problem coverage, 48.3% vs 67.8%, before vs after RL under fixed harness
task count, 895, total tasks across four scientific families
adaptation batches, 24, batches used for harness refinement
training time, ~2 hours, duration of model reinforcement learning
Limitations
The paper does not establish performance beyond the presented benchmark families or long‑term scalability of the recursive‑in‑recursive framework.
Validation accuracy increases from 31.1% to 51.1%, a gain of 20 percentage points with model weights fixed.Found in the source text, word for word.
Picked because: ScienceBuddy is released as an interactive workspace that integrates continuously improving scientific agents into researcher workflows, offering usable tooling.
Sara Vera Marjanović, Jiacheng Xu, Aleksandr Laptev and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.MA
Problem
Prior multi-agent systems assumed that adding more open-source models would improve reasoning performance, but in practice larger heterogeneous pools often hurt results, and simply adding any extra model does not yield gains.
Approach
The authors systematically evaluate eight model‑selection strategies using both before‑generation routing and after‑generation majority‑voting or LLM‑as‑a‑judge architectures. They extract pre‑evaluation signals (size, release date, family, domain specialization) and post‑evaluation signals (accuracy, correct‑answer diversity, error diversity) to rank models. Subsets are formed by ordering models according to each signal and selecting the top‑k candidates. Oracle MAS performance is measured as the maximum possible pass@1 score, and real MAS performance is compared against it. The study focuses on scientific reasoning benchmarks to assess how selection impacts overall system stability.
Result
The study finds a large gap between the theoretical oracle MAS potential and actual MAS performance; expanding the candidate pool frequently reduces accuracy below that of the single best base model. Selecting candidates within a single model family yields the best relative improvement over a standalone model, while heterogeneous pools introduce instability.
Why it matters
Researchers and engineers designing multi‑agent systems should prioritize careful model‑pool selection, especially favoring homogeneous families, to avoid performance degradation.
Method details
23 LM agents from 6 architecture families, released 2024‑2026, ranging 2B‑1.6T parameters.
5 generations per model are collected for each question.
Benchmarks: Humanity’s Last Exam (HLE), GPQA‑Diamond, Frontier Science‑Olympiad.
Training set of 15.5k questions (24.9% multiple‑choice) used for router training and performance calibration.
GPT‑OSS‑120b serves as the judge for open‑form answer equivalence, agreeing with human annotators on 93% of decisions (Cohen’s kappa 0.63).
Selection methods include size‑sorted, family‑grouped, and LLM‑chosen rankings.
Numbers
93% agreement with human annotators, Cohen’s kappa 0.63, compared against human judgment
15.5k training questions, compared against test benchmarks
24.9% multiple‑choice questions, compared against total training set
5 generations per model, compared against single generation baselines
23 models evaluated, compared against individual model baselines
6 architecture families, compared against cross‑family selection strategies
Limitations
The paper does not establish how its findings generalize beyond the three scientific reasoning benchmarks studied.
Expanding candidate pool sizes often degrades performance below that of the top performing base-model.Found in the source text, word for word.
Picked because: It systematically evaluates model‑selection strategies for multi‑agent systems and provides actionable guidance for building effective model pools.