arXiv digest

Wednesday

September 23, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

Jennifer Williams, Dave Farris, Jeff Farris and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing benchmarks either ignore inference engineering or focus only on isolated kernel generation and performance optimization, so they do not evaluate the full production inference serving stack.

Approach

SWE-Serve constructs 53 repository‑grounded tasks from recent SGLang production changes, each paired with hidden functional, regression, and where applicable end‑to‑end serving tests and calibrated performance gates. Agents receive a task instruction and the codebase at a base commit, produce a patch, and the verifier scores the patch. The benchmark runs in Harbor 0.13.1 sandboxes on CPU or a single H100 GPU, enforcing a closed‑book setting. Evaluation uses the mini‑SWE‑agent v2.4.3 harness, with no‑op and oracle controls plus adversarial verifier review to ensure task validity.

Result

The best‑performing configuration reaches a 75% mean pass@1, while the lowest model achieves 35% mean pass@1, showing a 40‑point spread. End‑to‑end serving tests reject roughly one‑third of patches that pass other tests (45.9% under the verifier versus 69.4% when E2E tests are excluded). Mean task wall‑clock time is 43.3 minutes, compared with 34.3 minutes on DeepSWE, and only 2.4% of attempts fail due to execution limits.

Why it matters

Researchers and engineers developing agentic code‑generation systems for inference serving should care, as SWE‑Serve quantifies the gap between local task completion and production‑correctness.

Method details
  • Evaluated 11 models including Claude Opus 5, Claude Sonnet 5, GPT‑5.6 Sol, Luna, Terra, Kimi K3, DeepSeek V4 Flash 0731, GLM‑5.2, Gemini 3.6 Flash, Laguna S 2.1, and Inkling S.
  • Ran 31 model‑effort configurations (low, medium, high, xhigh, max) across the 53 tasks, each repeated three times.
  • Agents were limited to 350 steps or 210 minutes per task, with individual commands timing out after 120 seconds.
  • Benchmark uses Harbor 0.13.1 to provide CPU or single H100 GPU resources for each sandbox.
  • Baseline comparisons include DeepSWE benchmark metrics for wall‑clock time and task distribution.
Numbers
  • mean pass@1, 75%, best‑performing configuration (Claude Opus 5 and GPT‑5.6 Sol)
  • pass@1 range, 40 percentage points, across 11 models (75% to 35%)
  • execution‑limit failures, 2.4% (42/1,749), across top‑per‑model configurations
  • mean task wall‑clock time, 43.3 minutes, versus DeepSWE 34.3 minutes
  • E2E test rejection rate, 45.9%, under verifier versus 69.4% with E2E tests excluded
  • cost per task, $0.33, $17.40, across models
Limitations

The paper does not establish how agents would perform in open‑book settings or in real‑world deployment beyond the benchmark.

Across 11 models and 31 model‑effort configurations, the best‑performing configuration achieves 75% mean pass@1.Found in the source text, word for word.

Picked because: Introduces SWE-Serve, a released benchmark suite for evaluating agents on real‑world production inference engineering tasks, giving engineers concrete metrics and artifacts.

Paper 2 of 5

SARA: SLO-Aware Resource Allocation for Disaggregated Agentic LLM Services

Shicong Liu, Xianghao Yu, Zhen Gao and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.PF

Problem

Existing disaggregated LLM serving systems rely on hardware profiling, configuration enumeration, or heuristic scheduling, which provides limited analytical guidance for cost‑efficient resource allocation. The obvious fix of more exhaustive profiling or heuristic tuning does not work because it cannot capture the complex interplay of compute, memory bandwidth, and latency SLOs across stages.

Approach

SARA first models the three inference stages-prefill, KV‑cache transfer, and decode-as an M/G/k queue, an M/G/1 queue, and a generalized birth‑death process respectively. Using queuing theory it derives closed‑form tail‑behaviour expressions for stage‑wise latency quantiles under both light‑ and heavy‑tailed workloads. These expressions map workload characteristics, model architecture, and hardware parameters to minimum resource requirements that satisfy SLO constraints. An optimization framework then allocates prefill and decode devices to maximize system goodput while respecting a deployment‑cost budget and quantile‑based SLOs. The resulting allocation respects the quadratic scaling of required devices with input and output lengths and accounts for heavy‑tailed workload amplification.

Result

Simulation and hardware experiments show that SARA predicts stage‑wise SLO latency with mean errors below 5% and achieves a 26.6% average increase in system goodput compared to state‑of‑the‑art baselines under the same deployment cost.

Why it matters

Data‑center planners and LLM service providers should care because SARA offers an analytical way to allocate compute and memory resources that boosts throughput while meeting latency SLOs.

Method details
  • LLaMA 3.1 8B, Qwen2.5 32B, and GPT‑3 175B models are evaluated
  • Inference experiments are run with the SGLang framework on multiple NVIDIA A100 GPUs
  • Heavy‑tailed workloads are generated with a log‑normal input‑length distribution and exponential output‑length distribution
  • Baseline comparisons use state‑of‑the‑art resource‑allocation methods
  • Cost per device is set to USD per hour and bandwidth cost to USD per GiB per hour
Numbers
  • mean error, below 5%, compared to actual stage‑wise SLO
  • goodput improvement, 26.6%, compared to state‑of‑the‑art baseline methods
Limitations

The paper does not explicitly state any limitations of the proposed framework.

improves system goodput by 26.6% on average over state-of-the-art baseline methods under the same deployment costFound in the source text, word for word.

Picked because: Presents SARA, an SLO‑aware analytic scheduler for disaggregated LLM services with open‑source implementation, directly applicable to cloud resource automation.

Paper 3 of 5

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

Trang Nguyen, Eulrang Cho, Bingqing Chen and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Agents needed millions of tokens but limited context windows forced compaction across sessions, and naïve approaches that repeatedly re‑compact already compacted output cause context drift and degrade performance.

Approach

CliffCompaction only truncates or drops tokens without rephrasing, and never compacts a compaction; each pass operates on the original content and discards prior compacted output. This preserves faithful information while keeping the context append‑only. The method is integrated directly into scaffolds or via an API‑proxy for closed systems. It is configured to match a mean peak context budget K across models. By avoiding KV‑cache invalidation, it reduces per‑step cost dramatically. The approach works across different model families and benchmarks.

Result

CliffCompaction matches or improves full‑context performance on Terminal‑Bench at 32K and 16K thresholds, adding over 10 percentage points for less than the cost of two full‑context runs, and reaches 69.7% practical resolution at 16K, surpassing Opus 4.7 (69.4%) and GPT 5.3 Codex (64.7%). On SWE‑bench it preserves most of the full‑context success rates at moderate thresholds. On KernelBench it achieves CUDA speedups of 2.23× after 200 steps and 3.58× after 400 steps, exceeding specialized search agents.

Why it matters

Developers of long‑horizon coding agents should adopt CliffCompaction to cut inference cost while maintaining or improving success rates on large‑context tasks.

Method details
  • Evaluated on SWE‑bench Verified, Terminal‑Bench 2.0/2.1 and KernelBench
  • Used Kimi K2.5, K2.6, K2.7 and GLM 5, 5.1, 5 Turbo, 5.3 Flash, 4.7 Flash models
  • Compared against full‑context runs and native summarization/condensers such as OpenHands’ native condenser
  • Integrated CliffCompaction as a scaffold‑agnostic API proxy for Claude Code
  • Tested token thresholds of 32K, 16K and 8K context budgets
Numbers
  • cost reduction up to 50% under a bounded context
  • Terminal‑Bench 16K CliffCompaction adds over 10 percentage points compared to full‑context
  • Kimi K2.6 + Cliff + SGV at 16K achieves 69.7% practical resolution vs Opus 4.7 69.4%
  • Three‑rollout cost $58.01, lower than single GPT 5.3 Codex run $64.63
  • KernelBench CUDA speedup 2.23× after 200 steps
  • KernelBench CUDA speedup 3.58× after 400 steps
Limitations

The paper does not evaluate beyond the presented benchmarks and explicitly does not compact already compacted output, so it cannot address scenarios requiring recursive compaction.

CliffCompaction makes test-time scaling cost-effective where naive scaling is not.Found in the source text, word for word.

Picked because: Describes CliffCompaction, an autocompaction method that halves context‑window cost while preserving performance, with code released for self‑hosted agents.

Paper 4 of 5

Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation

Lijuan Tang, Yuemeng Zheng · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Local serving stacks reject or mis‑parse tool calls, causing the harness to record capable models as non‑calls and report 0% fidelity; simply adding a text hint does not fix the issue because it harms models with native tool‑call support such as Llama‑3.2.

Approach

The authors isolate the serving layer as an experimental variable by probing four stacks (Ollama, llama.cpp, vLLM, SGLang) and measuring per‑seed fidelity under three protocols: native, native+text‑hint, and uniform text‑tools. They log rejection and retry failures as distinct outcomes instead of collapsing them into assistant messages. They compare turn‑pooled versus per‑instance estimates and report intervals per seed. The method includes a checklist to fix logging, fix the protocol, and hold the serving interface constant for standardized model comparisons.

Result

Native+hint recovers most of the measured fidelity for models accepted by the stack, while the uniform text protocol reduces fidelity for Llama‑3.2. Naïve analysis that ignores rejection metadata can report 0% fidelity for rejected models. Turn‑pooled estimates can differ from per‑instance estimates by up to about 55 points.

Why it matters

Benchmark designers and researchers evaluating local agent tool use must treat the serving stack as part of the measurement protocol, otherwise they risk measuring the stack rather than the model.

Method details
  • Models: Qwen2.5-Coder 0.5B/1.5B/3B/7B/14B, Llama-3.2-3B, Phi-3-mini, Gemma-3-4B, Gemma-3-270m (served via Ollama) and deepseek-v4-flash (cloud).
  • Datasets/tasks: aggregation (Table 1, 8 seeds), dependency‑chain (Table 3, 4 seeds), and a single‑turn HumanEval replication.
  • Baselines: native tool‑call channel, native+text‑hint channel, uniform text‑tools channel.
  • Ablations: adding text hint, using uniform text protocol, constrained decoding, per‑seed vs turn‑pooled aggregation.
  • Serving stacks probed: Ollama, llama.cpp, vLLM, SGLang with default and flag‑enabled configurations.
Numbers
  • 0% fidelity (naïve report for rejected models)
  • up to about 55 points difference between turn‑pooled and per‑instance estimates
  • 8 to 16 turns over 8 seeds (weaker models)
  • 8 seeds (aggregation task)
  • 4 seeds (dependency‑chain task)
  • Qwen2.5-Coder sizes 0.5B/1.5B/3B/7B/14B
Limitations

The study only probes four serving stacks, evaluates two tasks, and tests vLLM and SGLang on only Qwen‑0.5B and Phi‑3, so results may not generalize to other stacks, tasks, or model‑stack pairs.

turn-pooled versus per-instance estimates differ by up to about 55 points.Found in the source text, word for word.

Picked because: Shows how serving‑stack configurations can bias tool‑use evaluation, providing practical guidelines and tooling for verifying LLM agent calls in deployment.

Paper 5 of 5

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

Gaoyuan Du, Anam Nawaz Khan, Rex Zhou and 4 others · abstract · pdf

quote verifiedfigures checkedread: abstract onlycs.LG

Problem

Greedy decoding was assumed deterministic but the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. The obvious fix of using higher precision everywhere does not work because broader FP32 compute actually makes agreement worse.

Approach

The authors perform an empirical error‑propagation analysis and discover that whether a token flips depends primarily on the top‑two logit margin at the LM head relative to the directional perturbation between the top‑two candidates. Based on this insight they design a selective recomputation strategy that runs the LM head in FP32 only when the margin falls below a threshold. This selective FP32 LM head recomputation is applied during greedy decoding while the rest of the model runs in lower precision. The method therefore adds minimal overhead while targeting the critical decision point that causes divergence. Experiments confirm five predictions derived from the analysis, including that more FP32 compute worsens agreement. The approach is evaluated across multiple models, batch sizes, and hardware.

Result

Across the evaluations, 49-100% of prompts diverge between BF16 and FP16. The selective FP32 LM head recomputation yields +22-36 percentage‑point exact agreement on A10G and +12-21 percentage‑point on L4 and A100, with less than 4% latency overhead in low‑batch (batch size <=4) single‑stream inference. The benefit disappears at batch size >=8 or under end‑to‑end FP8.

Why it matters

Practitioners deploying LLM inference who need deterministic greedy outputs should consider the selective FP32 LM head recomputation, especially in low‑batch, single‑stream settings where latency overhead is minimal.

Method details
  • Evaluated six models ranging from 1.1B to 7B parameters across four families and also a 12B model
  • Three benchmarks were used to measure prompt divergence
  • Baseline is standard greedy decoding in BF16 or FP16 without intervention
  • Intervention: selective FP32 LM head recomputation triggered by low top‑two logit margin
  • Ablation: applying more FP32 compute (broader scope) makes agreement worse
  • Batch sizes examined include low‑batch (<=4) and larger batch sizes (>=8)
Numbers
  • 49-100% of prompts diverge
  • 22 layers of accumulated body error
  • +22-36 pp exact agreement on A10G
  • +12-21 pp on L4 and A100
  • less than 4% latency overhead
  • batch size <=4
Limitations

The method is a partial mitigation rather than a universal determinism guarantee; its benefit vanishes when body‑originated error dominates, including at batch size >=8 and under end‑to‑end FP8.

Greedy decoding from large language models is commonly treated as deterministic.Found in the source text, word for word.

Picked because: Demonstrates that greedy decoding is precision‑sensitive across BF16/FP16, offering actionable insights and scripts for robust inference on heterogeneous hardware.