arXiv digest

Wednesday

September 2, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

Olga Tsymboi, Dmitrii Stoianov, Ramil Latypov and 11 others · abstract · pdf

quote verified1 figure not in sourceread: full textcs.CL

Problem

Enterprises faced fragmented GPU pools and quality gaps in instruction following, function‑calling, and internal task distribution, and a naive joint optimization of all objectives caused cross‑domain reward interference that did not improve per‑domain performance.

Approach

The method builds on Qwen3‑32B with a Cyrillic‑dense tokenizer and operates in non‑reasoning mode. First a shared SFT stage mixes all domains, then the checkpoint is forked into three independent GRPO runs, each optimized for a specific reward (instruction following, function‑calling, or general alignment). The three expert checkpoints are merged into a single deployment model via a two‑stage sequential SLERP procedure. This avoids reward interference while addressing each failure mode with a domain‑specific fix.

Result

The final model outperforms the larger baseline on the in‑house Arena (69.6 vs 65.8), improves instruction‑following scores (0.85 vs 0.83) and function‑calling scores (0.79 vs 0.77), lifts general dialogue benchmarks, and now handles 50% of platform traffic (116M requests per month) at a fraction of the serving cost.

Why it matters

Enterprises with data‑residency constraints can consolidate many applications onto a single, cost‑effective LLM while meeting internal quality requirements.

Method details
  • Base model: Qwen3‑32B (32 B parameters)
  • Tokenizer: adapted Cyrillic‑dense tokenizer
  • Training pipeline: shared SFT followed by three separate GRPO runs per axis
  • Reward model: general‑domain RM retained (in‑house‑adapted RM showed no gain)
  • Merging: two‑stage sequential SLERP merging of expert checkpoints
  • Baselines: compared against a ~7× larger model by total parameters
Numbers
  • Arena score 69.6 vs 65.8 (baseline)
  • Instruction‑following 0.85 vs 0.83 (baseline)
  • Function‑calling 0.79 vs 0.77 (baseline)
  • Traffic handled 50% of platform
  • Requests per month 116M
  • Model ~7× larger baseline
Limitations

The paper only evaluates the model in non‑reasoning mode and does not address classification‑type failures or knowledge‑base errors.

In non‑reasoning mode the recipe surpasses a larger by total parameters baseline on the in‑house Arena with 69.6 to 65.8, instruction following with 0.85 to 0.83, and function‑calling with 0.79 to 0.77Found in the source text, word for word.
These figures do not appear in the source text: 32 B. Treat them as unverified.Number check failed.

Picked because: Describes a production‑grade pipeline for self‑hosting an LLM that handles a corporate request mix, with concrete deployment artifacts and performance trade‑offs engineers can adopt.

Paper 2 of 5

Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation

Kefeng Duan, Dewu Zheng, Yanlin Wang and 8 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Existing retrieval-augmented generation methods provide repository context only at the task level and do not identify the decisive tokens that need fine-grained context. The obvious fix of retrieving more global context does not address the fact that errors concentrate at a few critical token positions.

Approach

ACToR introduces a critical token discriminator that monitors hidden states during autoregressive generation to flag critical tokens. When a critical token is detected, the framework triggers on‑demand retrieval of additional repository snippets. Retrieved snippets are weighted by a position‑aware scoring function before being inserted into the prompt. The generator then continues generation conditioned on this dynamically refreshed context. The pipeline thus couples token‑level inference with a weighted dense retriever.

Result

ACToR achieves relative improvements of 8.4% on RepoExec and 15.4% on CoderEval over prior state‑of‑the‑art. On RepoExec it reaches Pass@1 11.27%, Pass@3 19.24% and Pass@5 23.10%; on CoderEval it reaches Pass@1 20.00%, Pass@3 28.70% and Pass@5 33.48%. Ablations show that removing any labeling rule or the dynamic inference module reduces these scores.

Why it matters

Researchers and engineers building repository‑level code generation systems should care because ACToR demonstrates that token‑aware dynamic retrieval can substantially boost correctness without large model changes.

Method details
  • Retriever: UniXcoder dense retriever with cosine similarity scoring.
  • Generators: DeepSeekCoder (DSCoder) 1.3B and 6.7B, and CodeLlama 7B and 13B.
  • Training data: RepoST-Train filtered to 102 repositories yielding 1,203 task samples.
  • Benchmarks: RepoExec (355 Python tasks) and CoderEval (230 Python tasks).
  • Inference settings: context length 1K tokens, max 10 snippets, temperature 0.6, max generated tokens 512, critical‑token thresholds 0.05 (attention influence) and 0.8 (uncertainty).
Numbers
  • relative improvement 8.4% on RepoExec compared to SOTA
  • relative improvement 15.4% on CoderEval compared to SOTA
  • Pass@1 11.27% on RepoExec (full ACToR)
  • Pass@3 19.24% on RepoExec (full ACToR)
  • Pass@1 20.00% on CoderEval (full ACToR)
  • Pass@5 33.48% on CoderEval (full ACToR)
Limitations

The paper does not discuss any limitations.

Experimental results show that ACToR consistently outperforms state-of-the-art methods, achieving relative improvements of 8.4% on RepoExec and 15.4% on CoderEval.Found in the source text, word for word.

Picked because: Presents Adaptive Critical Token‑Aware Retrieval to feed repository‑level context into code‑generation models, addressing real‑world repo size limits and releasing a retrieval system usable in CI/CD pipelines.

Paper 3 of 5

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

Kefeng Duan, Dewu Zheng, Yanlin Wang and 7 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Prior efficient benchmarking methods for SWE agents use only pass/fail outcomes in Item Response Theory, discarding the multi-step execution trajectories; simply adding more pass/fail data does not capture the process signals needed for accurate evaluation.

Approach

PTA-IRT treats historical agent trajectories as privileged information. It first converts each trajectory into a structured semantic summary, then learns trajectory-aware four-parameter logistic (4PL) item measurement characteristics. Using trajectory-aware Fisher information, it selects a difficulty‑stratified calibration subset. Ability estimation for a new agent is performed via Learning Using Privileged Information (teacher‑student training) and test‑time inference. The whole pipeline integrates trajectory representation, informative item selection, and LUPI‑based ability estimation.

Result

Across four SWE benchmarks, PTA-IRT achieves lower mean absolute error and higher Kendall's tau and Spearman's rho than all prior IRT baselines, demonstrating superior score and ranking recovery under low calibration budgets.

Why it matters

Researchers and practitioners building LLM‑based software‑engineering agents should care because PTA-IRT enables accurate performance estimation with far fewer costly task executions.

Method details
  • Datasets: SWE-bench Lite (300 tasks, 35 models), Verified (500 tasks, 70 models), Full (2,294 tasks, 14 models), Pro (730 tasks, 14 models).
  • Baselines compared: Classical IRT, Deep‑IRT, PSN‑IRT, AutoJudger.
  • Trajectory summaries are generated with DeepSeek-V4-Flash11 and embedded with all‑MiniLM‑L6‑v2.
  • Training uses four‑fold cross‑validation with 75% of models for training and 25% for testing.
  • Ablations include removing the trajectory‑aware scorer, removing LUPI, and replacing stratified selection with global Top‑K or clustering.
Numbers
  • MAE .045.004 vs .114.035 (PSN-IRT) on SWE-bench Lite
  • Kendall .836 vs .650 (PSN-IRT) on SWE-bench Lite
  • Spearman .950 vs .817 (PSN-IRT) on SWE-bench Lite
  • MAE .043.015 vs .114.035 (PSN-IRT) on SWE-bench Verified
  • Kendall .872 vs .786 (PSN-IRT) on SWE-bench Verified
  • Spearman .976 vs .941 (PSN-IRT) on SWE-bench Verified
Limitations

The paper does not evaluate PTA-IRT on domains outside software‑engineering benchmarks or in scenarios lacking historical execution trajectories.

Under low calibration budgets, PTA-IRT consistently outperforms prior IRT baselines on score and ranking recovery across four SWE benchmarks.Found in the source text, word for word.

Picked because: Introduces a trajectory‑aware evaluation framework for software‑engineering agents that reduces benchmarking cost while preserving fidelity, providing released benchmarks and tooling for rapid agent validation.

Paper 4 of 5

When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation

Peiying Zhu, Sidi Chang · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

The initial evaluation conflated guardrail interventions with changes to offer schemas and buyer choice procedures, so the reported welfare gains were scaffold-sensitive; simply fixing the schema revealed that the original effect was not due to the guardrail alone.

Approach

The authors introduce a construct‑validity contract that separates incentive validity, protocol isolation, stochastic stability, and welfare accounting. They audit the original study with four experiments: scaffold control (E1), repeated‑generation forensic analysis (E2), incentive manipulation check (E3), and scripted positive controls (E4). Each experiment isolates a component of the evaluation pipeline while keeping other factors fixed. The contract requires returning INVALID or INCONCLUSIVE before any substantive policy claim. This framework lets researchers pinpoint whether observed effects stem from the intended policy or from confounding protocol changes.

Result

When the schema and buyer chooser were held fixed, the paired welfare contrasts changed to +7.2, -13.9, and +23.8. The four largest 14B single‑generation effects averaged +229, but after averaging three generations per profile‑condition they averaged +37.6 with a 95% bootstrap interval of [-34.2, 109.3]; generation residuals explained 49.9% of the variation. The original estimate was deemed INVALID under protocol isolation, and the controlled study remained INCONCLUSIVE under incentive validity and stochastic stability.

Why it matters

Researchers and policymakers designing or evaluating LLM‑driven marketplaces should adopt the construct‑validity contract to ensure that reported welfare effects truly reflect the intended policy rather than confounding scaffolds.

Method details
  • Qwen2.5‑Instruct models at 1.5B, 3B, and 14B parameters, run in 4‑bit quantization
  • Inference temperature set to 0.2
  • 30 held‑out synthetic buyer profiles used in every model and policy cell (full panel 60 profiles, 43 negative utilities)
  • E1 scaffold control used 180 original dialogues and 360 controlled dialogues
  • E2 generated three repetitions per profile‑condition for four selected profiles (24 dialogues total)
  • E3 performed an 18‑dialogue smoke test comparing three seller prompts on three profiles with two generations each
Numbers
  • welfare gain, +87.4, across Qwen2.5 ladder
  • welfare gain, +35.0, across Qwen2.5 ladder
  • welfare gain, +28.8, across Qwen2.5 ladder
  • fixed‑schema contrast, +7.2, compared to original scaffold
  • fixed‑schema contrast, -13.9, compared to original scaffold
  • fixed‑schema contrast, +23.8, compared to original scaffold
  • single‑generation effect average, +229, for 14B models
  • three‑generation average, +37.6, with 95% CI [-34.2, 109.3]
  • generation residual variance, 49.9%, of total variation
  • original estimate status, INVALID, under protocol isolation
  • controlled study status, INCONCLUSIVE, under incentive validity and stochastic stability
Limitations

The audit is limited to a single instruction‑tuned model family, synthetic buyer profiles, and a small number of generations; it does not establish external validity or performance of cross‑family sellers.

The original estimate is INVALID under protocol isolation, while the controlled study remains INCONCLUSIVE under incentive validity and stochastic stability.Found in the source text, word for word.

Picked because: Audits LLM‑agent guardrails in market simulations, exposing construct‑validity failures and offering concrete diagnostics that engineers can apply to verify agent safety claims.

Paper 5 of 5

Can LLMs Design Video Coding Tools? A Case Study on Planar Mode

Yingwen Zhang, Meng Wang, Liqiang He and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.MM

Problem

Designing video coding tools is highly challenging because modifications to a tool tightly couple with downstream modules, syntax signaling, and encoder‑decoder matching, so simply replacing a hand‑crafted tool does not work.

Approach

The method runs an iterative generation‑and‑evaluation loop where an LLM is prompted to produce C++ code for the core Planar predictor, the code is compiled into the codec, and actual encoding runs compute BD‑rate feedback. Prompts contain a task description, historical feedback (design ideas, code, BD‑rate), strategy instructions (exploration or modification), an interface specification, and validity constraints. The feedback is used to rank candidates with an objective that adds a worst‑case regularization term, and the top‑k candidates are sampled as parents for the next prompt. This loop is applied to VVenC (faster preset) and to ECM with both mode replacement and insertion strategies.

Result

On the VVenC faster preset the LLM‑generated Planar predictor achieved 0.18% bitrate savings with only 0.4% complexity overhead, and both replacement and insertion strategies yielded coding gains in the low‑resolution ECM setting.

Why it matters

Video codec developers and researchers should care because the work shows that LLMs can automatically generate competitive coding tools, potentially expanding the design space beyond expert intuition.

Method details
  • LLM generates deterministic C++ code for xPredIntraPlanar_Core based on compact prompts.
  • Encoder used: Fraunhofer Versatile Video Encoder (VVenC) version 1.14.0 faster preset.
  • Evaluation metric: sequence‑level BD‑rate compared to the default Planar mode.
  • Parent pool size: top‑16 candidates are retained for subsequent prompts.
  • Candidates generated per iteration: 64.
  • Parallel evaluation on 256 AMD EPYC 9754 cores, one iteration takes about one hour.
Numbers
  • bitrate savings, 0.18%, vs default Planar mode
  • complexity overhead, 0.4%, vs default Planar mode
  • candidates evaluated before useful predictor, nearly 2000, in ECM experiment
  • parallel cores, 256, used for evaluation
  • iteration time, about one hour, per iteration on 256 cores
  • candidate pool size, top-16, selected for next generation
Limitations

The study does not demonstrate that predictors found at low resolution transfer well to higher resolutions, and the evaluation process remains computationally expensive.

achieving 0.18% bitrate savings with 0.4% complexity overhead on the standard benchmark.Found in the source text, word for word.

Picked because: Shows a generation‑and‑evaluation loop where LLMs design a video‑coding tool (Planar mode), delivering a reproducible case study and code that illustrates practical LLM‑agent tooling.