arXiv digest

Saturday

August 22, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Qian Kou, Xiaofeng Shi, Xiaosong Qiu and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. The obvious fix of adding a retriever is not always viable due to latency, privacy, or deployment constraints.

Approach

IAR is a three‑stage post‑training framework. Stage 1 (Inject) converts source documents into continuation, rewrite, and instruction‑conditioned reconstruction objectives. Stage 2 (Align) adapts the injected model with answer‑only QA supervision. Stage 3 (Recover) merges the domain‑adapted checkpoint with the original instruction model using post‑hoc model merging (e.g., SLERP, task arithmetic, TIES, DARE). The stages are trained sequentially and the final checkpoint is selected by balancing domain QA accuracy against general benchmark retention. This separation isolates document exposure, answer alignment, and general‑capability recovery.

Result

IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset‑model settings. It yields an average gain of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. The method also retains strong general capability while boosting domain accuracy, outperforming extended CC baselines such as LoRA and FAPM on the overall frontier.

Why it matters

Practitioners who need retrieval‑free question answering-e.g., for low‑latency or privacy‑sensitive deployments-should consider IAR to internalize bounded corpora without sacrificing general language abilities.

Method details
  • Models include Llama‑3.2‑3B, Phi‑4‑mini, Qwen3‑4B, and SmolLM3‑3B (plus Qwen3‑8B/14B/32B scaling ablations)
  • Datasets are Common Corpus (CC) with 14,258 training and 750 test QA pairs and CCI with 10,926 training and 575 test pairs
  • Training uses a three‑stage pipeline: Inject, Align, then Recover with a fixed 12‑candidate merging grid (SLERP, task arithmetic, TIES, DARE)
  • Baselines compared against include Vanilla SFT, CPT+SFT, LoRA, FAPM, SDFT, Replay, and BudgetMatch QA‑only SFT
  • Ablations examine stage‑level effect, token‑budget matched QA‑only SFT, pre‑recovery domain internalization, and Qwen scaling
Numbers
  • domain QA accuracy, +3.6 percentage points, over Vanilla SFT
  • mean general performance, +12.1 percentage points, over Vanilla SFT
  • settings improved, 7 of 8 dataset‑model settings, over Vanilla SFT
  • evaluated model‑answer instances, 242,255, evaluation size
Limitations

The paper does not state explicit limitations.

IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench.Found in the source text, word for word.

Picked because: Presents IAR, a three-stage post-training pipeline that lets a self-hosted LLM internalize a document collection for retrieval-free QA, with released code and scripts.

Paper 2 of 5

BreakGuard: Towards Detecting Dependency Breaking Changes with LLM-Generated Tests

Rachna Raj, Benoit Baudry, Diego Elias Costa · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Client-side test suites often fail to detect breaking changes because they have limited library coverage and do not exercise all library methods used in the client codebase, so simply running existing tests does not reveal the regressions.

Approach

BreakGuard first performs static analysis to extract every client method that calls the target library (focal method) and its call sites, then uses a large language model to generate a test for each focal method; the generated test is run on the pre-breaking version and the breaking version, and a failure on the latter signals a breaking change. The pipeline combines import filtering, AST extraction with Spoon, focal method grouping, prompt construction with varying context levels, and LLM inference to produce the test suite.

Result

Using the best configuration (GPT-4o with class context) BreakGuard detected 30.3% of breaking changes (27 of 89) at an average cost of about $0.90 USD per detected change, and it was more reliable for crash-type breaking changes than for behavioral ones.

Why it matters

Developers maintaining Java dependencies and researchers studying LLM-based test generation should note that automated test synthesis can expose a subset of breaking changes but still requires complementary techniques for full coverage.

Method details
  • Three LLMs were evaluated: GPT-4o, Qwen3-coder-480B, and GPT-OSS-120B, all with temperature=0 and top_p=1.0
  • Dataset consisted of 89 real-world breaking changes from the BUMP benchmark after filtering from an original 571 instances
  • Static analysis used Spoon to build ASTs, extract library method invocations and type references, and group them by enclosing focal method
  • Three context variants (minimal, method, class) were tested; class context provided the best detection trade-off
  • Baseline comparison across the three LLMs and three context levels identified GPT-4o with class context as the best configuration
Numbers
  • Detection rate, 30.3%, of breaking changes (27 of 89) compared to total 89 BUMP instances
  • Mean cost, $0.90 USD, per detected breaking change
  • Dataset size, 89, breaking changes after filtering from original 571 BUMP instances
  • LLM count, 3, models evaluated (GPT-4o, Qwen3-coder-480B, GPT-OSS-120B)
  • Context levels, 3, (minimal, method, class) evaluated
  • Library categories, 9, categories represented in the retained dataset
Limitations

The approach cannot reliably detect behavioral breaking changes and mainly catches runtime crashes, missing many non-crash regressions.

BreakGuard detects 30.3% of breaking changes (27 of 89) at a mean cost of roughly $0.90 USD per detected breaking change.Found in the source text, word for word.

Picked because: Introduces BreakGuard, a tool that generates LLM-based tests to automatically detect breaking API changes, offering a practical workflow for DevOps verification.

Paper 3 of 5

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

Atsuyuki Miyai, Kiyoharu Aizawa, Toshihiko Yamasaki · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Existing harness optimization evaluates a fixed validation set fully each iteration, incurring high cost and causing overfitting when the same subset is reused; simply fixing a subset does not work because it leads to overfitting, and naïve resampling makes raw subset scores incomparable across iterations.

Approach

Task-CoEvolve interleaves variance‑weighted task selection with sampling‑aware full‑set performance estimation. Phase 0 initializes history from two full‑set evaluations. Phase 1 computes a Bernoulli variance weight for each task and samples a new subset each iteration. Phase 2 estimates the full‑set score from the sampled tasks using either a Hájek estimator or an anchored‑difference estimator depending on the benchmark. Phase 3 selects the candidate with the highest estimated full‑set score. This loop repeats for each evolution iteration.

Result

Task‑CoEvolve attains the highest average test accuracy at both 7% and 20% validation budgets, reaching 47.6% with 7% budget (within 1 % of full‑set search) and 49.3% with 20% budget, surpassing all baselines while using far fewer evaluations.

Why it matters

Researchers seeking to optimize LLM harnesses can dramatically cut validation evaluation cost while preserving or improving downstream performance.

Method details
  • GPT‑OSS‑120B classifier LLM with temperature 0 is used for online text classification
  • Claude Opus 4.6 serves as the meta‑agent that writes harness code
  • Datasets: LawBench (215 classes), Symptom2Disease (22 classes), USPTO‑50k (180 classes) and Terminal‑Bench 2.1 (89 tasks)
  • 20 evolution iterations with three candidates per iteration (60 candidates total)
  • Baselines: Meta‑Harness Full Search, Naive fixed‑subset, Random‑Resample, Rotation
  • Ablations test the impact of subset rotation, full‑set estimation, and variance‑weighted selection
Numbers
  • Evaluation cost reduced by 80% compared with full‑set search
  • Full‑set search uses 7,800 validation evaluations; Task‑CoEvolve uses 480 evaluations at 7% budget (16 × fewer)
  • Task‑CoEvolve average accuracy 49.3% at 20% budget vs full‑set search 48.6%
  • Task‑CoEvolve average accuracy 47.6% at 7% budget vs full‑set search 48.6%
  • Meta‑Harness Full Search average 48.6% vs Naive 45.2% at 7% budget
  • Terminal‑Bench GPT‑5.6‑Luna: Task‑CoEvolve 61.8% vs Full Search 62.9%
Limitations

The paper evaluates only online text classification and Terminal‑Bench 2.1, so generality to other task families is not demonstrated.

Task-CoEvolve consistently outperforms fixed-subset baselines and matches the final performance of full-set search while reducing the number of evaluations during optimization by 80%.Found in the source text, word for word.

Picked because: Describes Task-CoEvolve, an adaptive validation-task selection method that speeds up LLM agent harness optimization without retraining the model, with an open-source implementation.

Paper 4 of 5

From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation

Zhijun Gao, Jing Chen · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Prior work assumed that technical documentation, designed for humans, would serve coding agents effectively, but there was no empirical evidence of how agents actually use documentation and why simply making docs more actionable or verifiable does not reliably improve agent behavior.

Approach

The authors conduct a behaviour‑grounded study by mining event logs from two public corpora (SWE‑chat and AIDev). They extract fine‑grained development events, apply a manually coded scheme to label documentation interactions, and use statistical models (cluster bootstrap and stage‑adjusted odds ratios) to quantify relationships between consultation and subsequent actions. From these traces they derive a descriptive two‑lobed cycle model of agent‑documentation interaction, contrasting it with the previously assumed linear journey.

Result

Agent‑facing artefacts (instruction files and working notes) dominate documentation interactions at 60.5% versus 1.3% for API references; the link between consultation and immediate code editing is ambiguous (adjacent probability 0.002, lift 1.05, adjusted OR 1.33); agents initiate documentation access 70.2% of the time and code changes precede documentation changes by a factor of 4.7.

Why it matters

Documentation designers and tool builders should prioritize agent‑instruction files and working notes, as these dominate agent interaction, while researchers should reconsider assumptions about actionability and verifiability for coding agents.

Method details
  • SWE‑chat: 557 sampled sessions containing 94,813 events and 3,033 documentation interactions (3.2%).
  • AIDev: 33,097 pull requests with 690,260 file‑level rows, of which 29,597 are documentation paths.
  • Statistical treatment includes adjacent transition probability (0.002), unadjusted lift (1.05 for code editing), and stage‑adjusted odds ratio (OR 1.33 [1.09,1.62]).
  • Interaction initiation: 70.2% of documentation interactions are self‑initiated.
  • Code‑first ordering: in pull requests changing both code and docs, code is touched first 4.7 × more often.
Numbers
  • agent‑facing docs, 60.5%, compared to classical technical docs 10.6% and API references 1.3%
  • adjacent transition probability, 0.002, indicating low immediate coupling
  • unadjusted lift for code editing after consultation, 1.05
  • stage‑adjusted odds ratio for code editing after consultation, 1.33 [1.09,1.62]
  • self‑initiated documentation interactions, 70.2% of all interactions
  • code touched before documentation in pull requests, 4.7 × more often
Limitations

The observational design cannot establish causality or prove that improving actionability or verifiability will change agent behavior, and no validation events were observed.

Agents’ documentation work is dominated by agent-facing artefacts: instruction files and working notes account for 60.5% of all documentation interactions.Found in the source text, word for word.

Picked because: Provides an empirical study of how coding agents consume and produce documentation, releasing a dataset of 557 agentic sessions that can guide tooling for agent-friendly docs.

Paper 5 of 5

Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

Gijs Kassenaar, Zhao Yang, Vincent François-Lavet · abstract · pdf

quote verifiedfigures checkedread: abstract onlycs.AI

Problem

Reasoning language models trained with reinforcement learning typically operate under a fixed token budget, which leads to over-computation on easy problems and insufficient computation on difficult ones. Simply increasing the token budget uniformly does not address the need for adaptive allocation and wastes compute on easy instances.

Approach

The model selects, as the first token of its response, one of three modes: NoThink, Short, or Long. This choice is learned end-to-end with Group Relative Policy Optimization (GRPO) using a shaped reward that makes each mode beneficial at a different response length. Hard per-mode token caps keep the modes distinct, eliminating the need for a separate routing network. The three modes emerge during training without collapsing to a single choice. The router therefore sorts problems by difficulty rather than randomly.

Result

Averaged over three seeds, the adaptive policy stays close to the base model's accuracy on the held-out MATH500 (0.782 vs. 0.796) while cutting the mean response length from 4,796 to 2,811 tokens (a 41% reduction). It also achieves a 76% token reduction on GSM8K with higher accuracy than baselines at similar response length.

Why it matters

Researchers building reasoning language models and RL fine‑tuning pipelines should care because the method provides a simple way to allocate compute adaptively, improving efficiency without sacrificing accuracy.

Method details
  • 1.5B distilled model trained on the MATH dataset
  • Uses Group Relative Policy Optimization (GRPO) for policy learning
  • Three response modes: NoThink, Short, Long with hard token caps
  • Evaluation on held-out MATH500 benchmark
  • Baseline comparison to the base model without adaptive routing
  • Transfer evaluation on GSM8K benchmark without retraining
Numbers
  • accuracy, 0.782 vs. 0.796 base
  • mean response length, 4,796 to 2,811 tokens
  • token reduction, 41% reduction
  • token reduction on GSM8K, 76% reduction
  • model size, 1.5B parameters
  • seeds, three
Limitations

The paper does not discuss any limitations.

the resulting policy stays close to the base model's accuracy on the held-out MATH500 ($0.782$ vs.\ $0.796$) while cutting the mean response length from $4{,}796$ to $2{,}811$ tokens (a $41\%$ reduction).Found in the source text, word for word.

Picked because: Shows how a language model can learn to choose between reasoning modes at inference time to allocate compute, offering a concrete technique for runtime resource management.