Qian Kou, Xiaofeng Shi, Xiaosong Qiu and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. The obvious fix of adding a retriever is not always viable due to latency, privacy, or deployment constraints.
Approach
IAR is a three‑stage post‑training framework. Stage 1 (Inject) converts source documents into continuation, rewrite, and instruction‑conditioned reconstruction objectives. Stage 2 (Align) adapts the injected model with answer‑only QA supervision. Stage 3 (Recover) merges the domain‑adapted checkpoint with the original instruction model using post‑hoc model merging (e.g., SLERP, task arithmetic, TIES, DARE). The stages are trained sequentially and the final checkpoint is selected by balancing domain QA accuracy against general benchmark retention. This separation isolates document exposure, answer alignment, and general‑capability recovery.
Result
IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset‑model settings. It yields an average gain of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. The method also retains strong general capability while boosting domain accuracy, outperforming extended CC baselines such as LoRA and FAPM on the overall frontier.
Why it matters
Practitioners who need retrieval‑free question answering-e.g., for low‑latency or privacy‑sensitive deployments-should consider IAR to internalize bounded corpora without sacrificing general language abilities.
Method details
Models include Llama‑3.2‑3B, Phi‑4‑mini, Qwen3‑4B, and SmolLM3‑3B (plus Qwen3‑8B/14B/32B scaling ablations)
Datasets are Common Corpus (CC) with 14,258 training and 750 test QA pairs and CCI with 10,926 training and 575 test pairs
Training uses a three‑stage pipeline: Inject, Align, then Recover with a fixed 12‑candidate merging grid (SLERP, task arithmetic, TIES, DARE)
Baselines compared against include Vanilla SFT, CPT+SFT, LoRA, FAPM, SDFT, Replay, and BudgetMatch QA‑only SFT
IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench.Found in the source text, word for word.
Picked because: Presents IAR, a three-stage post-training pipeline that lets a self-hosted LLM internalize a document collection for retrieval-free QA, with released code and scripts.
Rachna Raj, Benoit Baudry, Diego Elias Costa · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Client-side test suites often fail to detect breaking changes because they have limited library coverage and do not exercise all library methods used in the client codebase, so simply running existing tests does not reveal the regressions.
Approach
BreakGuard first performs static analysis to extract every client method that calls the target library (focal method) and its call sites, then uses a large language model to generate a test for each focal method; the generated test is run on the pre-breaking version and the breaking version, and a failure on the latter signals a breaking change. The pipeline combines import filtering, AST extraction with Spoon, focal method grouping, prompt construction with varying context levels, and LLM inference to produce the test suite.
Result
Using the best configuration (GPT-4o with class context) BreakGuard detected 30.3% of breaking changes (27 of 89) at an average cost of about $0.90 USD per detected change, and it was more reliable for crash-type breaking changes than for behavioral ones.
Why it matters
Developers maintaining Java dependencies and researchers studying LLM-based test generation should note that automated test synthesis can expose a subset of breaking changes but still requires complementary techniques for full coverage.
Method details
Three LLMs were evaluated: GPT-4o, Qwen3-coder-480B, and GPT-OSS-120B, all with temperature=0 and top_p=1.0
Dataset consisted of 89 real-world breaking changes from the BUMP benchmark after filtering from an original 571 instances
Static analysis used Spoon to build ASTs, extract library method invocations and type references, and group them by enclosing focal method
Three context variants (minimal, method, class) were tested; class context provided the best detection trade-off
Baseline comparison across the three LLMs and three context levels identified GPT-4o with class context as the best configuration
Numbers
Detection rate, 30.3%, of breaking changes (27 of 89) compared to total 89 BUMP instances
Mean cost, $0.90 USD, per detected breaking change
Dataset size, 89, breaking changes after filtering from original 571 BUMP instances
Library categories, 9, categories represented in the retained dataset
Limitations
The approach cannot reliably detect behavioral breaking changes and mainly catches runtime crashes, missing many non-crash regressions.
BreakGuard detects 30.3% of breaking changes (27 of 89) at a mean cost of roughly $0.90 USD per detected breaking change.Found in the source text, word for word.
Picked because: Introduces BreakGuard, a tool that generates LLM-based tests to automatically detect breaking API changes, offering a practical workflow for DevOps verification.
Atsuyuki Miyai, Kiyoharu Aizawa, Toshihiko Yamasaki · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Existing harness optimization evaluates a fixed validation set fully each iteration, incurring high cost and causing overfitting when the same subset is reused; simply fixing a subset does not work because it leads to overfitting, and naïve resampling makes raw subset scores incomparable across iterations.
Approach
Task-CoEvolve interleaves variance‑weighted task selection with sampling‑aware full‑set performance estimation. Phase 0 initializes history from two full‑set evaluations. Phase 1 computes a Bernoulli variance weight for each task and samples a new subset each iteration. Phase 2 estimates the full‑set score from the sampled tasks using either a Hájek estimator or an anchored‑difference estimator depending on the benchmark. Phase 3 selects the candidate with the highest estimated full‑set score. This loop repeats for each evolution iteration.
Result
Task‑CoEvolve attains the highest average test accuracy at both 7% and 20% validation budgets, reaching 47.6% with 7% budget (within 1 % of full‑set search) and 49.3% with 20% budget, surpassing all baselines while using far fewer evaluations.
Why it matters
Researchers seeking to optimize LLM harnesses can dramatically cut validation evaluation cost while preserving or improving downstream performance.
Method details
GPT‑OSS‑120B classifier LLM with temperature 0 is used for online text classification
Claude Opus 4.6 serves as the meta‑agent that writes harness code
Task‑CoEvolve average accuracy 49.3% at 20% budget vs full‑set search 48.6%
Task‑CoEvolve average accuracy 47.6% at 7% budget vs full‑set search 48.6%
Meta‑Harness Full Search average 48.6% vs Naive 45.2% at 7% budget
Terminal‑Bench GPT‑5.6‑Luna: Task‑CoEvolve 61.8% vs Full Search 62.9%
Limitations
The paper evaluates only online text classification and Terminal‑Bench 2.1, so generality to other task families is not demonstrated.
Task-CoEvolve consistently outperforms fixed-subset baselines and matches the final performance of full-set search while reducing the number of evaluations during optimization by 80%.Found in the source text, word for word.
Picked because: Describes Task-CoEvolve, an adaptive validation-task selection method that speeds up LLM agent harness optimization without retraining the model, with an open-source implementation.
Prior work assumed that technical documentation, designed for humans, would serve coding agents effectively, but there was no empirical evidence of how agents actually use documentation and why simply making docs more actionable or verifiable does not reliably improve agent behavior.
Approach
The authors conduct a behaviour‑grounded study by mining event logs from two public corpora (SWE‑chat and AIDev). They extract fine‑grained development events, apply a manually coded scheme to label documentation interactions, and use statistical models (cluster bootstrap and stage‑adjusted odds ratios) to quantify relationships between consultation and subsequent actions. From these traces they derive a descriptive two‑lobed cycle model of agent‑documentation interaction, contrasting it with the previously assumed linear journey.
Result
Agent‑facing artefacts (instruction files and working notes) dominate documentation interactions at 60.5% versus 1.3% for API references; the link between consultation and immediate code editing is ambiguous (adjacent probability 0.002, lift 1.05, adjusted OR 1.33); agents initiate documentation access 70.2% of the time and code changes precede documentation changes by a factor of 4.7.
Why it matters
Documentation designers and tool builders should prioritize agent‑instruction files and working notes, as these dominate agent interaction, while researchers should reconsider assumptions about actionability and verifiability for coding agents.
AIDev: 33,097 pull requests with 690,260 file‑level rows, of which 29,597 are documentation paths.
Statistical treatment includes adjacent transition probability (0.002), unadjusted lift (1.05 for code editing), and stage‑adjusted odds ratio (OR 1.33 [1.09,1.62]).
Interaction initiation: 70.2% of documentation interactions are self‑initiated.
Code‑first ordering: in pull requests changing both code and docs, code is touched first 4.7 × more often.
Numbers
agent‑facing docs, 60.5%, compared to classical technical docs 10.6% and API references 1.3%
unadjusted lift for code editing after consultation, 1.05
stage‑adjusted odds ratio for code editing after consultation, 1.33 [1.09,1.62]
self‑initiated documentation interactions, 70.2% of all interactions
code touched before documentation in pull requests, 4.7 × more often
Limitations
The observational design cannot establish causality or prove that improving actionability or verifiability will change agent behavior, and no validation events were observed.
Agents’ documentation work is dominated by agent-facing artefacts: instruction files and working notes account for 60.5% of all documentation interactions.Found in the source text, word for word.
Picked because: Provides an empirical study of how coding agents consume and produce documentation, releasing a dataset of 557 agentic sessions that can guide tooling for agent-friendly docs.
Reasoning language models trained with reinforcement learning typically operate under a fixed token budget, which leads to over-computation on easy problems and insufficient computation on difficult ones. Simply increasing the token budget uniformly does not address the need for adaptive allocation and wastes compute on easy instances.
Approach
The model selects, as the first token of its response, one of three modes: NoThink, Short, or Long. This choice is learned end-to-end with Group Relative Policy Optimization (GRPO) using a shaped reward that makes each mode beneficial at a different response length. Hard per-mode token caps keep the modes distinct, eliminating the need for a separate routing network. The three modes emerge during training without collapsing to a single choice. The router therefore sorts problems by difficulty rather than randomly.
Result
Averaged over three seeds, the adaptive policy stays close to the base model's accuracy on the held-out MATH500 (0.782 vs. 0.796) while cutting the mean response length from 4,796 to 2,811 tokens (a 41% reduction). It also achieves a 76% token reduction on GSM8K with higher accuracy than baselines at similar response length.
Why it matters
Researchers building reasoning language models and RL fine‑tuning pipelines should care because the method provides a simple way to allocate compute adaptively, improving efficiency without sacrificing accuracy.
Method details
1.5B distilled model trained on the MATH dataset
Uses Group Relative Policy Optimization (GRPO) for policy learning
Three response modes: NoThink, Short, Long with hard token caps
Evaluation on held-out MATH500 benchmark
Baseline comparison to the base model without adaptive routing
Transfer evaluation on GSM8K benchmark without retraining
Numbers
accuracy, 0.782 vs. 0.796 base
mean response length, 4,796 to 2,811 tokens
token reduction, 41% reduction
token reduction on GSM8K, 76% reduction
model size, 1.5B parameters
seeds, three
Limitations
The paper does not discuss any limitations.
the resulting policy stays close to the base model's accuracy on the held-out MATH500 ($0.782$ vs.\ $0.796$) while cutting the mean response length from $4{,}796$ to $2{,}811$ tokens (a $41\%$ reduction).Found in the source text, word for word.
Picked because: Shows how a language model can learn to choose between reasoning modes at inference time to allocate compute, offering a concrete technique for runtime resource management.