arXiv digest

Friday

September 4, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Xin He, Yanlin Wang, Mingwei Liu and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Existing repository-level benchmarks only check whether generated patches pass functional tests and ignore review-derived acceptance constraints, so they cannot measure whether a patch meets real-world code‑review requirements. Adding more functional tests does not capture these constraints because they are separate engineering intents not expressed by functional correctness alone.

Approach

SWE‑Gate builds a benchmark by first extracting atomic review suggestions from real pull‑request comments using LLM‑based processing. These constraint seeds are transferred into compatible repository contexts to create repair instances that include both functional and constraint tests. Each instance provides a mutant patch, a non‑compliant reference patch, and a gold patch that satisfies both test suites. Agents are evaluated under a common coding‑agent scaffold with either the constraint description visible (Constraint‑Provided) or omitted (Constraint‑Omitted). The pipeline validates patches in isolated containers to prevent leakage. Joint Success Rate (JSR) is used as the primary metric to capture full compliance.

Result

Across 303 instances, the agents produced 644 functionally successful patches but only 423 passed both functional and constraint tests, revealing 221 hidden failures. The best model, GPT‑5.5, achieved a functional success rate of 74.9% and a joint success rate of 52.8%, while its constraint‑following rate was 70.5%. Overall hidden failure rate was 34.3% of functional successes.

Why it matters

Developers of coding agents and benchmark designers should care because SWE‑Gate exposes a substantial gap between functional correctness and real‑world review compliance, guiding future improvements in agent capabilities.

Method details
  • Evaluated four LLM backends: GPT‑5.5, GPT‑5.4‑mini, DeepSeek‑V4‑Flash, GPT‑4o‑mini
  • Dataset comprises 303 repository‑level repair instances from 75 open‑source Python repositories
  • Inference uses Mini‑SWE‑Agent with a maximum of 100 interaction steps per repair
  • Baseline comparison is functional‑only evaluation (Constraint‑Omitted) versus dual evaluation (Constraint‑Provided)
  • Ablation studies report the impact of providing the natural‑language constraint to the agent
Numbers
  • Functional Success Rate (FSR) GPT‑5.5: 74.9%
  • Constraint Following Rate (CFR) GPT‑5.5: 70.5%
  • Joint Success Rate (JSR) GPT‑5.5: 52.8%
  • Total functionally successful repairs: 644
  • Hidden failures: 221
  • Overall Hidden Failure Rate (HFR): 34.3%
Limitations

The benchmark currently covers only Python repositories and relies on executable constraint tests, so it cannot evaluate non‑executable review requirements.

functional-only evaluation overestimates agents’ ability to satisfy the full requirements of repository-level repair tasks.Found in the source text, word for word.

Picked because: Introduces SWE‑Gate, a released benchmark that adds review‑constraint evaluation for coding agents, giving engineers a concrete tool to assess real‑world software‑engineering usefulness.

Paper 2 of 5

The Natural Language Interaction Protocol and Standard for AI Agents

Luyi Xing, Rasit Onur Topaloglu, Ranjan Sinha and 9 others · abstract · pdf

quote verifiedfigures checkedread: abstract onlycs.AI

Problem

Before this work AI agents could not interoperate because they were built on heterogeneous development frameworks, models, tool interfaces, protocols, and execution environments, and simply using existing transports does not provide a common communication layer.

Approach

The Natural Language Interaction Protocol (NLIP) defines a standards‑based application‑layer protocol that introduces a lightweight semantic message envelope. This envelope can be transmitted over existing transports such as HTTP/HTTPS, WebSocket, and AMQP. NLIP‑aware agents and gateways use the envelope to translate between clients, agents, local context stores, ontologies, tools, enterprise services, and heterogeneous underlying protocols. The design includes security‑by‑design considerations and a reference implementation. It also positions NLIP relative to emerging agent protocols such as MCP and A2A.

Result

The paper presents the motivation, design rationale, message model, transport bindings, security considerations, reference implementation, representative applications, adoption signals, and relationships to other protocols, demonstrating that NLIP has been standardized by Ecma International and is being adopted across companies and universities.

Why it matters

Organizations and developers building AI agents should care because NLIP offers a common protocol to enable interoperability across diverse systems.

Method details
  • Lightweight semantic message envelope for agent communication
  • Transport bindings support HTTP/HTTPS, WebSocket, and AMQP
  • Reference implementation is provided
  • Security‑by‑design considerations are incorporated
  • Relationship to emerging protocols MCP and A2A is described
Limitations

The paper does not state any explicit limitations.

NLIP provides a lightweight semantic message envelope that can be carried over existing transports such as HTTP/HTTPS, WebSocket, and AMQPFound in the source text, word for word.

Picked because: Presents the Natural Language Interaction Protocol (NLIP), an open standard with reference implementation for interoperable AI agents, directly applicable to building self‑hosted agent infrastructures.

Paper 3 of 5

PatchBench: Evaluating AI Agents for Vulnerability Patching

Chihao Shen, Jiacheng Li, Aastha Mahajan and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Existing evaluations only validate a patch by checking whether the provided PoC input still crashes, which lets agents either reproduce memorized historical patches or add superficial fixes that merely suppress the reported crash without fixing the root cause.

Approach

The authors introduce PatchBench, a benchmark that selects vulnerabilities whose developer patches lie off the sanitizer crash stack, transplants historical vulnerabilities into newer repository versions, and applies semantic-preserving code mutations. They also build a new validation pipeline that checks security (including sanitizer regression) and semantic correctness (output state and unit tests). By requiring agents to reason beyond the crash stack and preventing reuse of exact historical patches, PatchBench aims to provide a realistic evaluation of AI vulnerability patching.

Result

Across all agents the original PoC pass rate is 83.1% but only 45.3% of tasks pass the full security and semantic validation, indicating a 1.83× inflation. Even the strongest agents pass over 97% of PoCs yet solve only about 56‑59% of tasks. Semantic validation overall succeeds on 63.4% of patches, with sanitizer regression at 95.7% and output state and unit test checks at 75.5% and 79.9% respectively.

Why it matters

Researchers developing AI agents for automated vulnerability repair and benchmark designers should care because PatchBench reveals that PoC-only validation dramatically overestimates agent performance and highlights the need for deeper security and semantic evaluation.

Method details
  • Uses the ARVO dataset of reproducible C/C++ vulnerabilities as source of historical bugs.
  • Constructs 213 patching tasks by transplanting vulnerabilities and mutating code with NatGen's five transformations and CodeMorph.
  • Evaluates 11 state-of-the-art agents, including Codex + GPT-5.6 Sol, OpenHands + GPT-5.6 Sol, Claude Code + Claude Opus 4.8, and the AIxCC system Atlantis.
  • Measures performance with both the original PoC-only validation and the full PatchBench validation pipeline.
  • Reports total inference cost of approximately $6,500 for the evaluation.
Numbers
  • PoC pass rate, 83.1%, versus final solved rate 45.3%
  • Inflation factor, 1.83×, compared to PoC-only validation
  • Top agents pass over 97% of PoCs but solve 59.2% (Codex + GPT-5.6 Sol), 58.2% (OpenHands + GPT-5.6 Sol), 56.8% (Claude Code + Claude Opus 4.8)
  • Sanitizer regression pass rate, 95.7%, of semantic validation checks
  • Output state check pass rate, 75.5%, of semantic validation checks
  • Unit test check pass rate, 79.9%, of semantic validation checks
Limitations

The paper does not state any explicit limitations.

Averaged over all agents, 83.1% of the generated patches eliminate the original PoC crash, yet only 45.3% of the tasks are solved, i.e., pass both security and semantic validation, which shows a significant 1.83 inflation.Found in the source text, word for word.

Picked because: Provides PatchBench, an evaluation suite for AI‑driven vulnerability patching that includes reproducible artifacts engineers can use to validate security‑automation pipelines.

Paper 4 of 5

When Models Edit Too Much: On the Fidelity of Minimal Code Edits

Tongyao Zhu, Wei Hern Lim, Min-Yen Kan · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Existing code‑editing LLMs often produce correct patches but rewrite far more code than necessary, a behavior called over‑editing, and simply increasing model size or reasoning budget does not eliminate it.

Approach

The authors create an evaluation framework by injecting controlled AST‑level corruptions into 400 BigCodeBench solutions, giving each task a known minimal patch. They compare generic prompts with an explicit preservation instruction that asks the model to keep as much original code as possible. They evaluate many frontier models and then explore post‑training methods: supervised fine‑tuning (SFT), reinforced fine‑tuning (rSFT), direct preference optimization (DPO), and reinforcement learning (RL) to learn minimal‑edit behavior. RL is found to give the best out‑of‑domain edit‑fidelity versus performance trade‑off. The study also examines reasoning vs non‑reasoning variants and scaling across model sizes.

Result

With the preservation instruction the average excess Levenshtein distance drops from 0.195 to 0.131, added cognitive complexity falls by 26.6%, and Pass@1 rises by 2.3 points. Reinforcement learning then achieves the best out‑of‑domain trade‑off, attaining Pass@1 of 0.782 while keeping excess edit distance low (0.050) and preserving broader coding ability (LCB 33.2%).

Why it matters

Researchers and practitioners building code‑repair tools should care because the work shows that over‑editing is a measurable, steerable problem and that minimal‑edit behavior can be learned without sacrificing overall correctness.

Method details
  • 400 BigCodeBench problems with AST‑level corruptions are used as the benchmark
  • Frontier models include GPT‑5.5, Claude Opus 4.7, DeepSeek V3.2, Gemini, Qwen2.5‑Coder‑Instruct series (0.5B‑32B)
  • Two prompting styles: generic “Fix and complete my function.” and explicit “…but keep as much of the original code as possible”
  • Post‑training methods evaluated: SFT, rSFT, DPO, and RL on Qwen3‑4B‑Instruct‑2507
  • Baselines include the generic prompt and reasoning‑free variants; ablations cover preservation prompting, reasoning, and model‑size scaling
Numbers
  • excess Levenshtein distance reduced from 0.195 to 0.131
  • added cognitive complexity reduced by 26.6%
  • Pass@1 increased by 2.3 percentage points
  • RL out‑of‑domain Pass@1 0.782 vs SFT 0.458
  • RL out‑of‑domain excess Levenshtein 0.050 vs SFT 0.006
  • RL LCB 33.2 (0.6) compared to SFT 17.7 (14.9)
Limitations

The paper does not state any explicit limitations.

lowering average excess Levenshtein distance from 0.195 to 0.131, reducing added cognitive complexity by 26.6%, and increasing Pass@1 by 2.3 pointsFound in the source text, word for word.

Picked because: Studies over‑editing in code‑editing LLMs and releases an evaluation framework on 400 BigCodeBench tasks, offering actionable guidance for minimal, trustworthy code edits.

Paper 5 of 5

LabelMate: An LLM-Driven Framework for Refined Issue Report Labeling

Liam Johnston, Shayan Noei, Maram Assi and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Existing automated issue labeling approaches require extensive manual effort to design taxonomies, assign generic labels, and depend on pre‑labeled datasets, which limits scalability and domain specificity.

Approach

LabelMate first prompts instruction‑tuned LLMs to generate candidate labels from historical issue reports, then refines them via semantic clustering using a text embedding model to produce a coherent project‑specific taxonomy. For new issues, a retrieval‑augmented generation (RAG) step selects the most relevant subset of taxonomy labels, and the same LLMs assign labels to the issue. A larger LLM (deepseek‑r1:70b) evaluates assignment accuracy. The pipeline thus avoids any pre‑labeled training data while adapting to each project.

Result

LabelMate produced a taxonomy of 275 labels and achieved an average labeling accuracy of 89.84%, which the authors report as a statistically significant improvement over existing generic label‑assigning approaches.

Why it matters

Practitioners and researchers who need domain‑adapted, automated issue labeling without collecting labeled data should consider LabelMate for more accurate and scalable labeling.

Method details
  • Label Assigner LLMs: gemma-2-9b-it, Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct (7‑9B parameters each).
  • Label Evaluator LLM: deepseek-r1:70b (70 billion parameters).
  • Text Embedding Model: all-mpnet-base-v2 for semantic clustering and retrieval.
  • Dataset: 16,500 issue reports from 30 popular GitHub repositories, split 80%/20% into 13,210 training and 3,290 test reports.
  • RAG mechanism retrieves a relevant label subset before assignment.
Numbers
  • 275 labels, generated, compared to generic label sets
  • 89.84% labeling accuracy, achieved, compared to generic approaches
  • 16,500 issue reports, total, from 30 repositories
  • 13,210 training reports, used for taxonomy construction, 80% of data
  • 3,290 test reports, used for evaluation, 20% of data
  • 7‑9B parameters, size of assigner LLMs, compared to larger models
Limitations

The paper does not discuss any limitations of the proposed approach.

our approach generates a coherent list of 275 labels and achieves an average labeling accuracy of 89.84%Found in the source text, word for word.

Picked because: Offers LabelMate, an LLM‑driven system for automated issue‑report labeling with released code, addressing a common DevOps workflow.