Xin He, Yanlin Wang, Mingwei Liu and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Existing repository-level benchmarks only check whether generated patches pass functional tests and ignore review-derived acceptance constraints, so they cannot measure whether a patch meets real-world code‑review requirements. Adding more functional tests does not capture these constraints because they are separate engineering intents not expressed by functional correctness alone.
Approach
SWE‑Gate builds a benchmark by first extracting atomic review suggestions from real pull‑request comments using LLM‑based processing. These constraint seeds are transferred into compatible repository contexts to create repair instances that include both functional and constraint tests. Each instance provides a mutant patch, a non‑compliant reference patch, and a gold patch that satisfies both test suites. Agents are evaluated under a common coding‑agent scaffold with either the constraint description visible (Constraint‑Provided) or omitted (Constraint‑Omitted). The pipeline validates patches in isolated containers to prevent leakage. Joint Success Rate (JSR) is used as the primary metric to capture full compliance.
Result
Across 303 instances, the agents produced 644 functionally successful patches but only 423 passed both functional and constraint tests, revealing 221 hidden failures. The best model, GPT‑5.5, achieved a functional success rate of 74.9% and a joint success rate of 52.8%, while its constraint‑following rate was 70.5%. Overall hidden failure rate was 34.3% of functional successes.
Why it matters
Developers of coding agents and benchmark designers should care because SWE‑Gate exposes a substantial gap between functional correctness and real‑world review compliance, guiding future improvements in agent capabilities.
Method details
Evaluated four LLM backends: GPT‑5.5, GPT‑5.4‑mini, DeepSeek‑V4‑Flash, GPT‑4o‑mini
Inference uses Mini‑SWE‑Agent with a maximum of 100 interaction steps per repair
Baseline comparison is functional‑only evaluation (Constraint‑Omitted) versus dual evaluation (Constraint‑Provided)
Ablation studies report the impact of providing the natural‑language constraint to the agent
Numbers
Functional Success Rate (FSR) GPT‑5.5: 74.9%
Constraint Following Rate (CFR) GPT‑5.5: 70.5%
Joint Success Rate (JSR) GPT‑5.5: 52.8%
Total functionally successful repairs: 644
Hidden failures: 221
Overall Hidden Failure Rate (HFR): 34.3%
Limitations
The benchmark currently covers only Python repositories and relies on executable constraint tests, so it cannot evaluate non‑executable review requirements.
functional-only evaluation overestimates agents’ ability to satisfy the full requirements of repository-level repair tasks.Found in the source text, word for word.
Picked because: Introduces SWE‑Gate, a released benchmark that adds review‑constraint evaluation for coding agents, giving engineers a concrete tool to assess real‑world software‑engineering usefulness.
Before this work AI agents could not interoperate because they were built on heterogeneous development frameworks, models, tool interfaces, protocols, and execution environments, and simply using existing transports does not provide a common communication layer.
Approach
The Natural Language Interaction Protocol (NLIP) defines a standards‑based application‑layer protocol that introduces a lightweight semantic message envelope. This envelope can be transmitted over existing transports such as HTTP/HTTPS, WebSocket, and AMQP. NLIP‑aware agents and gateways use the envelope to translate between clients, agents, local context stores, ontologies, tools, enterprise services, and heterogeneous underlying protocols. The design includes security‑by‑design considerations and a reference implementation. It also positions NLIP relative to emerging agent protocols such as MCP and A2A.
Result
The paper presents the motivation, design rationale, message model, transport bindings, security considerations, reference implementation, representative applications, adoption signals, and relationships to other protocols, demonstrating that NLIP has been standardized by Ecma International and is being adopted across companies and universities.
Why it matters
Organizations and developers building AI agents should care because NLIP offers a common protocol to enable interoperability across diverse systems.
Method details
Lightweight semantic message envelope for agent communication
Transport bindings support HTTP/HTTPS, WebSocket, and AMQP
Reference implementation is provided
Security‑by‑design considerations are incorporated
Relationship to emerging protocols MCP and A2A is described
Limitations
The paper does not state any explicit limitations.
NLIP provides a lightweight semantic message envelope that can be carried over existing transports such as HTTP/HTTPS, WebSocket, and AMQPFound in the source text, word for word.
Picked because: Presents the Natural Language Interaction Protocol (NLIP), an open standard with reference implementation for interoperable AI agents, directly applicable to building self‑hosted agent infrastructures.
Chihao Shen, Jiacheng Li, Aastha Mahajan and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CR
Problem
Existing evaluations only validate a patch by checking whether the provided PoC input still crashes, which lets agents either reproduce memorized historical patches or add superficial fixes that merely suppress the reported crash without fixing the root cause.
Approach
The authors introduce PatchBench, a benchmark that selects vulnerabilities whose developer patches lie off the sanitizer crash stack, transplants historical vulnerabilities into newer repository versions, and applies semantic-preserving code mutations. They also build a new validation pipeline that checks security (including sanitizer regression) and semantic correctness (output state and unit tests). By requiring agents to reason beyond the crash stack and preventing reuse of exact historical patches, PatchBench aims to provide a realistic evaluation of AI vulnerability patching.
Result
Across all agents the original PoC pass rate is 83.1% but only 45.3% of tasks pass the full security and semantic validation, indicating a 1.83× inflation. Even the strongest agents pass over 97% of PoCs yet solve only about 56‑59% of tasks. Semantic validation overall succeeds on 63.4% of patches, with sanitizer regression at 95.7% and output state and unit test checks at 75.5% and 79.9% respectively.
Why it matters
Researchers developing AI agents for automated vulnerability repair and benchmark designers should care because PatchBench reveals that PoC-only validation dramatically overestimates agent performance and highlights the need for deeper security and semantic evaluation.
Method details
Uses the ARVO dataset of reproducible C/C++ vulnerabilities as source of historical bugs.
Constructs 213 patching tasks by transplanting vulnerabilities and mutating code with NatGen's five transformations and CodeMorph.
Evaluates 11 state-of-the-art agents, including Codex + GPT-5.6 Sol, OpenHands + GPT-5.6 Sol, Claude Code + Claude Opus 4.8, and the AIxCC system Atlantis.
Measures performance with both the original PoC-only validation and the full PatchBench validation pipeline.
Reports total inference cost of approximately $6,500 for the evaluation.
Numbers
PoC pass rate, 83.1%, versus final solved rate 45.3%
Inflation factor, 1.83×, compared to PoC-only validation
Top agents pass over 97% of PoCs but solve 59.2% (Codex + GPT-5.6 Sol), 58.2% (OpenHands + GPT-5.6 Sol), 56.8% (Claude Code + Claude Opus 4.8)
Sanitizer regression pass rate, 95.7%, of semantic validation checks
Output state check pass rate, 75.5%, of semantic validation checks
Unit test check pass rate, 79.9%, of semantic validation checks
Limitations
The paper does not state any explicit limitations.
Averaged over all agents, 83.1% of the generated patches eliminate the original PoC crash, yet only 45.3% of the tasks are solved, i.e., pass both security and semantic validation, which shows a significant 1.83 inflation.Found in the source text, word for word.
Picked because: Provides PatchBench, an evaluation suite for AI‑driven vulnerability patching that includes reproducible artifacts engineers can use to validate security‑automation pipelines.
Tongyao Zhu, Wei Hern Lim, Min-Yen Kan · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Existing code‑editing LLMs often produce correct patches but rewrite far more code than necessary, a behavior called over‑editing, and simply increasing model size or reasoning budget does not eliminate it.
Approach
The authors create an evaluation framework by injecting controlled AST‑level corruptions into 400 BigCodeBench solutions, giving each task a known minimal patch. They compare generic prompts with an explicit preservation instruction that asks the model to keep as much original code as possible. They evaluate many frontier models and then explore post‑training methods: supervised fine‑tuning (SFT), reinforced fine‑tuning (rSFT), direct preference optimization (DPO), and reinforcement learning (RL) to learn minimal‑edit behavior. RL is found to give the best out‑of‑domain edit‑fidelity versus performance trade‑off. The study also examines reasoning vs non‑reasoning variants and scaling across model sizes.
Result
With the preservation instruction the average excess Levenshtein distance drops from 0.195 to 0.131, added cognitive complexity falls by 26.6%, and Pass@1 rises by 2.3 points. Reinforcement learning then achieves the best out‑of‑domain trade‑off, attaining Pass@1 of 0.782 while keeping excess edit distance low (0.050) and preserving broader coding ability (LCB 33.2%).
Why it matters
Researchers and practitioners building code‑repair tools should care because the work shows that over‑editing is a measurable, steerable problem and that minimal‑edit behavior can be learned without sacrificing overall correctness.
Method details
400 BigCodeBench problems with AST‑level corruptions are used as the benchmark
Frontier models include GPT‑5.5, Claude Opus 4.7, DeepSeek V3.2, Gemini, Qwen2.5‑Coder‑Instruct series (0.5B‑32B)
Two prompting styles: generic “Fix and complete my function.” and explicit “…but keep as much of the original code as possible”
Post‑training methods evaluated: SFT, rSFT, DPO, and RL on Qwen3‑4B‑Instruct‑2507
Baselines include the generic prompt and reasoning‑free variants; ablations cover preservation prompting, reasoning, and model‑size scaling
Numbers
excess Levenshtein distance reduced from 0.195 to 0.131
added cognitive complexity reduced by 26.6%
Pass@1 increased by 2.3 percentage points
RL out‑of‑domain Pass@1 0.782 vs SFT 0.458
RL out‑of‑domain excess Levenshtein 0.050 vs SFT 0.006
RL LCB 33.2 (0.6) compared to SFT 17.7 (14.9)
Limitations
The paper does not state any explicit limitations.
lowering average excess Levenshtein distance from 0.195 to 0.131, reducing added cognitive complexity by 26.6%, and increasing Pass@1 by 2.3 pointsFound in the source text, word for word.
Picked because: Studies over‑editing in code‑editing LLMs and releases an evaluation framework on 400 BigCodeBench tasks, offering actionable guidance for minimal, trustworthy code edits.
Liam Johnston, Shayan Noei, Maram Assi and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Existing automated issue labeling approaches require extensive manual effort to design taxonomies, assign generic labels, and depend on pre‑labeled datasets, which limits scalability and domain specificity.
Approach
LabelMate first prompts instruction‑tuned LLMs to generate candidate labels from historical issue reports, then refines them via semantic clustering using a text embedding model to produce a coherent project‑specific taxonomy. For new issues, a retrieval‑augmented generation (RAG) step selects the most relevant subset of taxonomy labels, and the same LLMs assign labels to the issue. A larger LLM (deepseek‑r1:70b) evaluates assignment accuracy. The pipeline thus avoids any pre‑labeled training data while adapting to each project.
Result
LabelMate produced a taxonomy of 275 labels and achieved an average labeling accuracy of 89.84%, which the authors report as a statistically significant improvement over existing generic label‑assigning approaches.
Why it matters
Practitioners and researchers who need domain‑adapted, automated issue labeling without collecting labeled data should consider LabelMate for more accurate and scalable labeling.