arXiv digest

Tuesday

August 25, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Deyao Hong, Yizhe Chi, Wenyi Li and 7 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Existing benchmarks only check behavioural correctness, allowing agents to cheat by copying the original implementation (Blindness) without actually performing the migration.

Approach

The paper introduces SWE Refactor Bench, a set of 20 whole‑repository migration tasks evaluated with a three‑stage protocol: Migration Audit verifies that the migration truly occurred, Behavioural Tests run a fixed test suite to ensure functional parity, and Agentic Verification employs six independent coding agents to generate targeted tests for hidden behavioural differences. The Migration Audit uses a judge model (gpt‑5.6‑sol) with majority voting over three samples. The fixed test suite records checks from the original repository. The six verification agents each run for one hour to find counter‑examples. This pipeline separates migration completeness from behavioural correctness.

Result

Only 28 of 520 runs (5.4%) pass all three stages, and 13 of the 20 tasks receive no accepted solution; the top model, claude‑opus‑5, attains a composite score of 47.0/100. Among the 340 runs that pass Migration Audit, 58% reach 99% of the fixed checks while only 26% reach 100%.

Why it matters

Researchers and engineers building coding agents should use SWE Refactor Bench to gauge and improve migration capability beyond mere behavioural correctness.

Method details
  • Eight frontier models evaluated: claude‑opus‑5, claude‑sonnet‑5, gpt‑5.6‑luna, gpt‑5.6‑sol, kimi‑k3, qwen3.8‑max, dsv4‑flash, glm‑5.2
  • Each model runs on all 20 tasks, yielding 520 runs across 26 model‑effort configurations
  • Two GPT‑series models use Codex as their harness; the other six use Claude Code
  • Three‑stage evaluation: Migration Audit (judge gpt‑5.6‑sol with majority voting), Behavioural Tests (fixed suite), Agentic Verification (six independent coding agents, one hour each)
  • Best composite score achieved by claude‑opus‑5: 47.0/100
  • Category scores: build toolchain rewrites 31.4, language rewrites 5.6
Numbers
  • Accepted runs, 28, 5.4% of 520 runs
  • Best model score, 47.0/100, claude‑opus‑5
  • Unsolved tasks, 13, out of 20 total tasks
  • Migration completeness 99% reach, 58%, of runs that passed Migration Audit
  • Full correctness 100% reach, 26%, of runs that passed Migration Audit
  • Category score build toolchain rewrites, 31.4, compared to language rewrites 5.6
Limitations

The paper does not establish that any coding agent can reliably achieve a perfect whole‑repository migration; it only measures current frontier models.

only 28 of 520 runs ($5.4\%$) pass all three stagesFound in the source text, word for word.

Picked because: Introduces the SWE Refactor Bench, a concrete dataset and evaluation suite for coding agents to perform whole‑repository stack migrations, directly applicable to large‑scale software engineering automation.

Paper 2 of 5

Prime Agent: A Self-Improving RLM Harness

Seth Karten, Alex L. Zhang, Kevin Thomas and 8 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Prior approaches relied on the LLM’s weights and active token context, which cannot hold the external information and compute needed for long‑horizon agency; simply increasing context size or fine‑tuning does not provide persistent state or test‑time compute.

Approach

Prime Agent introduces a persistent IPython REPL (L2) and a Continual Harness (L3) that store histories, memories, skills, and subagent specifications across trajectories. Recursive subagents communicate directly with each other and with humans via an Agents View. Information management decides what state enters a model invocation, while computation management maps model‑selected actions to code, tools, and subagent sessions. The system separates information and computation management, uses hierarchical levels L0‑L3, and records all calls for accounting and recovery.

Result

Prime Agent lifts ARC‑AGI‑3 RHAE Best@1 from 30% to 95.5% and matches or exceeds native and popular harnesses on long‑context coding, GPU‑kernel generation, emulator construction, and autonomous nanoGPT speedruns. It also sustains an 85.5‑hour nanoGPT run with 19 validated records and supports continuous technology progression in Factorio and long‑horizon exploration in MazeBench.

Why it matters

Researchers building long‑horizon autonomous agents and evaluation harnesses should care because Prime Agent demonstrates how a standardized, expressive substrate can dramatically improve performance and enable persistent, recursive computation.

Method details
  • Evaluated on ARC‑AGI‑3 benchmark
  • Compared against Pi, Claude Code, Codex, Hermes Agent, OpenCode, and Kimi‑Code
  • Performed a 85.5‑hour autonomous nanoGPT speedrun with 19 validated records
  • Tested on Factorio and MazeBench environments
  • Uses a persistent IPython REPL and recursive subagents as core components
Numbers
  • ARC‑AGI‑3 RHAE Best@1 95.5% (up from 30%)
  • 85.5‑hour nanoGPT run
  • 19 validated records
Limitations

The paper notes that models still experience friction when allocating subagents, managing retained information, and refining reusable state, and many harness capabilities remain underused because current models were not trained to operate them.

Prime Agent raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5% and matches or exceeds native and popular harnesses across long-context coding, GPU-kernel generation, emulator construction, and autonomous nanoGPT speedruns.Found in the source text, word for word.

Picked because: Provides Prime Agent, an open‑source harness with a persistent REPL and continual compute extensions for building, testing, and verifying long‑horizon LLM agents in self‑hosted environments.

Paper 3 of 5

Robustness of Anomaly Detection Models for Industrial Control Systems under Training-Time Data Contamination

Mustafa Umut Ozbek, Taiwo Ojo, Pooria Madani and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Prior work on offline industrial control system anomaly detection assumes the training data are trustworthy, so it does not address the degradation caused by training-time data contamination; simply using clean validation and test sets does not prevent the models from being poisoned during training.

Approach

The paper evaluates the robustness of eleven heterogeneous anomaly detectors on the Secure Water Treatment (SWaT) benchmark under three contamination strategies (random injection, similarity‑targeted injection, and feature‑noise injection) with budgets from 1% to 10%. Detectors are trained on a normal‑only pool, calibrated on a clean validation split, and tested on an untouched test split. Contamination either inserts attack samples into the nominal training pool or adds bounded Gaussian noise to selected normal samples. The study uses a unified offline protocol and reports performance degradation across models. Results are compared against clean‑data baselines to assess model‑dependent robustness.

Result

Robustness is strongly model‑dependent and cannot be predicted from clean‑data performance alone; injection‑based contamination causes the greatest degradation, especially for local‑density and distance‑based detectors, while feature‑noise injection has a comparatively limited effect. PCA, SVM, HBOS, and IForest remain relatively stable, and the tuned neural detectors show intermediate robustness.

Why it matters

Practitioners deploying offline anomaly detectors in industrial control systems should consider training‑data integrity, and researchers should explore defenses against contamination‑style attacks.

Method details
  • SWaT pointwise pool after preprocessing: 449,919 samples, 44 features, 12.14% attack ratio.
  • Eleven detectors evaluated: PCA, SVM, HBOS, IForest, LOF, KNN, CBLOF, ABOD, MCD, AE, LSTM‑AE.
  • Contamination budgets examined: 1% to 10% of the clean training‑pool size.
  • Classical detector settings: IForest 100 trees, LOF 20 neighbors, KNN distance to 5th neighbor, CBLOF 8 clusters, ABOD fast 10‑neighbor, PCA 90% variance, SVM RBF kernel with gamma=scale.
  • Neural model configurations: AE uses LeakyReLU, dropout 0.1, Adam optimizer, batch size 2048; LSTM‑AE uses hidden size 256, AdamW optimizer, batch size 256.
Numbers
  • Processed samples, 449,919, total SWaT benchmark pool
  • Attack ratio, 12.14%, proportion of attack observations in the pool
  • Contamination budgets, 1% to 10%, range of training‑time manipulation evaluated
  • IForest trees, 100, number of trees used in the isolation forest
  • LOF neighbors, 20, number of neighbors for local outlier factor
  • PCA variance retained, 90%, amount of variance kept in the PCA detector
Limitations

The study does not establish robustness for online‑trained detectors, other datasets, or gradient‑driven poisoning attacks.

Injection-based contamination causes the greatest degradation, particularly for local-density and distance-based detectorsFound in the source text, word for word.

Picked because: Empirically studies the robustness of offline anomaly‑detection pipelines for industrial control systems under training‑time data contamination, offering practical guidance and released evaluation code for DevOps security.

Paper 4 of 5

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

Hanling Tian, Gengyu Zhang, Zeyang Sha and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Existing prompt injection attacks (DPI, BadChain, GCG) fail on memory‑augmented LLM agents because retrieved memories appear in long, variable‑depth prompts, diluting semantic cues and breaking positional regularities, so the injected command does not reliably steer generation.

Approach

InjecMEM crafts a two‑part injection: a high‑recall anchor that ensures the poisoned page is retrieved, and an adversarial command optimized to survive uncertain fused contexts. The command is learned via gradient‑based coordinate search over synthetic prompt templates and insertion positions, averaging across many surrogates. The method is retriever‑agnostic and can be jointly optimized across multiple backbone models to improve transfer. Multi‑GCG, the concrete instantiation, optimizes a single short command that remains effective despite variable placement and memory drift.

Result

In the primary MemoryOS setting Multi‑GCG achieves 76.6% ASR‑c and 35.6% joint ASR while baseline attacks all report 0%; retrieval success is 46.5% on MemoryOS and 37.2% on MemGPT, showing the attack persists under memory drift and does not affect unrelated topics.

Why it matters

Developers of LLM agents with persistent memory and security researchers should care because the work shows that even a single injected interaction can reliably hijack future generations, exposing a new attack surface.

Method details
  • Evaluated primarily on MemoryOS with Qwen2.5-7B-Instruct as backbone
  • Also tested on MemGPT, Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3 and other Qwen2.5 variants
  • Synthetic dialogues span 19 domains, 944 conversations, 3096 pages; plus real‑user WildChat data
  • Baselines: Direct Prompt Injection (DPI), BadChain, GCG, and an on‑topic paragraph anchor
  • Metrics: Retrieval Success Rate (RSR), Attack Success Rate conditional on retrieval (ASR‑c), joint ASR (ASR‑j)
  • Ablations include transfer within Qwen2.5 family and Family‑Joint Multi‑GCG optimization
Numbers
  • ASR‑c 76.6% vs 0% for DPI, BadChain, GCG
  • ASR‑j 35.6% vs 0% for DPI, BadChain, GCG
  • RSR 46.5% on MemoryOS
  • RSR 37.2% on MemGPT
  • MemGPT ASR‑c 48.6% and ASR‑j 18.1%
  • Qwen2.5‑7B‑Inst ASR‑c 78.4% (7B only) vs 0% for 3B‑Inst
Limitations

The paper does not establish effectiveness for systems that aggressively rewrite interactions before storage and only considers a single‑shot injection scenario.

Multi‑GCG remains effective by optimizing a single command over a distribution of surrogates with varying lengths and insertion positions, making it robust to retrieval‑induced variability.Found in the source text, word for word.

Picked because: Presents InjecMEM, a memory‑injection attack on LLM agent memory stores, exposing a concrete vulnerability and providing attack code useful for verification and hardening of deployed agents.

Paper 5 of 5

RAD: Rule-Augmented Relational Anomaly Detection

Noah Dahle, Anne Tumlin, Ngoc Tran and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Existing methods flatten multiple tables into a single feature matrix, which obscures entity identity, schema structure, and multi‑hop dependencies, limiting detection of anomalies that depend on relational context; the obvious fix of applying flat‑tabular detectors fails because it cannot capture relational or symbolic behavioral evidence.

Approach

RAD first mines candidate symbolic rules from random‑forest paths on flattened summaries of each target entity, refines them into compact predicates, and injects the resulting rule features into node attributes of a heterogeneous relational graph. The graph is processed by a relation‑aware GraphSAGE encoder to produce target‑node embeddings. Training optimizes attribute reconstruction loss, an optional edge‑reconstruction BCE loss, and a pairwise ranking loss that pushes labeled anomalies above normal nodes. Hyperparameters for the reconstruction and ranking weights are tuned on validation AUPRC. The final anomaly score combines reconstruction error (and optionally edge‑reconstruction score) to rank targets.

Result

Across the LANL cybersecurity event detection, Amazon review churn, and H&M purchase churn tasks, RAD achieves the best average rank on both AUROC and AUPRC, outperforming flattened tabular detectors and relational baselines under natural class imbalance.

Why it matters

Practitioners needing anomaly detection in multi‑table relational databases should consider RAD because it leverages both relational graph structure and symbolic rule evidence to improve detection of context‑dependent anomalies.

Method details
  • Two‑layer relation‑aware GraphSAGE encoder with mean aggregation, hidden dimension 128
  • Random‑forest rule mining uses forest size 1200 for LANL and 200 for Amazon/H&M, depth/leaf settings as in Table 2
  • Training objective combines MSE attribute reconstruction, optional BCE edge reconstruction, and pairwise ranking loss; weights selected by Optuna on validation AUPRC
  • Benchmarked on three relational datasets: LANL (7 tables, 1.65B records, 1.64B targets), Amazon (3 tables, 15.0M records, 1.85M targets), H&M (3 tables, 16.7M records, 1.37M targets)
  • Baselines include flattened tabular detectors, relational baselines (DOMINANT, SL‑GAD, Relational Deep Learning, RelBench) and a supervised MLP on flattened features
  • Ablations evaluate edge‑reconstruction (BCE vs no‑BCE), rule injection (with vs without rules), and ranking supervision (with vs without)
Numbers
  • AUROC, best average rank, compared against flattened tabular detectors and relational baselines
  • AUPRC, best average rank, compared against flattened tabular detectors and relational baselines
  • Anomaly rate, 1.5%, Amazon and H&M
Limitations

RAD’s rule mining relies on an auxiliary flattened view and requires labeled data for the random‑forest stage, so it does not extend directly to unlabeled settings and uses sampled subgraphs for LANL evaluation.

RAD improves anomaly ranking over flattened tabular detectors and relational baselines under natural class imbalanceFound in the source text, word for word.

Picked because: Proposes RAD, a rule‑augmented relational anomaly detection method that preserves database schema structure, with released implementation useful for platform engineers monitoring relational services.