Deyao Hong, Yizhe Chi, Wenyi Li and 7 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Existing benchmarks only check behavioural correctness, allowing agents to cheat by copying the original implementation (Blindness) without actually performing the migration.
Approach
The paper introduces SWE Refactor Bench, a set of 20 whole‑repository migration tasks evaluated with a three‑stage protocol: Migration Audit verifies that the migration truly occurred, Behavioural Tests run a fixed test suite to ensure functional parity, and Agentic Verification employs six independent coding agents to generate targeted tests for hidden behavioural differences. The Migration Audit uses a judge model (gpt‑5.6‑sol) with majority voting over three samples. The fixed test suite records checks from the original repository. The six verification agents each run for one hour to find counter‑examples. This pipeline separates migration completeness from behavioural correctness.
Result
Only 28 of 520 runs (5.4%) pass all three stages, and 13 of the 20 tasks receive no accepted solution; the top model, claude‑opus‑5, attains a composite score of 47.0/100. Among the 340 runs that pass Migration Audit, 58% reach 99% of the fixed checks while only 26% reach 100%.
Why it matters
Researchers and engineers building coding agents should use SWE Refactor Bench to gauge and improve migration capability beyond mere behavioural correctness.
Best composite score achieved by claude‑opus‑5: 47.0/100
Category scores: build toolchain rewrites 31.4, language rewrites 5.6
Numbers
Accepted runs, 28, 5.4% of 520 runs
Best model score, 47.0/100, claude‑opus‑5
Unsolved tasks, 13, out of 20 total tasks
Migration completeness 99% reach, 58%, of runs that passed Migration Audit
Full correctness 100% reach, 26%, of runs that passed Migration Audit
Category score build toolchain rewrites, 31.4, compared to language rewrites 5.6
Limitations
The paper does not establish that any coding agent can reliably achieve a perfect whole‑repository migration; it only measures current frontier models.
only 28 of 520 runs ($5.4\%$) pass all three stagesFound in the source text, word for word.
Picked because: Introduces the SWE Refactor Bench, a concrete dataset and evaluation suite for coding agents to perform whole‑repository stack migrations, directly applicable to large‑scale software engineering automation.
Seth Karten, Alex L. Zhang, Kevin Thomas and 8 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Prior approaches relied on the LLM’s weights and active token context, which cannot hold the external information and compute needed for long‑horizon agency; simply increasing context size or fine‑tuning does not provide persistent state or test‑time compute.
Approach
Prime Agent introduces a persistent IPython REPL (L2) and a Continual Harness (L3) that store histories, memories, skills, and subagent specifications across trajectories. Recursive subagents communicate directly with each other and with humans via an Agents View. Information management decides what state enters a model invocation, while computation management maps model‑selected actions to code, tools, and subagent sessions. The system separates information and computation management, uses hierarchical levels L0‑L3, and records all calls for accounting and recovery.
Result
Prime Agent lifts ARC‑AGI‑3 RHAE Best@1 from 30% to 95.5% and matches or exceeds native and popular harnesses on long‑context coding, GPU‑kernel generation, emulator construction, and autonomous nanoGPT speedruns. It also sustains an 85.5‑hour nanoGPT run with 19 validated records and supports continuous technology progression in Factorio and long‑horizon exploration in MazeBench.
Why it matters
Researchers building long‑horizon autonomous agents and evaluation harnesses should care because Prime Agent demonstrates how a standardized, expressive substrate can dramatically improve performance and enable persistent, recursive computation.
Method details
Evaluated on ARC‑AGI‑3 benchmark
Compared against Pi, Claude Code, Codex, Hermes Agent, OpenCode, and Kimi‑Code
Performed a 85.5‑hour autonomous nanoGPT speedrun with 19 validated records
Tested on Factorio and MazeBench environments
Uses a persistent IPython REPL and recursive subagents as core components
Numbers
ARC‑AGI‑3 RHAE Best@1 95.5% (up from 30%)
85.5‑hour nanoGPT run
19 validated records
Limitations
The paper notes that models still experience friction when allocating subagents, managing retained information, and refining reusable state, and many harness capabilities remain underused because current models were not trained to operate them.
Prime Agent raises ARC-AGI-3 RHAE Best@1 from 30% to 95.5% and matches or exceeds native and popular harnesses across long-context coding, GPU-kernel generation, emulator construction, and autonomous nanoGPT speedruns.Found in the source text, word for word.
Picked because: Provides Prime Agent, an open‑source harness with a persistent REPL and continual compute extensions for building, testing, and verifying long‑horizon LLM agents in self‑hosted environments.
Mustafa Umut Ozbek, Taiwo Ojo, Pooria Madani and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CR
Problem
Prior work on offline industrial control system anomaly detection assumes the training data are trustworthy, so it does not address the degradation caused by training-time data contamination; simply using clean validation and test sets does not prevent the models from being poisoned during training.
Approach
The paper evaluates the robustness of eleven heterogeneous anomaly detectors on the Secure Water Treatment (SWaT) benchmark under three contamination strategies (random injection, similarity‑targeted injection, and feature‑noise injection) with budgets from 1% to 10%. Detectors are trained on a normal‑only pool, calibrated on a clean validation split, and tested on an untouched test split. Contamination either inserts attack samples into the nominal training pool or adds bounded Gaussian noise to selected normal samples. The study uses a unified offline protocol and reports performance degradation across models. Results are compared against clean‑data baselines to assess model‑dependent robustness.
Result
Robustness is strongly model‑dependent and cannot be predicted from clean‑data performance alone; injection‑based contamination causes the greatest degradation, especially for local‑density and distance‑based detectors, while feature‑noise injection has a comparatively limited effect. PCA, SVM, HBOS, and IForest remain relatively stable, and the tuned neural detectors show intermediate robustness.
Why it matters
Practitioners deploying offline anomaly detectors in industrial control systems should consider training‑data integrity, and researchers should explore defenses against contamination‑style attacks.
Method details
SWaT pointwise pool after preprocessing: 449,919 samples, 44 features, 12.14% attack ratio.
Processed samples, 449,919, total SWaT benchmark pool
Attack ratio, 12.14%, proportion of attack observations in the pool
Contamination budgets, 1% to 10%, range of training‑time manipulation evaluated
IForest trees, 100, number of trees used in the isolation forest
LOF neighbors, 20, number of neighbors for local outlier factor
PCA variance retained, 90%, amount of variance kept in the PCA detector
Limitations
The study does not establish robustness for online‑trained detectors, other datasets, or gradient‑driven poisoning attacks.
Injection-based contamination causes the greatest degradation, particularly for local-density and distance-based detectorsFound in the source text, word for word.
Picked because: Empirically studies the robustness of offline anomaly‑detection pipelines for industrial control systems under training‑time data contamination, offering practical guidance and released evaluation code for DevOps security.
Hanling Tian, Gengyu Zhang, Zeyang Sha and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CR
Problem
Existing prompt injection attacks (DPI, BadChain, GCG) fail on memory‑augmented LLM agents because retrieved memories appear in long, variable‑depth prompts, diluting semantic cues and breaking positional regularities, so the injected command does not reliably steer generation.
Approach
InjecMEM crafts a two‑part injection: a high‑recall anchor that ensures the poisoned page is retrieved, and an adversarial command optimized to survive uncertain fused contexts. The command is learned via gradient‑based coordinate search over synthetic prompt templates and insertion positions, averaging across many surrogates. The method is retriever‑agnostic and can be jointly optimized across multiple backbone models to improve transfer. Multi‑GCG, the concrete instantiation, optimizes a single short command that remains effective despite variable placement and memory drift.
Result
In the primary MemoryOS setting Multi‑GCG achieves 76.6% ASR‑c and 35.6% joint ASR while baseline attacks all report 0%; retrieval success is 46.5% on MemoryOS and 37.2% on MemGPT, showing the attack persists under memory drift and does not affect unrelated topics.
Why it matters
Developers of LLM agents with persistent memory and security researchers should care because the work shows that even a single injected interaction can reliably hijack future generations, exposing a new attack surface.
Method details
Evaluated primarily on MemoryOS with Qwen2.5-7B-Instruct as backbone
Also tested on MemGPT, Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3 and other Qwen2.5 variants
Synthetic dialogues span 19 domains, 944 conversations, 3096 pages; plus real‑user WildChat data
Baselines: Direct Prompt Injection (DPI), BadChain, GCG, and an on‑topic paragraph anchor
Metrics: Retrieval Success Rate (RSR), Attack Success Rate conditional on retrieval (ASR‑c), joint ASR (ASR‑j)
Ablations include transfer within Qwen2.5 family and Family‑Joint Multi‑GCG optimization
Numbers
ASR‑c 76.6% vs 0% for DPI, BadChain, GCG
ASR‑j 35.6% vs 0% for DPI, BadChain, GCG
RSR 46.5% on MemoryOS
RSR 37.2% on MemGPT
MemGPT ASR‑c 48.6% and ASR‑j 18.1%
Qwen2.5‑7B‑Inst ASR‑c 78.4% (7B only) vs 0% for 3B‑Inst
Limitations
The paper does not establish effectiveness for systems that aggressively rewrite interactions before storage and only considers a single‑shot injection scenario.
Multi‑GCG remains effective by optimizing a single command over a distribution of surrogates with varying lengths and insertion positions, making it robust to retrieval‑induced variability.Found in the source text, word for word.
Picked because: Presents InjecMEM, a memory‑injection attack on LLM agent memory stores, exposing a concrete vulnerability and providing attack code useful for verification and hardening of deployed agents.
Noah Dahle, Anne Tumlin, Ngoc Tran and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
Existing methods flatten multiple tables into a single feature matrix, which obscures entity identity, schema structure, and multi‑hop dependencies, limiting detection of anomalies that depend on relational context; the obvious fix of applying flat‑tabular detectors fails because it cannot capture relational or symbolic behavioral evidence.
Approach
RAD first mines candidate symbolic rules from random‑forest paths on flattened summaries of each target entity, refines them into compact predicates, and injects the resulting rule features into node attributes of a heterogeneous relational graph. The graph is processed by a relation‑aware GraphSAGE encoder to produce target‑node embeddings. Training optimizes attribute reconstruction loss, an optional edge‑reconstruction BCE loss, and a pairwise ranking loss that pushes labeled anomalies above normal nodes. Hyperparameters for the reconstruction and ranking weights are tuned on validation AUPRC. The final anomaly score combines reconstruction error (and optionally edge‑reconstruction score) to rank targets.
Result
Across the LANL cybersecurity event detection, Amazon review churn, and H&M purchase churn tasks, RAD achieves the best average rank on both AUROC and AUPRC, outperforming flattened tabular detectors and relational baselines under natural class imbalance.
Why it matters
Practitioners needing anomaly detection in multi‑table relational databases should consider RAD because it leverages both relational graph structure and symbolic rule evidence to improve detection of context‑dependent anomalies.
Method details
Two‑layer relation‑aware GraphSAGE encoder with mean aggregation, hidden dimension 128
Random‑forest rule mining uses forest size 1200 for LANL and 200 for Amazon/H&M, depth/leaf settings as in Table 2
Training objective combines MSE attribute reconstruction, optional BCE edge reconstruction, and pairwise ranking loss; weights selected by Optuna on validation AUPRC
Baselines include flattened tabular detectors, relational baselines (DOMINANT, SL‑GAD, Relational Deep Learning, RelBench) and a supervised MLP on flattened features
Ablations evaluate edge‑reconstruction (BCE vs no‑BCE), rule injection (with vs without rules), and ranking supervision (with vs without)
Numbers
AUROC, best average rank, compared against flattened tabular detectors and relational baselines
AUPRC, best average rank, compared against flattened tabular detectors and relational baselines
Anomaly rate, 1.5%, Amazon and H&M
Limitations
RAD’s rule mining relies on an auxiliary flattened view and requires labeled data for the random‑forest stage, so it does not extend directly to unlabeled settings and uses sampled subgraphs for LANL evaluation.
RAD improves anomaly ranking over flattened tabular detectors and relational baselines under natural class imbalanceFound in the source text, word for word.
Picked because: Proposes RAD, a rule‑augmented relational anomaly detection method that preserves database schema structure, with released implementation useful for platform engineers monitoring relational services.