arXiv digest

Saturday

August 29, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench

Dewu Zheng, Yanlin Wang, Xiwen Wang and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Prior work reduces code review to a single-round static decision task, which fails to capture the iterative, multi-round interactions of real-world reviews; simply applying static models does not work because performance drops sharply as the number of interaction rounds grows.

Approach

The paper introduces MCR-Bench, a defect state‑aware benchmark that simulates realistic multi‑round code review. It constructs each task from PR histories, combining static PR metadata with dynamic review information (code diffs and discussion timeline). Ground truth is provided as Defect Cards containing a defect description, precise location, lifecycle state, and supplementary metadata. The benchmark pipeline consists of language and repository selection, PR data collection, LLM‑based state‑aware defect annotation, and manual cross‑validation. Evaluation uses both mainstream LLMs and adapted ACR baselines to assess defect detection and state tracking across rounds.

Result

Extensive experiments show that mainstream LLMs have limited overall capability in defect detection and defect lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increases; performance also varies widely across defect types and severity levels.

Why it matters

Researchers and tool developers focusing on automated code review should care because the benchmark reveals current LLM weaknesses in multi‑round, state‑aware review scenarios.

Method details
  • 2,269 multi‑round code review tasks are included.
  • Five programming languages are covered: Python, Java, JavaScript, TypeScript, C#.
  • Seven LLMs are evaluated: three commercial (GPT‑5.2, Claude‑Haiku‑4.5, Gemini‑3‑Flash) and four open‑source (DeepSeek‑V3.2, Qwen3‑Max, GLM‑4.7, Kimi‑k2).
  • Two ACR baselines are adapted: PR‑Agent and Hybrid‑Review.
  • Repositories must have >100 stars, >10 contributors, and >1,500 PRs, and use permissive licenses.
Numbers
  • tasks, 2,269, MCR‑Bench
  • languages, 5, most active on GitHub
  • models, 7, evaluated LLMs
  • baselines, 2, adapted ACR methods
  • repository stars, >100, required filter
  • required PR count, >1,500, selection criterion
Limitations

The paper does not state any explicit limitations.

experiments reveal that mainstream LLMs exhibit limited overall performance in defect detection and defect lifecycle state tracking, with performance degrading significantly as the number of interaction rounds increasesFound in the source text, word for word.

Picked because: Provides MCR-Bench, a released benchmark and tooling for realistic multi‑round code‑review automation, directly useful for building full‑stack code‑review assistants.

Paper 2 of 5

RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution

Junjie Zhang, Hui Liu, Kecheng Chen and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Existing automatic red‑teaming methods rely on fixed attacks or trajectory‑based retrieval that suffers from retrieval bias, unclear tool credit, and context overhead, limiting interpretability and effectiveness.

Approach

RedEvoAgent distills cross‑case attack trajectories into a concise, human‑readable attack skill document. The skill evolves via tool‑effectiveness profiling and Deciding‑Tool Attribution, guided by a validation ratchet that keeps only updates improving validation performance. The attacker agent architecture includes a toolbox of seven jailbreak tools, a skill document inserted into the system prompt, and an attack workflow that iteratively selects tool‑call or query actions. Experience is collected from past rollouts to inform skill updates. Validation‑guided skill evolution iterates over ratchet rounds to refine the skill.

Result

RedEvoAgent consistently outperforms fixed and agentic baselines, achieving the highest attack success rates and HarmScore while using fewer tool calls per case, demonstrating both higher effectiveness and efficiency.

Why it matters

Security evaluators of product‑level LLM agents should adopt RedEvoAgent for more effective and efficient automated red‑teaming of black‑box agents.

Method details
  • Attacker model is GPT‑4o mini (OpenAI, 2024).
  • Attack toolbox integrates seven tools: GCG, AutoDAN, AmpleGCG, Template, FlipAttack, RolePlay, Prompt Substitution.
  • Evaluated on Agent Security Bench (400 cases, train/val/test 80/40/280) and AgentHarm (208 cases, split 40/20/148).
  • Target models: MiniMax‑M2.5, DeepSeek‑V4‑Flash, Qwen3.5‑35B paired with Claude Code or Codex execution harnesses.
  • Baselines include No Jailbreak, six isolated tools (GCG, AmpleGCG, AutoDAN, Template, FlipAttack, RolePlay), RedCodeAgent, and MAJIC.
  • Ablations remove Tool‑Effectiveness Profile, trajectory collection, or Deciding‑Tool Attribution, showing ASR drops to 76.9%, 85.0%, and 87.5% respectively.
Numbers
  • ASR 93.2% on ASB MiniMax‑M2.5 with Claude Code (vs 77.5% No Skill).
  • ASR 99.3% on ASB DeepSeek‑V4‑Flash with Claude Code (vs 96.4% No Skill).
  • HarmScore 20.8% on AgentHarm MiniMax‑M2.5 with Claude Code (vs 15.4% No Skill).
  • Average tool calls 1.8 per case on ASB MiniMax‑M2.5 (vs 3.0 No Skill).
  • Zero‑shot transfer ASR 95.3% (+5.6) on Qwen3‑8B (vs 89.7% No Skill).
  • Zero‑shot transfer ASR 90.5% (+9.7) on Codex (vs 80.8% No Skill).
Limitations

The paper does not establish performance when adding completely new jailbreak tools beyond those evaluated.

RedEvoAgent achieves the highest attack performance in every evaluated setting.Found in the source text, word for word.

Picked because: Introduces RedEvoAgent, an open‑source red‑teaming LLM agent that evolves skills automatically, offering concrete artifacts for securing LLM‑driven services.

Paper 3 of 5

Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling

Maksim Utushkin, Andrei Ovsiannikov, Alexander D'yakonov · abstract · pdf

quote verifiedfigures checkedread: full textcs.IR

Problem

Previous industrial GNN systems either ignored trainable user IDs because a full embedding table would exceed hundreds of gigabytes, or used naive temporal neighbor sampling that scans entire adjacency lists, which is infeasible for users with tens of thousands of friends.

Approach

The method replaces the full per‑user ID table with a multi‑hash embedding layer that maps each user into a shared hash table, cutting memory dramatically. It stores edges in a timestamp‑sorted CSR format and selects temporal neighbors via a binary‑search cutoff, turning sampling from O(deg(v)+k) to O(log deg(v)+k). A GATv2 encoder with two role‑specific heads processes the multi‑hash IDs (and optional tabular features). Training runs on 8 A100 GPUs with DDP, while CPU workers perform the temporal sampling. The pipeline decouples sampling, training, and inference to fit on a single 8‑GPU host.

Result

Offline ranking quality reaches a per‑user ROC‑AUC of 0.6278, the highest among the evaluated models. In an online A/B test the system boosts friend additions from recommendations by 16 % and unique friend adders by 11.5 % over a strong production baseline.

Why it matters

Teams building large‑scale graph‑based recommender systems should consider multi‑hash ID embeddings and binary‑search temporal sampling to achieve production‑grade accuracy while staying within memory and latency budgets.

Method details
  • Graph size: 194M users and 28B edges
  • Training interactions: 1.4B train, 25M test
  • Baselines: Top‑pop, MF, WalkGNN
  • Ablations compare multi‑hash IDs only (ROC‑AUC 0.5997) vs full table (ROC‑AUC 0.6246) vs features+multi‑hash (ROC‑AUC 0.6278)
  • Temporal sampling cost reduced from full scan to binary‑search (per‑batch cost essentially unchanged)
  • Embedding table memory reduced by >98% compared to a full float32 table
Numbers
  • friend additions increase, 16%, over strong production baseline
  • unique friend adders increase, 11.5%, over strong production baseline
  • ROC‑AUC, 0.6278, Ours vs Top‑pop 0.5050
  • ROC‑AUC, 0.5997, multi‑hash IDs only vs full table 0.6246
  • embedding table size reduction, >98%, compared to full table
Limitations

The paper does not establish performance for cold‑start users, non‑real‑time serving, or applicability beyond the friendship graph.

our system increases friend additions from recommendations by 16 percent and unique friend adders by 11.5 percent over a strong production baselineFound in the source text, word for word.

Picked because: Describes a production‑grade system for scaling GNNs to hundreds of millions of nodes using multi‑hash embeddings and temporal neighbor sampling, with released code that can be adapted to large‑scale infrastructure.

Paper 4 of 5

Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification

Jinghan Xu, Yikai Zhang, Aili Chen and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing propose-and-verify methods score every candidate on a fixed task set, wasting rollouts on unrelated behaviors and allowing aggregate scores to hide specific regressions; simply increasing verification rollouts does not improve harness quality.

Approach

HarnessLens is a budget-aware framework that jointly explores task space and user-configurable components, derives candidate modifications from execution trajectories, and selectively verifies each candidate on behavior-relevant tasks using an attributable-evidence gate. It consists of three components: Context Exploration (characterizes tasks and editable components), Trajectory Diagnosis (processes trajectories from initial and verification rollouts), and Harness Evolution (selects proposals, constructs candidate harnesses, and verifies them before updating the current harness). The method uses behavior-aware batch selection to choose verification tasks and an evidence gate to accept only modifications with attributable improvement.

Result

HarnessLens achieves the best or tied-best pass@1 in eight of twelve harness‑benchmark pairs while using the smallest budget. For OpenCode it raises Retail from 75.00% to 85.00%, Banking from 20.90% to 25.37%, and BIRD Mini‑Dev from 37.50% to 45.83%. For Codex it improves BIRD Mini‑Dev from 37.50% to 47.22%. The average held‑out performance gain across benchmarks is reported as 7.6‑13.6%.

Why it matters

Researchers and engineers building LLM‑based agent harnesses should care because HarnessLens provides a sample‑efficient way to evolve harnesses under tight interaction budgets, delivering reliable performance gains.

Method details
  • Model: deepseek-v4-flash-preview (DeepSeek-AI) used for all LLM agent and evolution roles.
  • Harness versions: OpenCode v1.17.13, Codex CLI v0.144.4, Pi Coding Agent v0.80.10.
  • Benchmarks: Retail (Barres et al. 2025), Banking Knowledge (Shi et al. 2026), Terminal-Bench 2.0 (Merrill et al. 2026), BIRD Mini-Dev (Li et al. 2024).
  • Baselines: Self-Harness, Meta-Harness, HarnessFix.
  • Budget caps: HarnessLens 200 total units; Self-Harness 4,800 TRAIN rollouts; Meta-Harness 660 TRAIN rollouts; HarnessFix 300 TRAIN rollouts.
  • Ablations: Fixed Batch, Random Batch, RHO-based Batch, Metric-Only Gate.
Numbers
  • pass@1 85.00 vs 75.00 on OpenCode Retail
  • pass@1 25.37 vs 20.90 on OpenCode Banking
  • pass@1 45.83 vs 37.50 on OpenCode BIRD Mini‑Dev
  • pass@1 47.22 vs 37.50 on Codex BIRD Mini‑Dev
  • budget 200 total units for HarnessLens vs 4,800 rollouts for Self‑Harness
  • average improvement 7.6‑13.6% over baselines
Limitations

The evaluation is limited to three specific agent harnesses and four benchmarks, so generalization to other harnesses or domains is not demonstrated.

HarnessLens attains the best or tied-best pass rate in eight of the twelve harness-benchmark pairsFound in the source text, word for word.

Picked because: Presents HarnessLens, a budget‑aware verification framework that efficiently evaluates LLM agent harness behavior, delivering tooling engineers can integrate into CI pipelines.

Paper 5 of 5

When Context Gets Root: Privilege Escalation in LLM Harnesses

Xingbang He, Yuanwei Chen, Yi Qian and 6 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Instruction hierarchy was intended to prevent low‑privilege content from influencing model behavior, but agent harnesses can elevate low‑level content to higher instruction levels, breaking the protection; simply labeling messages does not stop this escalation.

Approach

The authors define instruction privilege escalation, where an attacker‑controlled artifact entered via a tool message is re‑presented by the harness at a higher instruction level (user or system‑effective) during context construction. They exploit multi‑agent delegation, custom subagent installation, persistent goals, and scheduled tasks to cause the model to follow malicious instructions that would otherwise be rejected. The attack is evaluated on six coding‑agent harnesses, each paired with a specific LLM, both with unrestricted execution and with automatic permission review enabled. Success is measured by achieving predefined security‑sensitive objectives.

Result

The attacks achieve all 13 objectives on all six harnesses when execution is unrestricted, and also achieve all 13 objectives on the three harnesses that provide automatic permission review.

Why it matters

Developers of LLM agent harnesses and security researchers should care because the findings show that current instruction‑hierarchy defenses can be bypassed across multiple models and harness designs.

Method details
  • Six coding‑agent harnesses: Claude Code, Codex, Gemini CLI, Qwen Code, Kimi, OpenCode.
  • Models used: Opus 4.8, GPT‑5.5, Gemini 3.1 Pro Preview, Qwen3.7 Max, Kimi 3, DeepSeek‑V4‑Pro‑0813.
  • 13 attack objectives spanning confidentiality, integrity, availability, and remote code execution.
  • Baseline comparison against existing prompt‑injection and role‑confusion attacks at the tool level.
  • Evaluation modes: full‑access execution and automatic permission review (Auto PR) on three harnesses.
Numbers
  • attack objectives, 13, achieved on all six harnesses
  • harnesses evaluated, 6, full‑access success
  • harnesses with Auto PR, 3, full success
  • attack objectives achieved, 13/13, all harnesses
Limitations

The paper does not state any explicit limitations.

Across six coding-agent harnesses and 13 attack objectives, our end-to-end attacks achieve every objective on every harness under full access.Found in the source text, word for word.

Picked because: Analyzes how LLM harness context construction can cause privilege escalation, exposing a concrete attack vector and mitigation strategies applicable to deployed agent platforms.