arXiv digest

Monday

August 24, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Natural-Language Workflows Are Not Software Yet: Artifact-Driven Compilation for Reliable Agent Execution

Xiangzhe Xu, Hanxi Guo, Guangyu Shen and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Natural-language workflows leave data dependencies implicit and agents often fail to follow long or branching instructions, and simply letting agents infer dependencies or using unconstrained LLM translation does not reliably fix these issues.

Approach

Artic compiles a natural-language workflow into an artifact-driven workflow where each step explicitly declares read/write artifacts, constraints gate artifacts, and control transfers are explicit. The compiler drafts a candidate workflow using an LLM, then a constraint checker runs lightweight program analysis to flag suboptimal regions. A faithfulness-validation component performs static, inductive, and dry-run checks on the candidate. If validation fails, diagnostics are fed back to the constrained‑optimization stage; otherwise the workflow is accepted. The resulting workflow is executed by a deterministic orchestrator that invokes subagents and manages artifact versions.

Result

Artic improves the task resolve rate by 28 percentage points over the original text workflow and yields workflows that are 32 and 56 percentage points more consistent in cross‑model and repeated‑execution setups respectively. In the average across domains, Artic achieves an 85% resolve rate versus 62% for plain text workflows.

Why it matters

Researchers and practitioners building domain‑specific agentic procedures should care because Artic provides a systematic way to make natural‑language workflows more enforceable and reliable.

Method details
  • Compiler models evaluated: GPT-5.4, Sonnet-4.6, and GLM-5.
  • Executor models span from 3B active parameters to 700B+ total parameters.
  • Benchmarks: 488 problem instances from 11 real-world domain workflows drawn from SOP-Bench and -Bench.
  • Baselines compared: textual execution, SkillCreator, direct code generation, and existing natural-language-to-workflow approaches.
  • Implementation size: 5.9K lines of Python code.
Numbers
  • task resolve rate improvement, +28 percentage points, over original text workflow
  • cross‑model consistency improvement, +32 percentage points, compared to non‑artifact workflows
  • repeated‑execution consistency improvement, +56 percentage points, compared to non‑artifact workflows
  • average resolve rate, 85%, Artic vs 62% Text
Limitations

The paper does not state any limitations.

it improves task resolve rate by 28 percentage points over the original text workflow.Found in the source text, word for word.

Picked because: Introduces an artifact-driven compilation system that turns natural-language workflows into reliable executable agents, providing concrete tooling for LLM‑based automation.

Paper 2 of 5

AI-to-AI Code Reviews of GitHub Pull Requests

Niruthiha Selvanayagam, Taher A. Ghaleb · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Prior empirical studies assumed pull requests were human-authored and reviewed, ignoring the growing presence of AI agents; simply filtering out AI events does not capture the closed-loop AI-to-AI interactions that now exist.

Approach

The authors link AI-attributed pull requests with AI-attributed review events from the CodAGE dataset, applying a two‑pass signature framework to reliably identify author and reviewer agents. They aggregate events per (PR, reviewer) pair, separating same‑product from cross‑product configurations. Using the CodeRabbit comment‑category classifier they quantify comment types, volume, and latency. Analyses are performed on three derived datasets (cross‑product, same‑product, and agent‑authored without AI review) to characterize prevalence and reviewer behavior.

Result

Cross‑product AI‑to‑AI review occurs in about 1.6% of identified agent‑authored PRs, amounting to 45k PRs, and grew by more than two orders of magnitude from 2025‑Q1 to 2025‑Q3. CodeRabbit labeled 35.0% of its comments on Claude‑Code PRs as refactor comments versus 10.5% on Copilot PRs. For three of four dual‑role reviewers, mean comments per PR were 58 to 65% higher in same‑product groups, and median latency was 1.2 minutes for cross‑product pairs versus 4.7 minutes for same‑product pairs.

Why it matters

Researchers and tool builders must account for AI‑to‑AI code review loops when sampling GitHub data and designing automated review systems, as they represent a growing but distinct portion of development activity.

Method details
  • Data source: CodAGE snapshot covering 2024-01-01 to 2026-04-15.
  • Attribution: strict signature framework requiring body or vendor‑login evidence.
  • Dataset size: 248,641 unique AI‑attributed PRs with at least one AI review.
  • Cross‑product PRs: 45,269; same‑product PRs: 208,145; both: 4,773.
  • Reviewer analysis: CodeRabbit comment‑category classifier applied to review comments.
Numbers
  • cross‑product proportion, 1.6%, of identified agent‑authored PRs
  • cross‑product PRs, 45,269, compared with same‑product PRs 208,145
  • growth, > two orders of magnitude, from 2025‑Q1 to 2025‑Q3
  • refactor comment rate, 35.0%, on Claude‑Code PRs vs 10.5% on Copilot PRs
  • median latency, 1.2 minutes, for cross‑product pairs vs 4.7 minutes for same‑product pairs
Limitations

The study does not establish review correctness or software quality and cannot separate reviewer behavior from PR characteristics due to limited timestamp availability.

Claude-Code PRs receive more refactor comments from CodeRabbit than Copilot PRs (35.0% vs. 10.5%)Found in the source text, word for word.

Picked because: Releases a large‑scale AI‑to‑AI code‑review dataset and analysis pipeline, enabling engineers to build and evaluate self‑hosted AI code‑review bots.

Paper 3 of 5

Beyond Fault Localization: A Trajectory-Level Study of LLM Agents for Microservice Root Cause Analysis

Qisheng Lu, Aoyang Fang, Junjielong Xu and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Prior automated RCA methods only measured endpoint correctness, ignoring the evidentiary basis and fault‑propagation route, so agents could localize the faulty service yet fail to reconstruct its impact path, and simply improving endpoint accuracy does not fix this evidence‑handling gap.

Approach

The paper introduces DiagGuard, a two‑stage defense‑in‑depth architecture that wraps a reasoning core (the ThinkDepth.ai Diagnostician) with a Grounder that surveys all available telemetry before localization and a Verifier that audits the diagnosis against the gathered evidence before committing. The Grounder ensures the agent is under‑grounded, while the Verifier provides inference‑time verification to catch misinterpretations or unsupported inferences. Together they address the three failure families identified in the trajectory analysis. The design is instantiated on the ThinkDepth.ai framework and evaluated on independent benchmarks.

Result

DiagGuard improves the top‑1 accuracy of root‑cause localization from 43.5% to 52.5% in an independent validation setting, demonstrating that trajectory‑level evaluation can expose hidden limitations of agents that achieve correct endpoint answers but poor diagnostic quality.

Why it matters

SREs and AIOps researchers should care because the approach reveals hidden diagnostic failures and provides a concrete defense that measurably improves automated root‑cause analysis for microservices.

Method details
  • Evaluated six agent frameworks: ThinkDepth.ai, AIQ, TaskWeaver, ClaudeCode, OpenRCA, mABC
  • Backbone models used include qwen3.5-plus-2026-02-15, claude-sonnet-4.6, doubao-seed-2.0-pro-2026-02-15
  • Benchmarks: RCABench (500‑case sample from 1,430 cases) and AIOps 2025 (400 incidents over ten services)
  • Baseline core is the ThinkDepth.ai Diagnostician; DiagGuard adds Grounder and Verifier defenses
  • DiagGuard was validated with a different model (Seed 2.0 Pro) and benchmark (AIOps 2025)
Numbers
  • Acc@1 43.5% baseline
  • Acc@1 52.5% with DiagGuard
  • 3,500 diagnostic trajectories analyzed
  • 500‑case sample from RCABench
  • 400 incidents in AIOps 2025 benchmark
  • six agent frameworks evaluated
Limitations

The paper does not state any limitations.

In an independent setting with a different model, benchmark, and service topology, DiagGuard raises Acc@1 from 43.5% to 52.5%.Found in the source text, word for word.

Picked because: Presents a trajectory‑level evaluation of LLM agents for microservice root‑cause analysis, with actionable insights and code for integrating such agents into production observability stacks.

Paper 4 of 5

Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking

Arulnidhi Karunanidhi · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Content‑only screening cannot reliably distinguish false assertions from true ones, and the obvious fix of adding a provenance weight fails because any weight strong enough to block a shaped attack also discards legitimate untrusted evidence.

Approach

The paper builds a four‑stage write‑time screening pipeline that flags content if any stage flags, and evaluates provenance‑weighted retrieval where a scalar weight modifies the similarity score. The pipeline combines deterministic regex, model‑based detectors (ProtectAI DeBERTa v2, Meta Llama Prompt Guard 2, LLM Guard), LLM‑as‑judge judges (GPT‑4o‑mini, Claude Haiku 4.5), and Aegis stages. Retrieval is plain nearest‑neighbor without reranking. The authors then argue for a bounded occupancy constraint on provenance rather than an additive weight.

Result

Poisoning 1.2% of the LongMemEval corpus drops accuracy from 0.850 to 0.300, a two‑thirds loss. The four‑stage screening pipeline attains 0.832 recall on indirect injection, flags 1.5% of benign trigger‑word text, and rejects none of 360 poisoned memories. The shipped provenance weight is statistically indistinguishable from no defense (p=0.80), while a stronger weight only improves accuracy to 0.7000 in a mixed‑provenance corpus but collapses it to 0.0417 when evidence is untrusted.

Why it matters

Developers of agent memory systems and security researchers should note that simple content screening and additive provenance weighting are insufficient defenses against memory poisoning.

Method details
  • Write‑path screening evaluates ten configurations including a naive regex baseline and three model‑based detectors.
  • Benchmarks use five corpora: direct injection (deepset/prompt‑injections), indirect injection (InjecAgent), Dolly‑15k, templated memory‑like entries, and NotInject.
  • Memory‑quality benchmark uses LongMemEval_S with 500 questions and ~115K tokens of chat history.
  • Ablation rescoring is done cumulatively per stage to isolate contribution of injection detection versus secrets detection.
  • Provenance weight is compared to a no‑defense baseline; statistical indistinguishability is reported with p=0.80.
Numbers
  • accuracy 0.850 → 0.300 after 1.2% poisoning
  • recall 0.832 on indirect injection
  • benign flag rate 1.5%
  • 0 of 360 poisoned memories rejected
  • p=0.80 for shipped provenance weight vs no defense
  • accuracy 0.3167 → 0.7000 in mixed‑provenance corpus
Limitations

The study only evaluates a non‑adaptive single‑pass attack on one memory implementation and does not implement or test the proposed bounded‑occupancy provenance constraint.

At 1.2% of the corpus, that attack removes two‑thirds of the memory’s value.Found in the source text, word for word.

Picked because: Demonstrates concrete memory‑poisoning attacks on LLMs and proposes mitigation stages, offering practical verification and security measures for self‑hosted models.

Paper 5 of 5

Ontology-supported AI Model and Dataset Management

Jan Novacek, Ali Ahari, Tobias Müller and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing AI model repositories like Hugging Face provide only semi‑structured natural language descriptions and current ML platforms (MLflow, H2O, Ray) add metadata that lacks interoperability, so assets cannot be reliably exchanged across companies.

Approach

The authors built the AIMDEP platform that registers AI assets and attaches metadata expressed in the AIMDEO ontology. Assets (datasets, models) are uploaded, automatically feature‑extracted, and annotated via the ontology. The ontology metadata is stored as OWL files and can be exported or accessed through a REST API. Search functions locate assets using ontology terms, and the platform can execute models directly for supported frameworks. The overall workflow links domain experts, AI experts, and end users through shared, semantically rich asset descriptions.

Result

The registered AI model’s metadata reports an average precision of 0.9794, demonstrating that the platform can capture and expose quantitative quality metrics for exchanged models.

Why it matters

Industrial AI developers and tool vendors should care because the platform enables semantically rich, interoperable exchange of models and datasets, reducing integration effort.

Method details
  • Model framework: scikit‑learn
  • Model name: RF Instruction Cache‑Line Access Classifier
  • Metric recorded: average precision 0.9794
  • Dataset includes cache replacement strategy (LRU) and cache size (2048 Byte)
  • Metadata exported as OWL via the AIMDEO ontology
Numbers
  • average precision, 0.9794, compared against none
  • Cache‑Size, 2048, compared against none
  • Replacement‑Strategy, LRU, compared against none
Limitations

The paper does not evaluate interoperability with other ontologies or provide a quantitative benchmark against existing model‑exchange platforms.

The platform incorporates an ontology that can foster a more profound common understanding of what is required in these tasksFound in the source text, word for word.

Picked because: Proposes and releases an ontology‑based system for managing AI models and datasets, supporting platform engineers in tracking and governing AI assets.