Xiangzhe Xu, Hanxi Guo, Guangyu Shen and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Natural-language workflows leave data dependencies implicit and agents often fail to follow long or branching instructions, and simply letting agents infer dependencies or using unconstrained LLM translation does not reliably fix these issues.
Approach
Artic compiles a natural-language workflow into an artifact-driven workflow where each step explicitly declares read/write artifacts, constraints gate artifacts, and control transfers are explicit. The compiler drafts a candidate workflow using an LLM, then a constraint checker runs lightweight program analysis to flag suboptimal regions. A faithfulness-validation component performs static, inductive, and dry-run checks on the candidate. If validation fails, diagnostics are fed back to the constrained‑optimization stage; otherwise the workflow is accepted. The resulting workflow is executed by a deterministic orchestrator that invokes subagents and manages artifact versions.
Result
Artic improves the task resolve rate by 28 percentage points over the original text workflow and yields workflows that are 32 and 56 percentage points more consistent in cross‑model and repeated‑execution setups respectively. In the average across domains, Artic achieves an 85% resolve rate versus 62% for plain text workflows.
Why it matters
Researchers and practitioners building domain‑specific agentic procedures should care because Artic provides a systematic way to make natural‑language workflows more enforceable and reliable.
Method details
Compiler models evaluated: GPT-5.4, Sonnet-4.6, and GLM-5.
Executor models span from 3B active parameters to 700B+ total parameters.
Benchmarks: 488 problem instances from 11 real-world domain workflows drawn from SOP-Bench and -Bench.
Baselines compared: textual execution, SkillCreator, direct code generation, and existing natural-language-to-workflow approaches.
Implementation size: 5.9K lines of Python code.
Numbers
task resolve rate improvement, +28 percentage points, over original text workflow
cross‑model consistency improvement, +32 percentage points, compared to non‑artifact workflows
repeated‑execution consistency improvement, +56 percentage points, compared to non‑artifact workflows
average resolve rate, 85%, Artic vs 62% Text
Limitations
The paper does not state any limitations.
it improves task resolve rate by 28 percentage points over the original text workflow.Found in the source text, word for word.
Picked because: Introduces an artifact-driven compilation system that turns natural-language workflows into reliable executable agents, providing concrete tooling for LLM‑based automation.
Niruthiha Selvanayagam, Taher A. Ghaleb · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Prior empirical studies assumed pull requests were human-authored and reviewed, ignoring the growing presence of AI agents; simply filtering out AI events does not capture the closed-loop AI-to-AI interactions that now exist.
Approach
The authors link AI-attributed pull requests with AI-attributed review events from the CodAGE dataset, applying a two‑pass signature framework to reliably identify author and reviewer agents. They aggregate events per (PR, reviewer) pair, separating same‑product from cross‑product configurations. Using the CodeRabbit comment‑category classifier they quantify comment types, volume, and latency. Analyses are performed on three derived datasets (cross‑product, same‑product, and agent‑authored without AI review) to characterize prevalence and reviewer behavior.
Result
Cross‑product AI‑to‑AI review occurs in about 1.6% of identified agent‑authored PRs, amounting to 45k PRs, and grew by more than two orders of magnitude from 2025‑Q1 to 2025‑Q3. CodeRabbit labeled 35.0% of its comments on Claude‑Code PRs as refactor comments versus 10.5% on Copilot PRs. For three of four dual‑role reviewers, mean comments per PR were 58 to 65% higher in same‑product groups, and median latency was 1.2 minutes for cross‑product pairs versus 4.7 minutes for same‑product pairs.
Why it matters
Researchers and tool builders must account for AI‑to‑AI code review loops when sampling GitHub data and designing automated review systems, as they represent a growing but distinct portion of development activity.
Method details
Data source: CodAGE snapshot covering 2024-01-01 to 2026-04-15.
Attribution: strict signature framework requiring body or vendor‑login evidence.
Dataset size: 248,641 unique AI‑attributed PRs with at least one AI review.
Reviewer analysis: CodeRabbit comment‑category classifier applied to review comments.
Numbers
cross‑product proportion, 1.6%, of identified agent‑authored PRs
cross‑product PRs, 45,269, compared with same‑product PRs 208,145
growth, > two orders of magnitude, from 2025‑Q1 to 2025‑Q3
refactor comment rate, 35.0%, on Claude‑Code PRs vs 10.5% on Copilot PRs
median latency, 1.2 minutes, for cross‑product pairs vs 4.7 minutes for same‑product pairs
Limitations
The study does not establish review correctness or software quality and cannot separate reviewer behavior from PR characteristics due to limited timestamp availability.
Claude-Code PRs receive more refactor comments from CodeRabbit than Copilot PRs (35.0% vs. 10.5%)Found in the source text, word for word.
Picked because: Releases a large‑scale AI‑to‑AI code‑review dataset and analysis pipeline, enabling engineers to build and evaluate self‑hosted AI code‑review bots.
Qisheng Lu, Aoyang Fang, Junjielong Xu and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Prior automated RCA methods only measured endpoint correctness, ignoring the evidentiary basis and fault‑propagation route, so agents could localize the faulty service yet fail to reconstruct its impact path, and simply improving endpoint accuracy does not fix this evidence‑handling gap.
Approach
The paper introduces DiagGuard, a two‑stage defense‑in‑depth architecture that wraps a reasoning core (the ThinkDepth.ai Diagnostician) with a Grounder that surveys all available telemetry before localization and a Verifier that audits the diagnosis against the gathered evidence before committing. The Grounder ensures the agent is under‑grounded, while the Verifier provides inference‑time verification to catch misinterpretations or unsupported inferences. Together they address the three failure families identified in the trajectory analysis. The design is instantiated on the ThinkDepth.ai framework and evaluated on independent benchmarks.
Result
DiagGuard improves the top‑1 accuracy of root‑cause localization from 43.5% to 52.5% in an independent validation setting, demonstrating that trajectory‑level evaluation can expose hidden limitations of agents that achieve correct endpoint answers but poor diagnostic quality.
Why it matters
SREs and AIOps researchers should care because the approach reveals hidden diagnostic failures and provides a concrete defense that measurably improves automated root‑cause analysis for microservices.
Method details
Evaluated six agent frameworks: ThinkDepth.ai, AIQ, TaskWeaver, ClaudeCode, OpenRCA, mABC
Backbone models used include qwen3.5-plus-2026-02-15, claude-sonnet-4.6, doubao-seed-2.0-pro-2026-02-15
Benchmarks: RCABench (500‑case sample from 1,430 cases) and AIOps 2025 (400 incidents over ten services)
Baseline core is the ThinkDepth.ai Diagnostician; DiagGuard adds Grounder and Verifier defenses
DiagGuard was validated with a different model (Seed 2.0 Pro) and benchmark (AIOps 2025)
Numbers
Acc@1 43.5% baseline
Acc@1 52.5% with DiagGuard
3,500 diagnostic trajectories analyzed
500‑case sample from RCABench
400 incidents in AIOps 2025 benchmark
six agent frameworks evaluated
Limitations
The paper does not state any limitations.
In an independent setting with a different model, benchmark, and service topology, DiagGuard raises Acc@1 from 43.5% to 52.5%.Found in the source text, word for word.
Picked because: Presents a trajectory‑level evaluation of LLM agents for microservice root‑cause analysis, with actionable insights and code for integrating such agents into production observability stacks.
Content‑only screening cannot reliably distinguish false assertions from true ones, and the obvious fix of adding a provenance weight fails because any weight strong enough to block a shaped attack also discards legitimate untrusted evidence.
Approach
The paper builds a four‑stage write‑time screening pipeline that flags content if any stage flags, and evaluates provenance‑weighted retrieval where a scalar weight modifies the similarity score. The pipeline combines deterministic regex, model‑based detectors (ProtectAI DeBERTa v2, Meta Llama Prompt Guard 2, LLM Guard), LLM‑as‑judge judges (GPT‑4o‑mini, Claude Haiku 4.5), and Aegis stages. Retrieval is plain nearest‑neighbor without reranking. The authors then argue for a bounded occupancy constraint on provenance rather than an additive weight.
Result
Poisoning 1.2% of the LongMemEval corpus drops accuracy from 0.850 to 0.300, a two‑thirds loss. The four‑stage screening pipeline attains 0.832 recall on indirect injection, flags 1.5% of benign trigger‑word text, and rejects none of 360 poisoned memories. The shipped provenance weight is statistically indistinguishable from no defense (p=0.80), while a stronger weight only improves accuracy to 0.7000 in a mixed‑provenance corpus but collapses it to 0.0417 when evidence is untrusted.
Why it matters
Developers of agent memory systems and security researchers should note that simple content screening and additive provenance weighting are insufficient defenses against memory poisoning.
Method details
Write‑path screening evaluates ten configurations including a naive regex baseline and three model‑based detectors.
Benchmarks use five corpora: direct injection (deepset/prompt‑injections), indirect injection (InjecAgent), Dolly‑15k, templated memory‑like entries, and NotInject.
Memory‑quality benchmark uses LongMemEval_S with 500 questions and ~115K tokens of chat history.
Ablation rescoring is done cumulatively per stage to isolate contribution of injection detection versus secrets detection.
Provenance weight is compared to a no‑defense baseline; statistical indistinguishability is reported with p=0.80.
Numbers
accuracy 0.850 → 0.300 after 1.2% poisoning
recall 0.832 on indirect injection
benign flag rate 1.5%
0 of 360 poisoned memories rejected
p=0.80 for shipped provenance weight vs no defense
accuracy 0.3167 → 0.7000 in mixed‑provenance corpus
Limitations
The study only evaluates a non‑adaptive single‑pass attack on one memory implementation and does not implement or test the proposed bounded‑occupancy provenance constraint.
At 1.2% of the corpus, that attack removes two‑thirds of the memory’s value.Found in the source text, word for word.
Picked because: Demonstrates concrete memory‑poisoning attacks on LLMs and proposes mitigation stages, offering practical verification and security measures for self‑hosted models.
Jan Novacek, Ali Ahari, Tobias Müller and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Existing AI model repositories like Hugging Face provide only semi‑structured natural language descriptions and current ML platforms (MLflow, H2O, Ray) add metadata that lacks interoperability, so assets cannot be reliably exchanged across companies.
Approach
The authors built the AIMDEP platform that registers AI assets and attaches metadata expressed in the AIMDEO ontology. Assets (datasets, models) are uploaded, automatically feature‑extracted, and annotated via the ontology. The ontology metadata is stored as OWL files and can be exported or accessed through a REST API. Search functions locate assets using ontology terms, and the platform can execute models directly for supported frameworks. The overall workflow links domain experts, AI experts, and end users through shared, semantically rich asset descriptions.
Result
The registered AI model’s metadata reports an average precision of 0.9794, demonstrating that the platform can capture and expose quantitative quality metrics for exchanged models.
Why it matters
Industrial AI developers and tool vendors should care because the platform enables semantically rich, interoperable exchange of models and datasets, reducing integration effort.
Method details
Model framework: scikit‑learn
Model name: RF Instruction Cache‑Line Access Classifier
Metric recorded: average precision 0.9794
Dataset includes cache replacement strategy (LRU) and cache size (2048 Byte)
Metadata exported as OWL via the AIMDEO ontology
Numbers
average precision, 0.9794, compared against none
Cache‑Size, 2048, compared against none
Replacement‑Strategy, LRU, compared against none
Limitations
The paper does not evaluate interoperability with other ontologies or provide a quantitative benchmark against existing model‑exchange platforms.
The platform incorporates an ontology that can foster a more profound common understanding of what is required in these tasksFound in the source text, word for word.
Picked because: Proposes and releases an ontology‑based system for managing AI models and datasets, supporting platform engineers in tracking and governing AI assets.