Ahmed Hereiz, Yingzhe Lyu, Hao Li and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Prior to this work, the maintenance and co‑evolution dynamics of Claude Code plugins were completely undocumented, and treating these plugins as conventional open‑source packages fails to capture their mixed natural‑language and script components.
Approach
The authors assembled a corpus of 1,926 GitHub repositories containing 8,351 plugins and 77,773 commits. They traced repository and component creation dates, classified plugin functionality with a Qwen model, and analyzed composition by counting declared components. Commit activity was measured with Mann‑Kendall and Sen's slope tests, while commit types were labeled using a two‑stage CCS pipeline aided by GPT‑5‑mini and manual validation. Finally, they examined cross‑component coupling via association‑rule mining on merged pull requests.
Result
Plugin‑touching commit activity grew 8.8× over the six months after the October 2025 launch. Feature commits accounted for 39.6% of commits versus 17.2% in conventional OSS, and Claude co‑authored 34.9% of all commits. Within skills directories, natural‑language instruction files and implementation scripts co‑evolved at above‑chance rates, with 78% of co‑changes being functionally coupled.
Why it matters
Researchers and practitioners building AI coding agents and plugin marketplaces should care because the study reveals distinct maintenance patterns and co‑evolution risks that differ from traditional software engineering.
Cross‑component coupling: association‑rule analysis on 10,679 plugin pull requests from 2,592 plugins.
Trend analysis: Mann‑Kendall test on six monthly commit counts with Sen’s slope estimation.
Numbers
commit activity growth, 8.8x, compared to launch baseline
feature commit share, 39.6%, versus 17.2% in traditional OSS
Claude co‑authored commits, 34.9%, of all commits
software‑engineering plugins, 61.3%, of total plugins
functionally coupled co‑changes, 78%, of observed co‑changes
Claude co‑authored commits, 34.9%, of all commits
Limitations
The paper does not state any limitations.
plugin-touching commit activity growing 8.8x over six months after the October 2025 launchFound in the source text, word for word.
Picked because: Provides an empirical analysis of how AI coding agent plugins evolve and are maintained, offering concrete insights and best‑practice recommendations for building and operating self‑hosted LLM‑powered tooling ecosystems.
Jingjing Nie, Jiawei Guo, Krishna Meda and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CR
Problem
There was no systematic understanding of how LLM‑based security agents are built, used, and assessed; the term "agent" was applied inconsistently and assessment protocols were incomparable.
Approach
The authors performed a systematic literature review of peer‑reviewed work from 2023 to 2026, defining a taxonomy that captures agent architecture, perception, memory, reasoning, planning, action space, orchestration, and self‑improvement, as well as applications and assessment protocols. They searched the literature with keyword queries and forward/backward snowballing, filtered papers by explicit inclusion/exclusion criteria, and ended with 100 selected studies. Each paper was then mapped (attributed) to the taxonomy items to summarize patterns across the field. The attribution results were analyzed to highlight gaps in modularity, safety, and audibility. The survey synthesizes these findings to outline limitations and future research directions.
Result
The review found that 91% of papers use a modular pipeline and 88% place the backbone LLM in a planner+actor role, indicating widespread modular design. Only 15% include a critic/verifier/judge component and 13% add a guardrail or policy‑enforcement layer, showing limited safety mechanisms. Single‑agent systems appear in 54% of works while multi‑agent systems appear in 46%, and mixture‑of‑agents designs are rare at 4%. These numbers reveal that agents can act but lack bounded authority and auditable behavior.
Why it matters
Researchers and practitioners developing LLM‑based security agents should care because the survey highlights critical gaps in safety, governance, and modular verification that need to be addressed.
Method details
Systematic literature search covering Jan 2023, Mar 2026.
Collected 100 peer‑reviewed papers after eligibility filtering.
Derived a three‑dimensional taxonomy (approach, application, assessment).
Performed paper attribution mapping each study to taxonomy items.
Analyzed attribution to identify architectural and assessment trends.
Numbers
modular pipeline, 91%, of surveyed papers
planner+actor, 88%, of surveyed papers
critic/verifier/judge, 15%, of surveyed papers
guardrail/policy enforcement layer, 13%, of surveyed papers
single‑agent, 54%, of surveyed papers
multi‑agent systems, 46%, of surveyed papers
Limitations
The paper does not establish agents with bounded authority or fully auditable behavior.
Table 4 shows that 91% of the surveyed papers use a modular pipeline, and 88% place the backbone LLM in a planner+actor role.Found in the source text, word for word.
Picked because: Describes practical designs, open‑source implementations, and evaluation of LLM‑based agents that automate software‑security workflows, giving engineers actionable patterns for integrating LLM agents into DevOps pipelines.
Remediation throughput is often the bottleneck, not the accuracy of vulnerability prioritisation, and simply improving ranking does not alleviate overload because most issue‑tracker queues operate at or above capacity.
Approach
The paper models remediation as a flow‑control problem. It constructs queues from Jira projects and Mozilla components, estimates utilisation, and classifies queues as supercritical or draining. It evaluates predictive models of queue context, compares them to simple project‑level baselines, and tests flow‑control interventions such as severity‑first sequencing and capacity reservation. Owner‑level analysis assesses whether spare capacity can be transferred across expertise‑connected owners. The overall method combines observational queue diagnostics, predictive discrimination, and counterfactual simulations of capacity allocation.
Result
Queue‑context models only achieve moderate discrimination (AUC 0.66 to 0.69) and are matched by simple baselines. Severity‑to‑speed discrimination varies across systems (AUC 0.505 to 0.785). Flow‑control interventions show larger effects: overloaded‑to‑draining transitions cut resolution time by 90 to 116 days, severity‑first sequencing cuts critical‑item dwell by 227 days in Apache and 59 days in Mozilla, and reserving 25.4% of capacity reduces mean sojourn to 137.5 days versus the observed 265 days.
Why it matters
Security and remediation teams should treat vulnerability remediation as a capacity‑allocation problem, using queue diagnostics and capacity‑reservation strategies to improve turnaround rather than relying solely on better prioritisation.
Method details
Datasets: 2,000 Apache Jira issues (1,672 resolved), 2,000 Mozilla Bugzilla Core bugs, 163 Red Hat CVE records, and an npm dependency graph of 1,021 packages.
Cross‑organisation sample: five public Jira organisations (Apache, MariaDB, Qt, MongoDB, Red Hat) each contributing 2,000 issues.
Predictive evaluation uses AUC on unseen data with bootstrap confidence intervals; queue‑context models achieve AUC 0.66 to 0.69.
Severity‑to‑speed discrimination measured by AUC: Apache 0.505, Mozilla 0.643, Red Hat 0.785.
Numbers
AUC 0.66 to 0.69, queue‑context models vs simple project‑level baselines
Severity‑to‑speed AUC Apache 0.505, Mozilla 0.643, Red Hat 0.785
Resolution days reduction 90 to 116 for overloaded‑to‑draining transitions
Critical‑item dwell reduction 227 days (Apache) and 59 days (Mozilla) with severity‑first sequencing
Capacity reservation 25.4% yields mean sojourn 137.5 days compared with observed 265 days
Owner‑level borrowing changes feasibility 0.0 to 85.3 percentage points across five Jira organisations
Limitations
The study is observational, relies on proxy measures, and its counterfactual analyses (e.g., draining‑transition, fast‑track replay) are based on limited treated queues and strong queueing assumptions.
severity-first sequencing reduces critical-item dwell by 227 days in Apache and 59 days in Mozilla.Found in the source text, word for word.
Picked because: Presents a data‑driven capacity‑allocation model for vulnerability remediation, backed by measurements from real issue‑tracking systems, enabling engineers to prioritize and automate remediation processes.
Pietro Tiberi, Gabriele Marcelli, Vitangelo Lasorella · abstract · pdf
quote verifiedfigures checkedread: full textcs.CR
Problem
Standard ERC‑20 transfers on a permissioned ledger expose sender, receiver and amount, revealing commercially sensitive bilateral flows; naive encryption would require trusted off‑chain note custody which defeats the trustless goal.
Approach
The protocol runs on Hyperledger Besu with QBFT consensus and implements three operations-shield, confidential transfer and unshield-using Groth16 zero‑knowledge proofs over BN254, Poseidon hash commitments stored in an incremental Merkle tree, multi‑recipient ECIES encryption for payloads, and an on‑chain NoteRegistry contract that appends encrypted notes to the ledger, eliminating off‑chain custodians while keeping the sender publicly identifiable.
Result
Proof verification incurs about 1 ms of node execution time (≈220k gas), proof generation requires 4 to 12 s on commodity ARM hardware, and the full settlement flow completes in 8 to 16 s, demonstrating practical feasibility on a five‑node deployment.
Why it matters
Regulators and banks designing permissioned CBDC settlement networks should consider this approach to obtain receiver and amount confidentiality while preserving sender accountability without relying on trusted off‑chain services.
Method details
Groth16 proving system over the BN254 curve
Poseidon hash used for note commitments in an incremental Merkle tree
ECIES multi‑recipient hybrid encryption for note payloads
Proof verification costs ~1 ms (220k gas) on‑chain
Proof generation takes 4 to 12 s on commodity ARM hardware
End‑to‑end settlement latency measured at 8 to 16 s
Numbers
proof verification time ~1 ms (220k gas)
proof generation time 4 to 12 s on commodity ARM hardware
end‑to‑end settlement latency 8 to 16 s
Limitations
Receiver confidentiality is not fully achieved in the proof‑of‑concept because the NoteRegistry is owner‑indexed, and the trusted setup uses a single contributor.
ZK proof verification takes approximately 1 ms of node execution time (220k gas) while proof generation takes 4 to 12 s in software on commodity ARM hardware, with end‑to‑end settlement completing in 8 to 16 sFound in the source text, word for word.
Picked because: Introduces a zero‑knowledge protocol for confidential inter‑bank settlement on permissioned Ethereum, with a released reference implementation, relevant for teams building self‑hosted blockchain or secure transaction services.
Hanzhang Jia, Liheng Zeng, Hao Cheng and 2 others · abstract · pdf
quote verified4 figures not in sourceread: full textcs.AI
Problem
Before this work, all plugins were hosted in a single process, making the process a single point of failure; a crash, memory exhaustion, or blocked event loop terminated every component and every co‑resident session at once. The obvious fix of restarting the process to recover a plugin also tears down all dependent components, so it does not provide isolation or fault tolerance.
Approach
Logos moves each plugin into its own operating‑system process and connects them via a shared bus. A Go router maintains only a routing table, while Python harnesses drive model inference and Node.js tools provide capabilities. All cross‑process state is stored in an append‑only transcript that records every step of a session. After any process dies, a new process rebuilds the session from the transcript (cold switching). The design is justified by four lemmas that show the spatiotemporal‑composability calculus holds across process boundaries.
Result
Measurements on one machine show twelve sessions resumed through six kills and eighty further sessions resumed after kills at the four boundaries of the tool‑call cycle with no repeated effect. A same‑fault comparison shows that in the single‑process configuration one fault interrupts every co‑resident session, whereas under the peer‑process construction one fault ends at a single node. Stress testing with 3,500 calls from up to two hundred concurrent callers exhibited no loss, duplication, or misattribution.
Why it matters
Researchers and engineers building LLM‑agent infrastructures should care because Logos demonstrates a fault‑tolerant, cross‑process composition that avoids the single‑process failure domain of existing frameworks.
Method details
Router implemented in Go, holding only a routing table.
Harnesses run in Python and invoke stateless language‑model inference.
Tools run as Node.js peers, each providing a capability as a separate process.
Shared mutable state is limited to an append‑only transcript owned by no process.
Baseline for comparison is a single‑process reference configuration.
Numbers
sessions resumed without repeated effect, 80, after kills at four boundaries of the tool‑call cycle
sessions resumed, 12, through six kills on a single machine
calls processed, 3,500, with no loss, duplication, or misattribution
concurrent callers, 200, with no loss
registration wins, 1, out of 100 simultaneous claims
explicit denials, 99, out of 100 simultaneous claims
Limitations
The paper does not yet provide a formal semantics of loss, partition, and reconnection, limiting its guarantees to the measured scenarios.
Eighty sessions resume with no repeated effect after kills placed at the four boundaries of the tool-call cycleFound in the source text, word for word.
These figures do not appear in the source text: 100, 12, 80, 99. Treat them as unverified.Number check failed.
Picked because: Defines a formal cross‑process bus for dynamically composing agent capabilities, along with a prototype harness, giving engineers a concrete framework for building verifiable, modular LLM‑agent systems.