Suman Raj, Hai Duc Nguyen, Haochen Pan and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.DC
Problem
Existing scientific workflow management systems rely on fixed, hand‑tuned rules, limiting autonomous adaptation; simply replacing these rules with LLMs is not straightforward because the integration points and risk boundaries are unclear.
Approach
Avatar introduces an actor‑based architecture with three components-the orchestrator, the executor, and the provenance monitor-each exposing a pluggable decision policy via a single adapter‑validated action catalog; policies can be rule‑based or LLM‑backed, allowing conventional and agentic control to run on the same core across different WMSs; the system is built on the Academy framework and evaluated on three distinct workloads to demonstrate representativeness, generalizability, and applicability.
Result
LLM‑backed Avatar reduces compute wastage by 55% and cuts GPU‑busy time by 40% compared with the rule‑based configuration, while the rule mode reproduces native workflow behavior across distinct systems.
Why it matters
Workflow system developers and researchers interested in integrating LLMs should care because Avatar shows a modular way to add agentic reasoning that can substantially improve resource efficiency.
Method details
Three actors (orchestrator, executor, provenance) with interchangeable policies
Action catalog validated by adapters to ensure safe execution
Implemented using the Academy framework
Evaluated on three workloads: Resilience (E1), Scaling (E2), Active Learning (E3)
Baselines: TaskVine for E1, Parsl for E2, Colmena for E3
LLM‑backed modes M1 and M2 compared against rule‑based mode M0
Numbers
compute wastage reduction, 55%, compared against rule mode
GPU‑busy time reduction, 40%, compared against rule mode
Limitations
The paper evaluates only three workloads and does not demonstrate performance on other workflow systems or at larger scales.
LLM-backed Avatar reports a reduction of compute wastage by $55\%$ and cuts GPU-busy time by $40\%$.Found in the source text, word for word.
Picked because: Presents Avatar, an actor‑based LLM‑agent architecture that autonomously orchestrates scientific workflows, offering concrete components (orchestrator, executor, provenance monitor) engineers can adopt for end‑to‑end automation.
Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Benchmarks score only advertised model identifiers, ignoring the serving route, precision, output contract, and harness, which leads to measurement error. Simply reporting the model identifier does not work because route-specific limits can change observable capability.
Approach
The paper introduces the IB2 protocol, which consists of three components: a gold‑blind capability‑binding preflight that verifies a route can execute the evaluation contract before any task is sent; a reliability‑inclusive first‑pass scoring rule that retains failures in the score while excluding unsupported capability; and a structurally score‑blind adjudication stage. IB2 also releases algorithms, classification tables, a request contract, and manifest schemas. The protocol treats the measurement procedure as the artifact rather than the task corpus. By binding routes to a strict contract, IB2 makes capability availability reportable.
Result
The protocol measured that changing the serving arm increased precision from 77.38 to 82.54 with a paired interval of [0.11,10.60]. The leader‑minus‑runner‑up difference was 0.97 displayed points, but its task‑set sensitivity interval straddles zero. Out of 55 pairwise interval comparisons, 46 excluded zero, indicating significant separations among most systems.
Why it matters
Enterprises deploying AI services and benchmark designers should care because IB2 reveals how serving routes affect measured capability, which standard model‑only evaluations miss.
Method details
Reference instantiation uses 128 locked tasks and 987 assertions across document, spreadsheet, chart, tool and database work.
The evaluation contract requires 13 images per request, a catalog of 25 tools, a completion budget of 65,536 tokens, and strict JSON output schema.
Serving‑arm choice moved one declared revision and precision from 77.38 to 82.54, paired interval [0.11,10.60].
Two complete single‑route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run.
Four of seven suites saturate under a six‑system band, with spread coming mainly from governed database work and multi‑tab joins.
Qwen3.8‑27B interval margin 1.6% of interval width
128 tasks, 987 assertions
Limitations
The paper does not establish that route limits are the only source of measurement error and acknowledges that the binding gate was incomplete for one case.
Serving-arm choice moved one declared revision and precision from 77.38 to 82.54, paired interval [0.11,10.60]Found in the source text, word for word.
Picked because: Introduces the IBIB protocol for measuring AI services by their serving route rather than model ID, providing a practical framework and tooling to verify production deployments.
Bokang Zeng, Zheng Gao, Xiaoyu Li and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CR
Problem
Prior behavioral watermarking only gave a global ownership signal and could not reveal which part of a trajectory was altered; simply adding visible actions would break the agent's behavior and does not expose local inconsistencies.
Approach
TrajMark introduces two complementary layers: a sparse owner layer that rewrites a keyed subset of natural READ actions into masked linear equations encoding a six‑bit identifier, and a localization layer that inserts linked Q12 ordinary, group, and terminal seals to commit to protected critical‑action segments. The owner layer provides robust, aggregated ownership evidence without adding actions, while the integrity layer adds overt read‑only seals that are fragile to local edits. During generation the wrapper rewrites eligible READs and emits seals before the next decision, enabling the verifier to solve for the owner ID and to localize any tampered segment. This separation allows ownership to accumulate across trajectories and local modifications to perturb nearby keyed commitments, exposing the affected protocol region.
Result
TrajMark recovers the exact owner in all evaluated clean full‑watermark batches, detects 95.5%‑100% of edits under exhaustive single‑site attacks, localizes 95.8% of modified sites to an accepted protocol region, and achieves Pass@1 of 26.9% versus 26.3% for unwatermarked runs.
Why it matters
Developers of coding agents and provenance systems can use TrajMark to obtain robust ownership attribution and fine‑grained tamper localization without degrading task performance.
Method details
Evaluated three coding‑agent frameworks: SWE‑agent, OpenHands, and OpenDev.
Three LLM providers: DeepSeek V4 Flash, GPT‑5 mini, and MiniMax M3.
Task families: SWE‑bench Python, SWE‑PolyBench Java, and SWE‑PolyBench JavaScript, each with 50‑instance lists.
Baselines compared: No‑WM control, ActHook‑style, AgentMark‑U, and Owner‑only ablation.
TrajMark is training‑free, symmetric‑key, and operates by rewriting READ actions and emitting Q12 seals during inference.
Numbers
owner recovery, 100%, exact owner recovered in all clean full‑watermark batches
edit detection, 95.5%-100%, under exhaustive eligible single‑site attacks
localization accuracy, 95.8%, of modified sites to an accepted protocol region
Pass@1, 26.9%, versus 26.3% for unwatermarked runs
deployment identifier size, six‑bit, used for owner channel
Limitations
The mechanism provides no public verifiability or non‑repudiation and does not protect against a dishonest embedding service, verifier, or key compromise.
detects 95.5%-100% of edits, and under random single-action corruption it localizes 95.8% of modified sites to an accepted protocol region rather than to the individual action.Found in the source text, word for word.
Picked because: Describes TrajMark, a behavioral watermarking system for coding‑agent trajectories that enables fine‑grained ownership attribution and tamper detection, useful for LLM‑agent verification.
Ivana Clairine Irsan, Ratnadira Widyasari, Huihui Huang and 6 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Static analysis tools like CodeQL require extensive manual effort to craft high‑coverage query suites, and directly applying LLMs to scan whole repositories is computationally and financially prohibitive.
Approach
The method mines vulnerability patterns from CVE entries in the MoreFixes dataset, then uses large language models to synthesize CodeQL queries from those patterns. A Syntax Refiner Agent iteratively repairs uncompilable queries up to three attempts. A hybrid configuration employs Gemini 3 Flash Preview as a syntax specialist to refine outputs from more expensive models such as GPT 5.2 Codex and Claude Sonnet 4.5. The resulting queries are executed with CodeQL v2.17.3 using the no‑build mode. This pipeline replaces manual query authoring with automated LLM‑driven generation and refinement.
Result
Kimi K2.5 generated queries achieved the highest average F1‑score of 24.04% and a detection rate of 44.64%, representing an 82% improvement over the CodeQL baseline average F1 of 13.20% and a three‑fold increase in detection rate.
Why it matters
Security engineers and static analysis tool developers should care because LLM‑generated CodeQL queries markedly boost vulnerability detection while keeping costs manageable.
Method details
LLM architectures evaluated include Gemini 3 Flash Preview, DeepSeek R1, Grok Code Fast 1, Kimi K2 Thinking, Kimi K2.5, Llama 3.3, Minimax M2.1, Qwen 3 Coder, GPT 5.2 Codex and Claude Sonnet 4.5
Dataset comprises CVEs from the MoreFixes collection filtered to 10 CWE categories, yielding 80 mining seeds and 112 test CVE‑code segment pairs
Baseline comparisons are against the default CodeQL query set and a PDBERT model
Syntax Refiner Agent performs up to three iterative repair attempts per query
Hybrid runs pair Sonnet 4.5 or GPT 5.2 Codex with Gemini 3 Flash Preview for syntax correction
Implementation uses OpenRouter API for model access and monitors monetary cost per experiment
Numbers
Avg F1 24.04% (Kimi K2.5) vs 13.20% (CodeQL baseline)
Detection Rate 44.64% (Kimi K2.5) vs 16.96% (CodeQL)
Compilable Queries 123 (Gemini 3 FP with syntax refiner) vs 17 (without)
The study is limited to Java repositories and does not demonstrate cross‑language transferability of the mined patterns.
Kimi K2.5 achieved the highest overall average F1-score of 24.04%.Found in the source text, word for word.
Picked because: Shows how large language models can automatically generate executable CodeQL security queries, delivering a reproducible pipeline for scalable vulnerability detection.
quote verifiedfigures checkedread: full textq-fin.TR
Problem
Researchers had to manually decode Uniswap logs, resolve token metadata, convert integer quantities and join execution metadata, leading to inconsistent units, signs and ordering. Using generic extraction tools does not solve this because they still require these protocol‑specific transformations.
Approach
dexamine provides a reusable Python session that orchestrates four components: a JSON‑RPC client retrieves transaction receipts and block data; a metadata resolver caches pool and token information via web3.py calls; protocol parsers decode Uniswap v2 and v3 event topics and data fields; and an output layer joins parsed events with execution metadata into flat rows. The session batches RPC requests, yields results incrementally, and can be run offline with recorded responses. Users supply block number, transaction index, protocol choice and optional pool address, and receive a table of events with receipt_log_index, token amounts, price, tick and virtual reserves.
Result
The parser successfully extracted four Uniswap v3 swaps from transaction 31 in block 12,561,528, ordered them by receipt_log_index, and reported post‑swap prices, ticks and virtual reserves as shown in Tables 1 and 2.
Why it matters
Empirical finance researchers who need consistent, reproducible Uniswap event data for microstructure analysis should use dexamine to avoid manual decoding and unit inconsistencies.
Method details
Supports Uniswap v2 and v3 swaps, mints and burns on Ethereum mainnet.
Version v1.1.0 is installed via pip from the GitHub repository.
Outputs can be flat rows including block/transaction indices, receipt_log_index and post‑event pool state.
Metadata cache grows with distinct contracts; batch size and endpoint latency affect processing rates.
Numbers
Price, 2802.766431, USDC per WETH after swap at receipt position 3
Tick, 196936, USDC/WETH pool after swap at receipt position 3
Virtual reserve 0, 137336604.311, USDC/WETH pool after swap at receipt position 3
Virtual reserve 1, 49000.374, USDC/WETH pool after swap at receipt position 3
Price, 0.998748, DAI per USDC after swap at receipt position 6
Tick, (none), DAI/USDC pool after swap at receipt position 6
Limitations
The library does not allocate gas costs to individual trades and cannot reconstruct internal call paths or trader identity.
Table 1 shows the four swaps in receipt order.Found in the source text, word for word.
Picked because: Releases dexamine, a Python package that parses Uniswap event data on Ethereum, giving engineers a ready‑to‑use tool for building self‑hosted blockchain analytics pipelines.