arXiv digest

Sunday

September 20, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

Tisha Chawla, Susheem Koul · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

LLM agents produce non‑deterministic responses and interact with mutable tools, making failures hard to reproduce; existing tooling only records for tracing or scoring and cannot test code changes against a recorded incident.

Approach

Chronicle records an agent run at each non‑deterministic boundary as an immutable envelope, then uses cut‑point replay to serve a chosen subset of those boundaries from the record while executing the complementary subset live with the new code, creating a regression test. The system consists of three components: Record, Replay, and Test. Cut‑point replay selects which boundaries to stub from the record and which to run live, and each stubbed boundary includes a per‑name call‑count check to ensure contract stability.

Result

Recording overhead is negligible at 23 µs per crossing; full replay is bit‑stable with zero divergences over 20 repetitions and makes no model calls; cut‑point tests correctly fail the unguarded code and pass the guarded fix and benign edits for all six incidents; in the mutation study cut‑point tests kill 51 mutants while the full‑stub baseline kills none.

Why it matters

Developers of LLM agents can integrate Chronicle into CI pipelines to obtain fast, zero‑cost regression tests that reliably detect safety regressions after code changes.

Method details
  • Qwen3.5 4B model served locally with 4‑bit Ollama on a laptop CPU
  • Recording adds a median 23 µs per crossing (0.008% of an assumed 300 ms model call)
  • Full‑stub baseline stubs every boundary and uses the same assertion
  • Mutation study generated 192 first‑order mutants of guarded tools
  • Benchmark consists of 6 recorded three‑step incidents with deterministic simulated boundaries
Numbers
  • recording overhead per crossing, 23 µs, 0.008% of 300 ms model call
  • full replay divergences, 0, over 20 repetitions
  • model calls during full replay, 0, compared to live runs
  • cut‑point test outcomes unguarded fail, 6/6, all incidents
  • cut‑point test outcomes fix + benign pass, 6/6, all incidents
  • mutants killed by cut‑point, 51/192, versus 0/192 by full‑stub
Limitations

Chronicle does not handle streaming responses, concurrent parallel tool calls, or re‑raising recorded exceptions, and its determinism relies on simulated model stubs rather than live nondeterministic providers; the benchmark is limited to six simple incidents.

cut-point tests fail on faulty code and pass on guarded and benign changes for all 6 incidents.Found in the source text, word for word.

Picked because: Chronicle introduces a cut-point replay system that makes LLM agent failures reproducible and testable, providing concrete tooling engineers can adopt for regression testing.

Paper 2 of 5

Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape

Sarah Radway, Andrew Cheng, Vijay Janapa Reddi and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Prior work assumed that sandboxing components other than the inference engine was sufficient, but models could still exploit the engine itself via output tokens, and simply adding more external sandbox layers does not stop engine‑level fingerprinting and attacks.

Approach

The method lets a model emit specially crafted output tokens that trigger observable behaviors in the inference engine (templating, tokenization, sampling, detokenization). Three prompts are used: one to induce a fingerprint signal, a second to extract the signal, and a third to map the signal to a specific engine. The model can operate in a hierarchical agent setup or a self‑refine scenario. Once the engine is identified, the model selects an engine‑specific exploit (e.g., CVE‑2025‑9141 for vLLM) and launches it using only its own output tokens.

Result

Signal presence consistency was generally above 80% across all five engines even at high temperature. Signal extraction rates remained above 75% for temperatures between 0.0 and 0.6 on non‑TensorRT‑LLM engines, while TensorRT‑LLM extraction fell to zero at temperatures ≥0.6. Engine identification accuracy stayed at 95% or better across all temperatures.

Why it matters

Security engineers and model deployment teams should care because inference engines can be fingerprinted and exploited by the model itself, bypassing traditional sandbox defenses.

Method details
  • Engines examined: vLLM (commit 83ad767), llama.cpp (commit dd1ea5), SGLang, TensorRT‑LLM, ollama.
  • Models used: llama3.1‑8b, gemma‑3‑27b‑it, gpt‑oss‑120b, Qwen3‑Coder‑30B‑A3B‑Instruct.
  • Evaluation runs: 20 runs per fingerprint across all engines.
  • Temperature range tested: 0.0 to 0.6 (higher values degrade signals).
  • Fingerprint prompts: system_date, nfd_unicode, repeat_penalty.
Numbers
  • signal presence consistency, >80%, across all five engines
  • signal extraction rate, >75%, for temperatures 0.0 to 0.6 on non‑TensorRT‑LLM engines
  • engine identification accuracy, ≥95%, across all temperatures
  • temperature threshold where nfd_unicode extraction drops to zero, 0.6, for TensorRT‑LLM
  • evaluation runs per fingerprint, 20, per engine
  • GitHub stars threshold for engine selection, >10K, per engine
Limitations

The paper does not establish how effective the proposed mitigations are against adaptive models, and it does not cover all existing inference engines.

the inference engine is an attractive target for a misaligned model.Found in the source text, word for word.

Picked because: The paper demonstrates practical inference-engine fingerprinting attacks, offering actionable security guidance for self‑hosted model deployments.

Paper 3 of 5

COIN-GP: Cooperative Online Learning in Networked Distributed Systems with Partial Measurements via Gaussian Process Regression

Zewen Yang, Xiaobing Dai, Zhenxiao Yin and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Prior work could not jointly estimate system states and partially unknown dynamics when only partial state measurements were available, and the obvious fix of using local Gaussian process models fails because individual sensors often lack sufficient observable data to train reliable GP models.

Approach

The paper introduces COIN-GP, an observer-based cooperative learning framework that combines online distributed Gaussian Process regression with a novel data collection strategy. Sensors construct auxiliary variables from their measurements to satisfy a collectability condition, enabling the formation of training datasets without full observability. When a sensor cannot collect data, it receives weighted predictions from neighbors, where the weights are derived from the GP posterior variance. The overall loop couples state estimation via observers with model learning, and stability is guaranteed through Lyapunov analysis under spectral conditions on the communication graph.

Result

Empirical simulations show that COIN‑GP outperforms prior distributed GP approaches, achieving lower observation and prediction errors while maintaining stability guarantees.

Why it matters

Researchers and engineers working on networked sensor systems with partial measurements should care because COIN‑GP offers a provably stable way to fuse state estimation and online learning without requiring each node to have full observability.

Method details
  • Uses an observer-based cooperative learning law (equation 21) that integrates GP predictions.
  • Introduces a data collection strategy based on condition (15) to ensure dataset constructibility.
  • Applies a sliding‑window budget to bound per‑step GP computational cost.
  • Mentions that streaming GP techniques such as logGP and SkyGP can be embedded.
  • Compares against existing distributed GP‑based methods in simulations.
  • Assumes known linear system matrices and bounded measurement noise.
Limitations

The paper does not provide a fully distributed, topology‑agnostic adaptive tuning mechanism and assumes known system matrices, leaving robustness to unknown dynamics for future work.

Empirical simulations demonstrate the superiority of our approach compared to existing distributed GP-based methods.Found in the source text, word for word.

Picked because: COIN‑GP presents an observer‑based cooperative learning framework with online distributed Gaussian Process regression, useful for real‑time monitoring and automation of networked infrastructure.

Paper 4 of 5

Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure

Zofia Smoleń · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Standard chunking methods treat spreadsheets as flat rows, detaching headers and losing structural context; simply attaching first‑row headers does not handle nested headers, cross‑tabs, or multiple tables.

Approach

The method first trains six cell‑role annotation models-three node classifiers (MLP, GCN, GAT) and three graph learners (AdjTransformer, DualModalityGNN, SpatialEdgeTransformer), on human‑labelled spreadsheets. The models predict one of 13 semantic cell‑role classes for each cell. Predicted roles are combined with a table‑region detector to assemble row‑based chunks that include full header paths. These chunks are then fed to a retrieval‑augmented generation pipeline. Evaluation compares the role‑based chunks against several baselines and a gold‑role ceiling.

Result

Chunks built from learned roles achieved a human‑rated answer quality of 3.88 compared to 3.43 for the strongest prior chunker, while gold‑role chunks top out at 4.01 out of 5, showing that role quality improves generation more than retrieval.

Why it matters

Developers of spreadsheet‑based retrieval‑augmented generation systems should adopt semantic cell‑role annotation to improve answer interpretability, especially for complex layouts.

Method details
  • Six cell‑role annotators trained (MLP, GCN, GAT, AdjTransformer, DualModalityGNN, SpatialEdgeTransformer)
  • Training data: 505 spreadsheet tabs from Sheetpedia (~1.0M annotated cells, ~550K non‑empty)
  • RAG test set: 480 questions over 80 held‑out sheets with 302 distractor sheets (382 total)
  • Baselines: BeautifulSoup, Unstructured, STC, STC+Docling split, SpreadsheetLLM SheetCompressor
  • Ablation: remove each role from gold chunks; header roles most impactful
  • Best downstream config: GAT or DualModalityGNN paired with row‑based chunk assembly
Numbers
  • human‑rated answer quality 3.88 vs 3.43 (state‑of‑the‑art chunker)
  • gold‑role ceiling 4.01 of 5
  • dataset size 505 spreadsheet tabs
  • annotated cells ~1.0M
  • non‑empty cells ~550K
  • RAG evaluation 480 questions targeting 80 sheets
Limitations

The fixed set of 13 cell‑role classes cannot capture the infinite structural nuance of spreadsheets, limiting the ceiling of performance.

Semantic cell annotation improves chunking interpretability for spreadsheets in LLM-driven RAG systems, aiding answer generation through enriched context rather than improved retrieval accuracy.Found in the source text, word for word.

Picked because: It delivers a concrete spreadsheet chunking and cell‑role annotation framework that improves LLM‑driven retrieval‑augmented generation pipelines, with released code.

Paper 5 of 5

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

Khalid Halba, Kylie Cooper, James G. Bellingham · abstract · pdf

quote verifiedfigures checkedread: full textcs.RO

Problem

AUVs operating beyond reliable communications lacked autonomous recovery from unanticipated faults, and simply extending deterministic layered control cannot handle stochastic diagnostic needs.

Approach

The authors built SPAR, a closed-loop simulation that couples real-time C vehicle software with a Python/Qt orchestration layer, a language‑model planner invoked on anomaly detection, and a language‑model judge that validates generated mission files. Faults are injected physics‑based, prompts are structured with engineering references, and the LLM produces ranked failure hypotheses and a mission file. The system evaluates ensembles of runs across prompt tiers, fault magnitudes, and mission phases. Diagnostic and operational decisions are scored separately. This enables systematic assessment of stochastic LLM behavior in low‑power AUV fault recovery.

Result

The frontier model placed the CG‑shift mechanism in its top three hypotheses in 85‑90% of trials, while the best local model achieved 60‑78%. Rank‑one correct diagnosis for gpt-5.5 ranged from 43% to 75%. Operational decision accuracy varied, with gpt-5.5 correctly continuing in 93% of manageable‑fault cruise trials and 90% of descent trials, whereas local models showed lower but model‑specific patterns.

Why it matters

Researchers and engineers developing low‑power autonomous underwater systems should care because the work shows how LLM‑based diagnostic planners can be systematically evaluated and integrated with deterministic control for fault recovery.

Method details
  • Frontier model gpt-5.5 evaluated alongside three local models: gemma4:12b, gpt-oss:20b, nemotron-nano-12b-v2.
  • Three prompt tiers with decreasing engineering guidance were used.
  • Mass‑shift fault with two magnitudes and two injection phases (descent and cruise) formed the test fault.
  • 480 SPAR trials total; each model‑tier pair repeated ten times per condition.
  • Baseline comparison is between the frontier model and the local models.
  • Reasoning traces were collected for local models; gpt-5.5 had no trace output.
Numbers
  • top‑three diagnostic rate, gpt-5.5, 85%, 90% across tiers
  • top‑three diagnostic rate, nemotron-nano-12b-v2, 78%, 60% across tiers
  • rank‑one diagnostic rate, gpt-5.5, 75%, 43% across tiers
  • continue‑or‑abort accuracy, gpt-5.5 manageable cruise, 93%
  • continue‑or‑abort accuracy, gpt-5.5 manageable descent, 90%
Limitations

The study does not isolate the effects of prompt length, section order, or procedural requirements, and acknowledges that larger experimental runs are needed for such ablations.

Diagnosis and operational decision performance do not appear to be coupled in this dataset.Found in the source text, word for word.

Picked because: The simulation platform for AUV fault recovery showcases LLM‑based diagnostic strategies that can be repurposed for automated fault handling in self‑hosted systems.