Paper 1 of 5
Chronicle: Cut-Point Replay for Regression Testing of LLM Agents
Problem
LLM agents produce non‑deterministic responses and interact with mutable tools, making failures hard to reproduce; existing tooling only records for tracing or scoring and cannot test code changes against a recorded incident.
Approach
Chronicle records an agent run at each non‑deterministic boundary as an immutable envelope, then uses cut‑point replay to serve a chosen subset of those boundaries from the record while executing the complementary subset live with the new code, creating a regression test. The system consists of three components: Record, Replay, and Test. Cut‑point replay selects which boundaries to stub from the record and which to run live, and each stubbed boundary includes a per‑name call‑count check to ensure contract stability.
Result
Recording overhead is negligible at 23 µs per crossing; full replay is bit‑stable with zero divergences over 20 repetitions and makes no model calls; cut‑point tests correctly fail the unguarded code and pass the guarded fix and benign edits for all six incidents; in the mutation study cut‑point tests kill 51 mutants while the full‑stub baseline kills none.
Why it matters
Developers of LLM agents can integrate Chronicle into CI pipelines to obtain fast, zero‑cost regression tests that reliably detect safety regressions after code changes.
Method details
- Qwen3.5 4B model served locally with 4‑bit Ollama on a laptop CPU
- Recording adds a median 23 µs per crossing (0.008% of an assumed 300 ms model call)
- Full‑stub baseline stubs every boundary and uses the same assertion
- Mutation study generated 192 first‑order mutants of guarded tools
- Benchmark consists of 6 recorded three‑step incidents with deterministic simulated boundaries
Numbers
- recording overhead per crossing, 23 µs, 0.008% of 300 ms model call
- full replay divergences, 0, over 20 repetitions
- model calls during full replay, 0, compared to live runs
- cut‑point test outcomes unguarded fail, 6/6, all incidents
- cut‑point test outcomes fix + benign pass, 6/6, all incidents
- mutants killed by cut‑point, 51/192, versus 0/192 by full‑stub
Limitations
Chronicle does not handle streaming responses, concurrent parallel tool calls, or re‑raising recorded exceptions, and its determinism relies on simulated model stubs rather than live nondeterministic providers; the benchmark is limited to six simple incidents.
cut-point tests fail on faulty code and pass on guarded and benign changes for all 6 incidents.Found in the source text, word for word.
Picked because: Chronicle introduces a cut-point replay system that makes LLM agent failures reproducible and testable, providing concrete tooling engineers can adopt for regression testing.