Paper 1 of 5
Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
Problem
Black‑box LLM observers on shared endpoints failed to reproduce rankings: same‑window repeats achieved Spearman 0.400 (required 0.90) and next‑day byte‑identical replays achieved 0.78 (required 0.99). Simple fixes such as metric substitution, scaling the sample grid, waiting, or switching providers did not close the gap.
Approach
The authors conducted preregistered audits with a fully frozen measurement instrument: they captured exact request bytes, configuration snapshots, and gate definitions before any calls. A deterministic simulation of the legacy D2 estimator (500 replicates, seed 20260730) used frozen D1‑S readouts as probability proxies and applied a ranking gate requiring within‑task Spearman median 0.90. Observer calls were issued at multiple call volumes and the outcomes (SE, gate median, gate q95, pass rate) were recorded. Follow‑up experiments (waiting, provider swaps, self‑hosting, constructed errors) probed the mechanisms behind instability. The evidence was distilled into design rules, a snapshot‑identity ladder, and a reporting checklist.
Result
Across 52,988 audited request attempts, same‑window repeat rankings reached Spearman 0.400 versus the required 0.90, and byte‑identical next‑day replays reached 0.78 versus the required 0.99. No configuration passed the gate (0/500 passes) at any call volume. Waiting did not improve stability (0.805 vs 0.800) and provider medians ranged only from 0.74 to 0.88, well below the 0.90 threshold.
Why it matters
Researchers and practitioners who rely on black‑box LLM judges for evaluation, benchmarking, or leaderboard scoring must treat the observer as a noisy instrument and preregister its reliability before freezing any evaluation gate.
Method details
- Deterministic simulation of the legacy D2 estimator with 500 replicates, seed 20260730.
- Frozen D1‑S readouts used as probability proxies.
- Ranking gate defined with within‑task Spearman median 0.90.
- Observer calls evaluated at call volumes 8, 16, 32, 64, 100, 200, 500.
- Four external providers measured, with all exposed metadata fields recorded.
- Self‑hosted batch‑invariant kernel serving tested under quiet load.
Numbers
- Spearman same‑window repeat ranking, 0.400, required 0.90
- Spearman next‑day replay ranking, 0.78, required 0.99
- Pass rate across 500 replicates, 0/500, required >0
- Waiting experiment median Spearman, 0.805, compared to 0.800
- Provider median Spearman range, 0.74‑0.88, compared to required 0.90
- Gate median at 500 calls, 0.566, compared to gate threshold 0.90
Limitations
The study only measures externally observable behavior on shared serving infrastructure and does not identify the internal cause of instability (e.g., batching vs kernel scheduling vs deployment rotation).
same‑window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte‑identical next‑day replays agreed at 0.78 against a required 0.99Found in the source text, word for word.
Picked because: Shows a reproducible audit of LLM judge reliability and releases scripts/data, giving engineers concrete methods to verify and monitor LLM‑based measurement tools.