arXiv digest

Sunday

September 6, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

Haoyaun Zhu, Jie Zhang · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Black‑box LLM observers on shared endpoints failed to reproduce rankings: same‑window repeats achieved Spearman 0.400 (required 0.90) and next‑day byte‑identical replays achieved 0.78 (required 0.99). Simple fixes such as metric substitution, scaling the sample grid, waiting, or switching providers did not close the gap.

Approach

The authors conducted preregistered audits with a fully frozen measurement instrument: they captured exact request bytes, configuration snapshots, and gate definitions before any calls. A deterministic simulation of the legacy D2 estimator (500 replicates, seed 20260730) used frozen D1‑S readouts as probability proxies and applied a ranking gate requiring within‑task Spearman median 0.90. Observer calls were issued at multiple call volumes and the outcomes (SE, gate median, gate q95, pass rate) were recorded. Follow‑up experiments (waiting, provider swaps, self‑hosting, constructed errors) probed the mechanisms behind instability. The evidence was distilled into design rules, a snapshot‑identity ladder, and a reporting checklist.

Result

Across 52,988 audited request attempts, same‑window repeat rankings reached Spearman 0.400 versus the required 0.90, and byte‑identical next‑day replays reached 0.78 versus the required 0.99. No configuration passed the gate (0/500 passes) at any call volume. Waiting did not improve stability (0.805 vs 0.800) and provider medians ranged only from 0.74 to 0.88, well below the 0.90 threshold.

Why it matters

Researchers and practitioners who rely on black‑box LLM judges for evaluation, benchmarking, or leaderboard scoring must treat the observer as a noisy instrument and preregister its reliability before freezing any evaluation gate.

Method details
  • Deterministic simulation of the legacy D2 estimator with 500 replicates, seed 20260730.
  • Frozen D1‑S readouts used as probability proxies.
  • Ranking gate defined with within‑task Spearman median 0.90.
  • Observer calls evaluated at call volumes 8, 16, 32, 64, 100, 200, 500.
  • Four external providers measured, with all exposed metadata fields recorded.
  • Self‑hosted batch‑invariant kernel serving tested under quiet load.
Numbers
  • Spearman same‑window repeat ranking, 0.400, required 0.90
  • Spearman next‑day replay ranking, 0.78, required 0.99
  • Pass rate across 500 replicates, 0/500, required >0
  • Waiting experiment median Spearman, 0.805, compared to 0.800
  • Provider median Spearman range, 0.74‑0.88, compared to required 0.90
  • Gate median at 500 calls, 0.566, compared to gate threshold 0.90
Limitations

The study only measures externally observable behavior on shared serving infrastructure and does not identify the internal cause of instability (e.g., batching vs kernel scheduling vs deployment rotation).

same‑window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte‑identical next‑day replays agreed at 0.78 against a required 0.99Found in the source text, word for word.

Picked because: Shows a reproducible audit of LLM judge reliability and releases scripts/data, giving engineers concrete methods to verify and monitor LLM‑based measurement tools.

Paper 2 of 5

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Jie Wu, Zhenru Zhang, Beichen Zhang and 11 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Realistic, executable terminal environments are scarce despite abundant agent trajectories, and simply generating environments from scratch does not leverage the latent workspace information present in the trajectories.

Approach

Terminal-Universe first replays the file operations recorded in a trajectory to deterministically restore each file before it was modified, producing a partial workspace. A completion agent then fills in missing files and dependencies to obtain a task-sufficient environment. The recovered workspace is used to reconstruct the original intent task and to synthesize new tasks via re‑querying. Breadth expansion mines directional dependency relations to create cross‑workspace queries, while depth expansion adds multi‑round user feedback through a user agent. Verifier filtering selects high‑quality trajectories for supervised fine‑tuning.

Result

Fine‑tuning Qwen3.5-27B on the Full Mixture yields 58.1% Avg. Pass@1 on Terminal‑Bench 2.1, a 11.9‑point gain over the base model, and improves MT@4 on EvoCode‑Bench v2 to 20.1%, a 13.8‑point gain, with a case score of 76.1%. The framework also reconstructs 37.3k task‑sufficient environments.

Why it matters

Researchers building terminal‑based code agents and SFT datasets should adopt Terminal‑Universe to obtain scalable, executable environments that boost both single‑turn and multi‑turn agent performance.

Method details
  • Fine‑tune Qwen3.5-27B for two epochs on the Terminal‑Universe corpus.
  • Full Mixture data contains 32.0k records after verifier filtering.
  • Baseline Qwen3.5-27B scores 41.6% on Terminal‑Bench 2.0 and 46.2% on 2.1 under Terminus2‑XML.
  • Compared against task synthesis methods such as TerminalTraj‑32B (22.0% TB2.0) and Nemotron‑Terminal‑32B (27.4% TB2.0).
  • Expansion‑Axis ablation shows environment expansion yields the largest performance gain under a fixed data budget.
Numbers
  • Avg. Pass@1 on Terminal‑Bench 2.1: 58.1% (up from 46.2% base)
  • MT@4 on EvoCode‑Bench v2: 20.1% (up by 13.8 points)
  • Case Score on EvoCode‑Bench v2: 76.1%
  • Task‑sufficient environments produced: 37.3k
  • Single‑round improvement on Terminal‑Bench 2.1: +11.9 points
  • Multi‑round improvement on EvoCode‑Bench v2 MT@4: +13.8 points
Limitations

The approach uses generic Ubuntu 24.04 containers rather than repository‑specific environments and is limited by the domain and toolchain coverage of the source trajectories, with a single teacher potentially restricting task diversity.

On EvoCode-Bench v2, the Full Mixture raises MT@4 from to and Case score from to , showing that Full Mixture training also improves performance on persistent tasks with cumulative requirements.Found in the source text, word for word.

Picked because: Provides a framework to turn recorded agent trajectories into reusable terminal environments, with released code useful for building and testing self‑hosted LLM agents.

Paper 3 of 5

Hardware-Aware FP4 FlashAttention-4

Robert Hu · abstract · pdf

quote unverified2 figures not in sourceread: full textcs.LG

Problem

Blackwell's 4-bit floating-point (FP4) tensor cores do not automatically make attention faster because softmax conversion and on-chip dependencies dominate once its matrix products shrink, and the obvious fix of simply using FP4 for all tensors leads to divergent training trajectories.

Approach

The paper introduces Direct-P for noncausal inference and a causal path that passes the forward quantization directly into backward. Direct-P maps normalized scores straight to FP4 probabilities, shortening the critical path, and computes the denominator from the same rounded payloads. The causal path reconstructs probabilities from saved quantized queries and keys and uses FP8 gradient operands. Both methods retain HAO's outer schedule and only modify the interval from FP32 score fragment to FP4 probability operand. A guard is added only for layers with extreme logits to preserve numerical stability.

Result

Direct-P achieves up to 2.13× the BF16 forward throughput on an NVIDIA GB200, while the causal path speeds up a full single‑GPU 8‑billion‑parameter update by up to 1.14×. Fixed‑input downstream evaluations show TK NV/MX fast providing speedups of 1.169× to 1.783× across ViT sizes with minimal loss in task score.

Why it matters

Engineers targeting inference acceleration on FP4 tensor cores can gain significant throughput improvements, but must avoid using MXFP4 for training due to instability.

Method details
  • 1.3B and 14B parameter Wan models are evaluated with the proposed methods.
  • Vision Transformer (ViT) and BERT are used for fixed-input downstream evaluations.
  • Datasets include COCO validation images for ViT-MAE, SST-2, and BERT MLM blocks.
  • Noncausal forward benchmark uses D128 shape with batch size, sequence length, and query-head count as per HAO's FP4 FA4 README.
  • Baselines compared against HAO NV/NV, ThunderKittens NV/MX fast and accurate, and BF16 reference.
  • A sampled guard (Algorithm 3) is applied only to layers with extreme logits.
Numbers
  • 2.13× the bfloat16 (BF16) forward throughput on an NVIDIA GB200
  • 1.14× acceleration for a complete single‑GPU 8‑billion‑parameter update
  • TK NV/MX fast speedup 1.169× on ViT S256
  • TK NV/MX fast speedup 1.344× on ViT S1024
  • TK NV/MX fast speedup 1.783× on ViT S4096
  • HAO NV/NV speedup 0.858× on ViT S256
Limitations

The paper does not establish stable training with MXFP4 probabilities/values; every tested MXFP4 probability/value training trajectory diverges.

The model could not quote the source for its claim, so nothing above has been checked against the paper. Read the abstract before trusting it.Citation check failed.
These figures do not appear in the source text: 1.3B, 14B. Treat them as unverified.Number check failed.

Picked because: Introduces FlashAttention‑4 with FP4 support and open‑source kernels, enabling engineers to cut attention latency and memory on commodity hardware.

Paper 4 of 5

Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

Sergii Kozyrev, Davyd Maiboroda · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Previous 4-bit quantizations of Qwen3.8-27B kept the Gated DeltaNet block in 8- or 16-bit precision because it was believed that quantization errors would accumulate over the long recurrent context, so the obvious fix of quantizing the GDN block was thought to fail.

Approach

The authors quantize every linear layer of the hybrid LLM, including all GDN gates, using NVFP4 W4A4 (4‑bit weights and activations). They then study four mechanisms: block‑wise scaling that localizes outliers, gate nonlinearities that compress GEMM error, the delta‑rule recurrence that bounds and quickly erases injected noise, and the fact that per‑token quantization cost washes out over long context. These mechanisms together explain why full‑model 4‑bit quantization does not degrade performance.

Result

Minima matches BF16 within seed noise across all six accuracy suites, uses only 17.5 GiB of VRAM, and achieves the fastest prefill time (4.03 s for a 32K token prompt) while maintaining comparable perplexity, with the gap shrinking at longer context lengths.

Why it matters

Developers of hybrid LLMs can quantize the entire model, including the recurrent GDN half, to 4‑bit precision and gain substantial memory and speed benefits without sacrificing accuracy.

Method details
  • Model: Qwen3.8-27B with 48 GDN layers and 16 softmax‑attention layers (64 total).
  • Quantization: NVFP4 W4A4 applied to all 496 linear layers (240 GDN, 64 attention, 192 MLP projections).
  • Baselines: BF16 reference, Unsloth (Dynamic v3) and RadixArk (ModelOpt) NVFP4 checkpoints that keep GDN and attention at FP8/W8A8.
  • Evaluation suites: WikiText‑2 perplexity (4K and 32K), MMLU‑Pro, GSM8K, AIME’25, GPQA‑Diamond, LiveCodeBench v6, RULER retrieval at 32K and 64K.
  • Serving setup: vLLM 0.27.1, TP=1, RTX PRO 6000 (96 GB), FP8 KV cache, GPU utilization 0.85, 32K generation cap.
  • Ablations: per‑module NVFP4 calibration vs fused‑GEMM scaling, and calibrated FP8 KV‑cache scales.
Numbers
  • PPL@4K: 7.67 (Minima) vs 6.95 (BF16)
  • PPL@32K: 10.84 (Minima) vs 10.35 (BF16)
  • 5‑task avg: 85.10 (Minima) vs 85.62 (BF16)
  • VRAM usage: 17.53 GiB (Minima) vs 50.13 GiB (BF16)
  • Decode throughput: 1,154 tok/s (Minima) vs 621 tok/s (BF16)
  • TTFT @32K: 4.03 s (Minima) vs 6.90 s (BF16)
Limitations

The study is limited to a single model family (Qwen3.8‑27B), a single quantization format (NVFP4), and evaluation up to 32K‑token perplexity and 64K retrieval, without longer‑context stress tests.

All quantized recipes match BF16 task accuracy within seed noise.Found in the source text, word for word.

Picked because: Explains why Gated DeltaNet survives 4‑bit quantization and supplies quantization recipes and code, helping practitioners deploy hybrid LLMs on limited resources.

Paper 5 of 5

Instruction Duplication as an Inference-Time Control Primitive

Victor Lavrenko · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Procedural instruction following was insufficiently observable for downstream controllers, and simply repeating the whole prompt changes content rather than exposing the procedural state.

Approach

The method repeats only the procedural instruction at inference time, creating multiple copies of the same instruction without altering the query or model parameters. A factorial placement design varies the number of copies (zero to three) and their positions (system message, before question, after question). This black‑box control requires no retraining, no decoding changes, and isolates the effect of instruction exposure on trajectory state. The experiment compares one‑copy versus two‑copy conditions across several models. Results are evaluated on deterministic All‑8 completion, TF‑IDF recall, and downstream Answer Engineering performance.

Result

Adding a second copy of the procedural instruction increased deterministic All‑8 completion from 90.22% to 93.17% and TF‑IDF recall from 73.44% to 74.81% while final‑answer accuracy stayed at 60.21%; premature commitment rose from 1.52% to 2.30%. In downstream Answer Engineering, duplication raised the SSNHL endpoint from 25.1% to 97.1% and the conductive branch from 58.9% to 73.8%.

Why it matters

Developers of trajectory‑controlling language‑model systems should consider instruction duplication as a low‑complexity, placement‑sensitive knob that can improve observable protocol state for downstream editors, especially in high‑stakes domains like medical QA.

Method details
  • Seven instruction‑tuned models: Gemma 3 12B, Llama 3.3 70B Instruct, Llama 4 Scout, Ministral 3 14B Instruct, Mistral Large 3, Qwen3 30B‑A3B Instruct, Qwen3 235B‑A22B Instruct
  • 300 medical multiple‑choice questions from MedQA, MedXpertQA, and AfriMed‑QA
  • Inference used temperature 0, cell‑specific seeds, pinned model/provider routes, and model‑specific output ceilings
  • Four‑copy factorial includes zero, three one‑copy, three two‑copy, and one three‑copy conditions (16,800 scheduled generations)
  • Baselines are the one‑copy condition; ablations examine copy count and placement effects
Numbers
  • All‑8 deterministic completion 90.22% → 93.17% (+2.95 pp)
  • Pre‑provisional TF‑IDF recall 73.44% → 74.81% (+1.38 pp)
  • Final‑answer accuracy 60.21% (unchanged)
  • Premature commitment 1.52% → 2.30%
  • SSNHL AE duplication 25.1% → 97.1%
  • Conductive branch AE duplication 58.9% → 73.8%
Limitations

The study is limited to seven contemporary instruction‑tuned models and three medical multiple‑choice benchmarks; results may not transfer to other domains, instruction families, or model populations.

moving from one to two copies raises the deterministic All-8 diagnostic responses passing all eight observable tests from 90.22% to 93.17%Found in the source text, word for word.

Picked because: Presents instruction duplication as a lightweight inference‑time control primitive, with implementation details that can be directly applied to improve LLM tool‑use safety.