arXiv digest

Wednesday

August 19, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents

Yining Hua, Hongbin Na, Yifan Zhou and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Agents could not reliably locate and edit evidence because they either only accessed native PDFs without searchable indexing or only used a parsed cache that did not reflect the full document, and simple fixes like using the parsed cache as a read surface or grepping raw PDFs failed.

Approach

StagedWorkspace binds each native file to a hash-keyed parsed cache and tracks file‑edit diffs, providing synchronized native and parsed views. SW‑Agent accesses both views (dual) and can see tracked diffs before submission. The system uses a ReAct‑style interleaved reasoning loop with a fixed 250‑tool‑call budget, keeping parser, retriever, grader, and prompts constant across ablations. Workspace updates trigger re‑parsing and re‑indexing so the parsed view stays aligned with the sandbox files.

Result

Dual synchronized views gave the highest performance on both benchmarks: OfficeQA Pass@1 improved by 8.3 to 12.1 points over artifact‑only and APEX mean rubric score improved by 4.7 to 9.2 points over parsed‑only. SW‑Agent achieved 63.9% Pass@1 with Gemini 3.1 Pro on OfficeQA and 42.1% mean rubric with GPT‑5.4 Nano on APEX, far above the published same‑model scores of 29.3% and 25.5% respectively.

Why it matters

Researchers and engineers building knowledge‑work agents should adopt a versioned workspace with synchronized native and parsed views, as it yields large, reproducible gains in retrieval and execution performance.

Method details
  • Benchmarks: OfficeQA Pro (U.S. Treasury Bulletin PDFs 1939 to 2025) and APEX‑Agents (33 worlds, 480 rubric‑graded tasks).
  • Models evaluated include Gemini 3.1 Pro, Gemini 3 Flash, GPT‑5.4 Nano, GPT‑5.4 Mini, and GPT‑5.4 (xHigh).
  • Baseline arms: artifact‑only (native files only), parsed‑only (parsed index only), and published public rows for the same models.
  • Ablations: read‑axis (dual vs artifact‑only vs parsed‑only) and review‑axis (diffs hidden vs diffs visible) on the same SW‑Agent harness.
  • Inference setup: ReAct‑style reasoning, up to 250 tool calls, fixed prompts, parser, retriever, grader, file tracker, and tool budget.
Numbers
  • OfficeQA Pass@1 63.9% vs published 29.3% (Gemini 3.1 Pro)
  • OfficeQA Pass@1 64.7 vs 56.4 (+8.3 points) for GPT‑5.4 dual vs best published Full row
  • APEX mean rubric 42.1 vs 25.5 (+16.6 points) for GPT‑5.4 Nano dual vs published
  • Lift dual vs artifact‑only OfficeQA GPT‑5.4 +12.1 (7.6)
  • Lift dual vs parsed‑only APEX GPT‑5.4 Nano +9.2 (5.9)
  • Review‑axis lift GPT‑5.4 Mini +8.5 points when diffs visible vs hidden
Limitations

The study excludes tasks that depend on external APIs (e.g., EDGAR) and evaluates only the two public benchmarks, so broader generalization is not demonstrated.

dual views means SW-Agent exposes both native files and parsed search over that full corpus.Found in the source text, word for word.

Picked because: Introduces StagedWorkspace, a released versioned workspace system enabling reliable state contracts for AI agents that manipulate code and documents, directly useful for building self‑hosted LLM‑agent tooling.

Paper 2 of 5

The Polyglot's Dilemma: Conformance Testing a Dozen Specs in as Many Languages

A. Jesse Jiryu Davis, Jeremy Mikola, Jeff Yemin · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Before this work each driver used its own ad‑hoc test format, leading to duplicated test code and inconsistent behavior across the dozen libraries; simply sharing code was not feasible because drivers are implemented natively in different languages.

Approach

The authors created a domain‑specific test language in YAML and a Unified Test Format (UTF) validated by a JSON Schema. Each driver implements a language‑specific UTF test runner that reads the YAML, creates entity maps, executes operations, and uses command monitoring and fail points to observe wire‑level behavior. The test runner synchronously runs one command at a time to ensure determinism. Tests are written once and run across all drivers, allowing centralized updates when specifications change.

Result

Adopting the Unified Test Format allowed the team to delete more than 22,000 lines of test code, grow the test corpus to 606 files with 124,168 lines, and achieve up to an 86% reduction in non‑conformance bugs for drivers that switched to the YAML‑based tests.

Why it matters

Teams that maintain multiple language implementations of the same API should adopt a declarative, cross‑language test format to reduce duplication and improve consistency.

Method details
  • Tests are authored in YAML and validated against a JSON Schema (UTF).
  • Drivers are written in nine native languages (C, C#, Go, Java, Node.js, Python, Ruby, Rust, Swift) plus wrappers in C++, PHP, Scala, Kotlin.
  • The test corpus grew from 7 files (1,046 lines) to 606 files (124,168 lines) by 2026.
  • Over 22,000 lines of redundant test‑runner code were deleted after migrating to UTF.
  • Non‑conformance bug rates fell up to 86% in drivers that adopted the YAML tests.
Numbers
  • over 22,000 lines of test code deleted
  • up to 86% reduction in nonconformance bugs
  • 7 files and 1,046 lines of YAML initially
  • 606 files with 124,168 lines later
  • 6,038 LOC saved in the Java driver
  • eleven years of industrial experience
Limitations

The paper notes limits of unification but does not establish how the approach scales beyond MongoDB drivers or to other domains.

The rate of nonconformance bugs fell up to 86% in drivers that adopted YAML testsFound in the source text, word for word.

Picked because: Presents a specification‑based conformance testing framework for client libraries in a dozen languages, with open‑source artifacts that help engineers ensure cross‑language consistency in platform software.

Paper 3 of 5

Minimizing Commit Rules for DAG-based Atomic Broadcast

Petr Kuznetsov, Maxence Perion, Sara Tucci-Piergiovanni · abstract · pdf

quote verifiedfigures checkedread: full textcs.DC

Problem

Existing DAG-based atomic broadcast protocols use commit rules that are not provably minimal, leading to unnecessary latency, and simply weakening those rules would break safety guarantees.

Approach

The paper defines commit rules on an uncertified round‑based DAG and introduces a sub‑rule relation to compare them. It identifies the weakest sufficient commit rule for both eventually synchronous and asynchronous models. The minimal rule requires only f edges to the same vertex in the previous round and that concurrent leader vertices are already decided. These minimal rules are instantiated in two protocols called S‑Minnow and A‑Minnow. The approach relies on a deterministic leaders function and a local DAG that records causal edges.

Result

The paper provides theoretical proofs that the identified commit rules are minimal for both models; no empirical measurements are reported.

Why it matters

Researchers designing Byzantine atomic broadcast protocols should consider the minimal commit rules to reduce latency and increase throughput.

Method details
  • Uses an uncertified round‑based DAG construction where each vertex references at least f+1 vertices from the previous round
  • Minimal commit rule for eventual synchrony requires f edges pointing to the same vertex in the previous round
  • Minimal commit rule for asynchrony also requires f edges and decided concurrent leaders
  • Instantiates the minimal rules in protocols S‑Minnow (eventual synchrony) and A‑Minnow (asynchrony)
  • Assumes a leaders function that returns a consistent sequence of leader slots across correct processes
Numbers
  • edges required,f,minimal commit rule for eventual synchrony
  • edges required,f,minimal commit rule for asynchrony
  • processes,n,total number of processes in the system
  • Byzantine processes,f,fault tolerance bound
Limitations

The work is purely theoretical and does not include implementation or experimental evaluation of Minnow.

we introduce Minnow, a new protocol for DAG-based atomic broadcast, which can be instantiated in both eventually synchronous (S-Minnow) and asynchronous networks (A-Minnow).Found in the source text, word for word.

Picked because: Proposes a minimal commit‑rule design for DAG‑based atomic broadcast, offering concrete protocol optimizations that can be adopted in distributed infrastructure and cloud services.

Paper 4 of 5

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

Bin Li, Dongdong Wang, Siyang Lu · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Existing language model based log anomaly detectors assign excessive confidence to incorrect predictions, especially under severe class imbalance. Conventional calibration metrics appear good but do not reduce the high confidence on erroneous predictions, so simple post‑hoc scaling fails to fix the reliability gap.

Approach

LoRD learns a separate autoencoder for each prediction route (normal and abnormal) using latent representations of correctly classified validation samples. Reconstruction distance from the route‑specific autoencoder serves as a reliability indicator. A calibration policy maps this distance to adjust the original confidence, selectively recalibrating high‑risk predictions while leaving reliable ones unchanged. The method operates as a lightweight post‑hoc step that does not require retraining the base detector. Route‑specific modeling, a reject region, and distance‑aware soft calibration together form the complete framework.

Result

LoRD consistently yields the lowest error confidence (CoE) across datasets and detectors, reducing CoE from values near one to around 0.5 while preserving high confidence on correctly classified anomalous samples (CoC). Conventional metrics such as ECE and Brier remain extremely small, indicating that LoRD improves reliability without harming detection performance.

Why it matters

Operators of large‑scale computing systems can adopt LoRD to obtain more trustworthy confidence scores for anomaly alerts, reducing false confidence without sacrificing detection accuracy.

Method details
  • Evaluated on four large‑scale log datasets: BGL, Spirit, Liberty, Thunderbird with anomaly ratios 0.49% to 32.01%.
  • Base detectors include TextCNN, LogRobust, LightLog, NeuralLog, and GPT2.
  • LoRD trains route‑specific autoencoders on latent representations of correctly classified samples for each route.
  • Compared against five post‑hoc baselines: Temperature Scaling, Logistic Scaling, Beta Scaling, Selective Scaling, and Ensembling.
  • Ablation studies remove the reject region and the soft calibration component to assess their impact.
  • Model complexity experiment varies MLP parameters; accuracy and F1 remain >0.999 while CoE rises from 0.90 to nearly 0.99.
Numbers
  • CoE 0.509 for LoRD on Spirit with LogRobust (vs 0.504 w/o Reject)
  • CoE 0.405 for LoRD on Liberty with LogRobust (vs 0.415 w/o Soft)
  • CoE 0.513 for LoRD on Spirit with NeuralLog (vs 0.507 w/o Reject)
  • CoE 0.680 for LoRD on Liberty with NeuralLog (vs 0.572 w/o Reject)
  • Accuracy >0.999 and F1 >0.999 across all model sizes
  • CoE increases from 0.90 to nearly 0.99 as model size grows
Limitations

The paper does not evaluate LoRD on unseen log domains or quantify its runtime overhead.

LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.Found in the source text, word for word.

Picked because: Shows how to calibrate language‑model‑based log anomaly detectors, providing practical methods and code to improve reliability of DevOps monitoring pipelines.

Paper 5 of 5

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

Sher Badshah, Ali Emami, Hassan Sajjad · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Reference-free LLM judges for factual tasks can hallucinate or lack evidence, and neither pure parametric evaluation nor retrieval augmentation alone provides formal control over the risk of accepted verdicts.

Approach

The method calibrates uncertainty thresholds on a held‑out set using finite‑sample Clopper-Pearson intervals to bound the false discovery rate (FDR). It employs a two‑mode routing: a parametric mode with a calibrated threshold, and if the judge is not confident, it routes the instance to a retrieval‑augmented mode with a second calibrated threshold. The joint two‑threshold policy inherits the same finite‑sample guarantee without extra assumptions. Coverage is increased by allowing retrieval to rescue instances that would otherwise be abstained.

Result

Across all 32 configurations the observed FDR stays at or below the specified target, confirming the guarantee, and the adaptive retrieval mode yields substantially higher coverage than single‑mode baselines.

Why it matters

Practitioners who need automated, risk‑controlled factual evaluation of LLM outputs should adopt this framework to obtain formal error bounds while retaining higher coverage.

Method details
  • Datasets: TriviaQA, Natural Questions, HotpotQA, and PopQA (four open‑domain QA benchmarks).
  • Candidate models: Qwen3‑8B and LLaMA‑3.1‑70B.
  • Judge models: Qwen3‑4B, Qwen3‑8B, Qwen3‑14B, and LLaMA‑3.1‑8B‑Instruct.
  • Uncertainty measure: predictive entropy computed from token‑level log‑probabilities at the verdict token.
  • Retrieval: top‑k web results concatenated to the judge prompt (k varied in ablations).
  • Calibration: 100 random 50/50 splits with grid search over unique uncertainty values to satisfy the Clopper-Pearson UCB constraint.
Numbers
  • sample size, 2,000 instances, per benchmark
  • calibration splits, 100 random 50/50 splits, used for threshold selection
  • datasets, 4 open‑domain QA benchmarks, evaluated
  • candidate models, 2 models, Qwen3‑8B and LLaMA‑3.1‑70B
  • judge models, 4 models, Qwen3‑4B, Qwen3‑8B, Qwen3‑14B, LLaMA‑3.1‑8B‑Instruct
  • evaluation runs, mean ± std over 100 splits, reported for FDR
Limitations

The paper does not establish guarantees for graded or subjective evaluation settings and does not cover black‑box judges without access to logits.

the framework maintains the target error rate while achieving substantially higher coverage than single-mode baselines.Found in the source text, word for word.

Picked because: Offers a provably risk‑bounded LLM judging approach with released evaluation tools, giving engineers a verifiable way to automate model output assessment.