Yining Hua, Hongbin Na, Yifan Zhou and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Agents could not reliably locate and edit evidence because they either only accessed native PDFs without searchable indexing or only used a parsed cache that did not reflect the full document, and simple fixes like using the parsed cache as a read surface or grepping raw PDFs failed.
Approach
StagedWorkspace binds each native file to a hash-keyed parsed cache and tracks file‑edit diffs, providing synchronized native and parsed views. SW‑Agent accesses both views (dual) and can see tracked diffs before submission. The system uses a ReAct‑style interleaved reasoning loop with a fixed 250‑tool‑call budget, keeping parser, retriever, grader, and prompts constant across ablations. Workspace updates trigger re‑parsing and re‑indexing so the parsed view stays aligned with the sandbox files.
Result
Dual synchronized views gave the highest performance on both benchmarks: OfficeQA Pass@1 improved by 8.3 to 12.1 points over artifact‑only and APEX mean rubric score improved by 4.7 to 9.2 points over parsed‑only. SW‑Agent achieved 63.9% Pass@1 with Gemini 3.1 Pro on OfficeQA and 42.1% mean rubric with GPT‑5.4 Nano on APEX, far above the published same‑model scores of 29.3% and 25.5% respectively.
Why it matters
Researchers and engineers building knowledge‑work agents should adopt a versioned workspace with synchronized native and parsed views, as it yields large, reproducible gains in retrieval and execution performance.
Method details
Benchmarks: OfficeQA Pro (U.S. Treasury Bulletin PDFs 1939 to 2025) and APEX‑Agents (33 worlds, 480 rubric‑graded tasks).
Models evaluated include Gemini 3.1 Pro, Gemini 3 Flash, GPT‑5.4 Nano, GPT‑5.4 Mini, and GPT‑5.4 (xHigh).
Baseline arms: artifact‑only (native files only), parsed‑only (parsed index only), and published public rows for the same models.
Ablations: read‑axis (dual vs artifact‑only vs parsed‑only) and review‑axis (diffs hidden vs diffs visible) on the same SW‑Agent harness.
Inference setup: ReAct‑style reasoning, up to 250 tool calls, fixed prompts, parser, retriever, grader, file tracker, and tool budget.
Numbers
OfficeQA Pass@1 63.9% vs published 29.3% (Gemini 3.1 Pro)
OfficeQA Pass@1 64.7 vs 56.4 (+8.3 points) for GPT‑5.4 dual vs best published Full row
APEX mean rubric 42.1 vs 25.5 (+16.6 points) for GPT‑5.4 Nano dual vs published
Lift dual vs artifact‑only OfficeQA GPT‑5.4 +12.1 (7.6)
Lift dual vs parsed‑only APEX GPT‑5.4 Nano +9.2 (5.9)
Review‑axis lift GPT‑5.4 Mini +8.5 points when diffs visible vs hidden
Limitations
The study excludes tasks that depend on external APIs (e.g., EDGAR) and evaluates only the two public benchmarks, so broader generalization is not demonstrated.
dual views means SW-Agent exposes both native files and parsed search over that full corpus.Found in the source text, word for word.
Picked because: Introduces StagedWorkspace, a released versioned workspace system enabling reliable state contracts for AI agents that manipulate code and documents, directly useful for building self‑hosted LLM‑agent tooling.
A. Jesse Jiryu Davis, Jeremy Mikola, Jeff Yemin · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
Before this work each driver used its own ad‑hoc test format, leading to duplicated test code and inconsistent behavior across the dozen libraries; simply sharing code was not feasible because drivers are implemented natively in different languages.
Approach
The authors created a domain‑specific test language in YAML and a Unified Test Format (UTF) validated by a JSON Schema. Each driver implements a language‑specific UTF test runner that reads the YAML, creates entity maps, executes operations, and uses command monitoring and fail points to observe wire‑level behavior. The test runner synchronously runs one command at a time to ensure determinism. Tests are written once and run across all drivers, allowing centralized updates when specifications change.
Result
Adopting the Unified Test Format allowed the team to delete more than 22,000 lines of test code, grow the test corpus to 606 files with 124,168 lines, and achieve up to an 86% reduction in non‑conformance bugs for drivers that switched to the YAML‑based tests.
Why it matters
Teams that maintain multiple language implementations of the same API should adopt a declarative, cross‑language test format to reduce duplication and improve consistency.
Method details
Tests are authored in YAML and validated against a JSON Schema (UTF).
Drivers are written in nine native languages (C, C#, Go, Java, Node.js, Python, Ruby, Rust, Swift) plus wrappers in C++, PHP, Scala, Kotlin.
The test corpus grew from 7 files (1,046 lines) to 606 files (124,168 lines) by 2026.
Over 22,000 lines of redundant test‑runner code were deleted after migrating to UTF.
Non‑conformance bug rates fell up to 86% in drivers that adopted the YAML tests.
Numbers
over 22,000 lines of test code deleted
up to 86% reduction in nonconformance bugs
7 files and 1,046 lines of YAML initially
606 files with 124,168 lines later
6,038 LOC saved in the Java driver
eleven years of industrial experience
Limitations
The paper notes limits of unification but does not establish how the approach scales beyond MongoDB drivers or to other domains.
The rate of nonconformance bugs fell up to 86% in drivers that adopted YAML testsFound in the source text, word for word.
Picked because: Presents a specification‑based conformance testing framework for client libraries in a dozen languages, with open‑source artifacts that help engineers ensure cross‑language consistency in platform software.
Petr Kuznetsov, Maxence Perion, Sara Tucci-Piergiovanni · abstract · pdf
quote verifiedfigures checkedread: full textcs.DC
Problem
Existing DAG-based atomic broadcast protocols use commit rules that are not provably minimal, leading to unnecessary latency, and simply weakening those rules would break safety guarantees.
Approach
The paper defines commit rules on an uncertified round‑based DAG and introduces a sub‑rule relation to compare them. It identifies the weakest sufficient commit rule for both eventually synchronous and asynchronous models. The minimal rule requires only f edges to the same vertex in the previous round and that concurrent leader vertices are already decided. These minimal rules are instantiated in two protocols called S‑Minnow and A‑Minnow. The approach relies on a deterministic leaders function and a local DAG that records causal edges.
Result
The paper provides theoretical proofs that the identified commit rules are minimal for both models; no empirical measurements are reported.
Why it matters
Researchers designing Byzantine atomic broadcast protocols should consider the minimal commit rules to reduce latency and increase throughput.
Method details
Uses an uncertified round‑based DAG construction where each vertex references at least f+1 vertices from the previous round
Minimal commit rule for eventual synchrony requires f edges pointing to the same vertex in the previous round
Minimal commit rule for asynchrony also requires f edges and decided concurrent leaders
Instantiates the minimal rules in protocols S‑Minnow (eventual synchrony) and A‑Minnow (asynchrony)
Assumes a leaders function that returns a consistent sequence of leader slots across correct processes
Numbers
edges required,f,minimal commit rule for eventual synchrony
edges required,f,minimal commit rule for asynchrony
processes,n,total number of processes in the system
Byzantine processes,f,fault tolerance bound
Limitations
The work is purely theoretical and does not include implementation or experimental evaluation of Minnow.
we introduce Minnow, a new protocol for DAG-based atomic broadcast, which can be instantiated in both eventually synchronous (S-Minnow) and asynchronous networks (A-Minnow).Found in the source text, word for word.
Picked because: Proposes a minimal commit‑rule design for DAG‑based atomic broadcast, offering concrete protocol optimizations that can be adopted in distributed infrastructure and cloud services.
Existing language model based log anomaly detectors assign excessive confidence to incorrect predictions, especially under severe class imbalance. Conventional calibration metrics appear good but do not reduce the high confidence on erroneous predictions, so simple post‑hoc scaling fails to fix the reliability gap.
Approach
LoRD learns a separate autoencoder for each prediction route (normal and abnormal) using latent representations of correctly classified validation samples. Reconstruction distance from the route‑specific autoencoder serves as a reliability indicator. A calibration policy maps this distance to adjust the original confidence, selectively recalibrating high‑risk predictions while leaving reliable ones unchanged. The method operates as a lightweight post‑hoc step that does not require retraining the base detector. Route‑specific modeling, a reject region, and distance‑aware soft calibration together form the complete framework.
Result
LoRD consistently yields the lowest error confidence (CoE) across datasets and detectors, reducing CoE from values near one to around 0.5 while preserving high confidence on correctly classified anomalous samples (CoC). Conventional metrics such as ECE and Brier remain extremely small, indicating that LoRD improves reliability without harming detection performance.
Why it matters
Operators of large‑scale computing systems can adopt LoRD to obtain more trustworthy confidence scores for anomaly alerts, reducing false confidence without sacrificing detection accuracy.
Method details
Evaluated on four large‑scale log datasets: BGL, Spirit, Liberty, Thunderbird with anomaly ratios 0.49% to 32.01%.
Base detectors include TextCNN, LogRobust, LightLog, NeuralLog, and GPT2.
LoRD trains route‑specific autoencoders on latent representations of correctly classified samples for each route.
Compared against five post‑hoc baselines: Temperature Scaling, Logistic Scaling, Beta Scaling, Selective Scaling, and Ensembling.
Ablation studies remove the reject region and the soft calibration component to assess their impact.
Model complexity experiment varies MLP parameters; accuracy and F1 remain >0.999 while CoE rises from 0.90 to nearly 0.99.
Numbers
CoE 0.509 for LoRD on Spirit with LogRobust (vs 0.504 w/o Reject)
CoE 0.405 for LoRD on Liberty with LogRobust (vs 0.415 w/o Soft)
CoE 0.513 for LoRD on Spirit with NeuralLog (vs 0.507 w/o Reject)
CoE 0.680 for LoRD on Liberty with NeuralLog (vs 0.572 w/o Reject)
Accuracy >0.999 and F1 >0.999 across all model sizes
CoE increases from 0.90 to nearly 0.99 as model size grows
Limitations
The paper does not evaluate LoRD on unseen log domains or quantify its runtime overhead.
LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.Found in the source text, word for word.
Picked because: Shows how to calibrate language‑model‑based log anomaly detectors, providing practical methods and code to improve reliability of DevOps monitoring pipelines.
Sher Badshah, Ali Emami, Hassan Sajjad · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Reference-free LLM judges for factual tasks can hallucinate or lack evidence, and neither pure parametric evaluation nor retrieval augmentation alone provides formal control over the risk of accepted verdicts.
Approach
The method calibrates uncertainty thresholds on a held‑out set using finite‑sample Clopper-Pearson intervals to bound the false discovery rate (FDR). It employs a two‑mode routing: a parametric mode with a calibrated threshold, and if the judge is not confident, it routes the instance to a retrieval‑augmented mode with a second calibrated threshold. The joint two‑threshold policy inherits the same finite‑sample guarantee without extra assumptions. Coverage is increased by allowing retrieval to rescue instances that would otherwise be abstained.
Result
Across all 32 configurations the observed FDR stays at or below the specified target, confirming the guarantee, and the adaptive retrieval mode yields substantially higher coverage than single‑mode baselines.
Why it matters
Practitioners who need automated, risk‑controlled factual evaluation of LLM outputs should adopt this framework to obtain formal error bounds while retaining higher coverage.
evaluation runs, mean ± std over 100 splits, reported for FDR
Limitations
The paper does not establish guarantees for graded or subjective evaluation settings and does not cover black‑box judges without access to logits.
the framework maintains the target error rate while achieving substantially higher coverage than single-mode baselines.Found in the source text, word for word.
Picked because: Offers a provably risk‑bounded LLM judging approach with released evaluation tools, giving engineers a verifiable way to automate model output assessment.