arXiv digest

Tuesday

September 1, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering

Gopi Krishnan Rajbahadur, Amir M. Ebrahimi, Boyuan Chen and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Existing post‑training pipelines suffered low usable supervision (only 3,412 valid solutions from 4,064 synthetic problems) and a fixed mixture budget, so simply adding more distilled data failed because new examples must displace existing ones, leading to diminishing returns and data debt.

Approach

The authors apply yield engineering to create a coverage‑first synthetic patch (FDS‑3K‑Cov) that selects high‑yield examples via testcase rectification and constraint injection, then replaces 3.5% of the rehearsal mixture with these examples. The patch is evaluated by running a single fixed checkpoint 16 times per benchmark and reporting bootstrap confidence intervals. This approach directly addresses zero‑sum mixture design and improves the fraction of usable supervision while staying within the fixed compute budget.

Result

Yield engineering raised accepted supervision from 3,412 to 9,697 solutions (2.84× increase) and improved downstream metrics: CodeForces pass@1 +2.59 points, pass@3 +3.11 points; LiveCodeBench v6 pass@1 +6.11 points, pass@3 +8.05 points, all statistically significant across 16 stochastic runs.

Why it matters

Industrial teams that maintain large language models under fixed compute and data budgets should care, as the yield‑engineered patch shows a practical way to improve code‑generation performance without expanding resources.

Method details
  • CodeForces benchmark with 65 tasks and LiveCodeBench v6 with 175 tasks are used for evaluation
  • Baseline (Base) uses a checkpoint trained on a 269,198‑example rehearsal mixture alone
  • Distill‑3K produces 3,412 Tree‑sitter‑valid synthetic solutions without yield engineering
  • FDS‑3K‑Cov is a 3,412‑example coverage‑first subset generated with yield engineering
  • Patch size is 3.5% of the continued‑training mixture (9,697 of 278,895 examples)
  • Each condition is evaluated 16 stochastic generations per task with 95% bootstrap CIs
Numbers
  • CodeForces pass@1 +2.59 points vs Base
  • CodeForces pass@3 +3.11 points vs Base
  • LiveCodeBench v6 pass@1 +6.11 points vs Base
  • LiveCodeBench v6 pass@3 +8.05 points vs Base
  • Accepted supervision increased 2.84× (3,412 to 9,697)
  • Patch replaces 3.5% of mixture (9,697 of 278,895 examples)
Limitations

The study is limited to code‑generation benchmarks (CodeForces and LiveCodeBench v6) and does not demonstrate applicability to other tasks or long‑term maintenance scenarios.

the yield-engineered patch improved CodeForces pass@1 by +2.59 points (+3.11 pass@3) and held-out LiveCodeBench v6 pass@1 by +6.11 (+8.05 pass@3)Found in the source text, word for word.

Picked because: Provides concrete brownfield post‑training workflows, mixture‑patch tooling and released artifacts for maintaining LLMs in production environments.

Paper 2 of 5

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

Milad Rezaei Hajidehi, Qitong Wang, Stratos Idreos · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Agents repeatedly open large documents, consuming up to a million tokens per question, and pre‑structuring everything in advance is infeasible because documents contain far more possible structure than any workload will use and the useful structure is unknown until queries arrive.

Approach

Agentic data cracking adds a cracking sub‑agent that forks from the answer agent when a document is opened, reusing the KV cache at marginal cost. The sub‑agent uses semantic reasoning, instructions, and the data model to speculate about reusable entity sets, attributes, and relations beyond the current query. Extracted structure is normalized, validated, and stored as cracked objects. Future queries first consult the cracked‑object store, avoiding document prefill. The cracking branch runs in parallel and does not delay the answer path.

Result

Cracking reduces mean prefill tokens from 189K to 87K on FanOutQA and from 565K to 161K on the case study, while decode stays below 3K tokens; mean cost per question drops from $0.26 to $0.12 on FanOutQA and from $0.81 to $0.27 on the case study, and LLM‑as‑judge accuracy changes only from 42% to 43% with no statistically significant difference.

Why it matters

Teams building LLM agents over unstructured corpora should care because adaptive structuring can halve token and API cost without degrading answer quality.

Method details
  • Base model: Claude‑Haiku‑4.5.
  • Related‑question generator: Claude‑opus‑5.
  • Datasets: FanOutQA benchmark and a Hitchcock cinema case study.
  • Baseline: same model without the cracking interface, using only RAG.
  • Prompt caching was enabled for all experiments.
  • Acronym Adc denotes the agentic data cracking system.
Numbers
  • prefill tokens, 189K → 87K, FanOutQA
  • prefill tokens, 565K → 161K, case study
  • decode tokens, <3K, all settings
  • mean cost per question, $0.26 → $0.12, FanOutQA
  • mean cost per question, $0.81 → $0.27, case study
  • LLM‑as‑judge accuracy, 42% vs 43%, baseline vs cracking
Limitations

The system assumes a static corpus and only exposes constrained read functions rather than the full SQL space.

cracking cuts mean prefill from 189K to 87K tokens on FanOutQAFound in the source text, word for word.

Picked because: Introduces a token‑efficient agent architecture for reasoning over unstructured data with open‑source code, directly reducing inference cost for LLM‑based services.

Paper 3 of 5

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing

Adrians Skapars, Edoardo Manino · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Automated auditors are highly sample‑inefficient, rarely surfacing rare behaviours because they lack optimisation pressure, and simply increasing compute (e.g., best‑of‑) does not close the gap.

Approach

The method builds on the BLOOM auditing pipeline and adds two rollout extensions: G‑PAIR, which iteratively refines the auditor's input messages across turns, and LogitTilt, which reweights the target model's next‑token distribution by mixing its own logits with those from a behaviour‑eliciting prompt. LogitTilt uses a log‑linear combination controlled by a strength parameter and applies a naturalness floor to avoid low‑probability tokens. The two components can be used independently or together; their combination is called WILT.

Result

WILT outperforms the baseline auditor in 30 of the 32 settings and raises average behaviour presence from 51% to 100% for self‑harm encouragement on Qwen3.5‑4B. LogitTilt alone achieves saturating presence and improves geometric‑mean output probability over vanilla best‑of‑, while preserving naturalness via its floor.

Why it matters

Safety‑focused researchers and practitioners should care because WILT provides a compute‑cheap, logit‑only way to elicit rare harmful behaviours for more reliable model auditing.

Method details
  • Target models: Llama-3.2-3B-Instruct, Phi-4-mini-instruct, Qwen3.5-4B, Gemma-4-E4B.
  • Auditor model: Gemma-4-26B-A4B, a sparse mixture‑of‑experts model quantised to FP8‑Dynamic for GPU efficiency.
  • Baselines: vanilla BLOOM zero‑shot, vanilla BLOOM best‑of‑, BEAST‑in, BEAST‑out, FLRT, TokenBias.
  • Behaviours evaluated: eight, including self‑harm encouragement, racial bias, political bias, etc.
  • Evaluation grid: 4 target models × 8 behaviours = 32 settings.
  • Ablations: G‑PAIR alone, LogitTilt alone, both combined (WILT), and variations of the sampling mixture described in Table 3.
Numbers
  • behaviour presence increase, 51% → 100%, compared to baseline methods
  • settings where WILT beats baseline, 30/32, compared to baseline auditor
Limitations

The paper does not evaluate WILT on unseen models or on longer multi‑turn dialogues beyond the tested scenarios.

WILT raises average behaviour presence from 51% to 100% when eliciting self-harm encouragement from Qwen3.5-4BFound in the source text, word for word.

Picked because: Presents the BLOOM‑WILT logit‑tilting method and an open audit framework for systematically eliciting and testing LLM behaviours in deployment.

Paper 4 of 5

DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening

Yung Wei Shueh, Zhi-Jie Chen, Chia-Hsuan Hsu and 9 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Prior LLM‑based clinical decision support suffered from hallucinated facts, unsupported recommendations, and citation errors; simple fixes such as calibrating probabilities or adding retrieval‑augmented generation do not eliminate these risks because hallucinations and citation drift persist.

Approach

DiaSentinel orchestrates three LLM‑enabled modules, a Risk Function, a Synthesizer, and an entailment check, within a LangGraph pipeline. The Risk Function is a LoRA‑fine‑tuned Qwen2.5‑14B model whose raw logits are calibrated with Platt scaling to produce a one‑year T2DM risk probability. Deterministic components extract clinical signals and retrieve ADA guideline chunks, which are combined by Reciprocal Rank Fusion. A hybrid verification layer first applies rule‑based checks (risk summary, numerical agreement, trend agreement, citation source) and then uses an LLM‑based entailment judge to reject unsupported claims. The calibrated risk scores, retrieved evidence, and verified report are presented via a real‑time batch‑screening dashboard and an interactive patient report interface.

Result

The calibrated risk model achieved an AUROC of 0.737 and a Brier score of 0.054, demonstrating moderate discrimination and good calibration. Deterministic verification checks attained 100% sensitivity and 100% specificity on synthetic clean and error‑injected cases. The LLM‑based guideline‑entailment judge reached 80% sensitivity and 100% specificity on synthetic claim‑evidence pairs.

Why it matters

Hospital clinicians and health‑IT teams should care because DiaSentinel offers an auditable, on‑premise LLM pipeline that mitigates hallucination and calibration issues, enabling safer deployment of AI‑assisted diabetes risk screening.

Method details
  • Risk Function uses a LoRA fine‑tuned Qwen2.5‑14B model.
  • Probability calibration is performed with Platt scaling.
  • Guideline retrieval employs Reciprocal Rank Fusion over ADA guidelines.
  • Hybrid verification combines deterministic rule‑based checks with an LLM entailment judge.
  • Model is trained and evaluated on de‑identified real‑world EHR data from a single hospital with a strict train/validation/test split.
Numbers
  • AUROC, 0.737, model performance
  • Brier score, 0.054, calibration quality
  • Deterministic verification sensitivity, 100%, clean vs injected cases
  • Deterministic verification specificity, 100%, clean vs injected cases
  • Guideline entailment sensitivity, 80%, synthetic claim‑evidence pairs
  • Guideline entailment specificity, 100%, synthetic claim‑evidence pairs
Limitations

The evaluation is limited to synthetic, manually constructed cases and does not reflect real‑world error distributions; fairness and subgroup robustness were not fully evaluated.

The model achieved an AUROC of 0.7370, indicating moderate discriminative ability.Found in the source text, word for word.

Picked because: Describes DIASENTINEL, a fully on‑premise multi‑agent system with released components for reliable, guideline‑grounded health screening pipelines.

Paper 5 of 5

Context-Aware Interleaved Batching for WhisperX

Carlos Bain, Max Bain · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

WhisperX speeds up transcription by batching audio chunks but isolates each chunk, losing historical context needed for coherent punctuation and proper‑noun transcription. The obvious fix of simply feeding previous text as a prefix fails because WhisperX processes batches in parallel, so the preceding chunk’s tokens are not available at inference time.

Approach

The paper introduces Context‑Aware Interleaved Batching (IB) that interleaves VAD‑derived chunk boundaries with text conditioning. VAD segments long audio into natural silences, producing variable‑length chunks up to 30 seconds. Chunks are grouped into batches for parallel inference, but the method injects the decoded text from earlier chunks as a conditioning prefix for later chunks within the same batch. This preserves continuous historical context across chunk boundaries while retaining the throughput benefits of parallel processing. The approach relies on Whisper’s existing encoder‑decoder Transformer and its condition_on_previous_text mechanism, activated only when the preceding text is available.

Result

On the Earnings‑21 benchmark, WhisperX+IB achieves a WER of 8.2% and a Proper‑Noun Score of 79.3%, improving over standard WhisperX (8.4% WER, 76.5% Pn) while matching the LLM Punctuation Score of 2.9 and delivering an 8.4× speedup. On the MedicalLessons corpus, WhisperX+IB reduces WER from 3.5% to 3.3% and raises the Proper‑Noun Score from 83.6% to 84.7% compared with WhisperX.

Why it matters

Researchers and engineers building high‑throughput speech transcription systems should care because the method restores context‑aware accuracy without sacrificing most of the speed gains of parallel batching.

Method details
  • Uses the standard Whisper encoder‑decoder Transformer architecture
  • Relies on a lightweight VAD model [9] to create chunk boundaries
  • Processes batches of size 4 during inference
  • Evaluated on Earnings‑21 and MedicalLessons datasets
  • Compared against baseline WhisperX and openai/whisper models
  • Ablation includes enabling or disabling condition_on_previous_text
Numbers
  • WER 8.2 vs WhisperX 8.4
  • Pn Score 79.3 vs WhisperX 76.5
  • Speed 8.4× vs WhisperX 11.8×
  • WER 3.3 vs WhisperX 3.5
  • Pn Score 84.7 vs WhisperX 83.6
Limitations

When the number of VAD‑segmented chunks is less than or equal to the batch size, context cannot be propagated without reducing the batch size, which trades throughput for accuracy.

By successfully propagating preceding text context to subsequent batches, WhisperX+IB reduces the overall WER from 3.5% to 3.3%.Found in the source text, word for word.

Picked because: Offers a context‑aware interleaved batching technique for WhisperX that can be integrated into speech‑processing pipelines to improve throughput without sacrificing context.