arXiv digest

Sunday

August 30, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Yufan Wu, Yinghui He, Zhengyi Hu and 4 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Prior inference-time scaling methods improve reasoning but rely on repeated generation or external verification, which is costly; simply generating more samples does not address the inefficiency.

Approach

CritICL leverages structured failure modes observed in weaker models of the same family as guidance for stronger models. It builds a CritBank of failure mode labels and critiques from small-scale models. Two variants are introduced: CritICL-dynamic predicts input‑specific failure modes and retrieves relevant critiques, while CritICL-static uses an aggregated global failure mode profile to retrieve stable examples. During inference the selected critique examples are provided as in‑context demonstrations to the target model. This transfers weak‑model failure information to improve strong‑model reasoning without extra generations.

Result

CritICL-static achieves 93.6% on GSM8K and 59.2% on MATH, raising the overall in‑distribution average to 76.4% and the out‑of‑distribution average to 21.3%, which surpasses standard ICL and matches or exceeds test‑time scaling methods while using far fewer generations.

Why it matters

Researchers seeking efficient inference‑time reasoning improvements should consider CritICL, as it provides strong accuracy gains without the token overhead of repeated sampling.

Method details
  • Evaluated on GSM8K (7.4k train, 1.3k test) and MATH (7.5k train, 5k test) plus AMC23, AIME24, AIME25 OOD benchmarks
  • CritBank constructed from responses of Qwen2.5‑1.5B‑Instruct, Qwen2.5‑3B‑Instruct, Qwen2.5‑7B‑Instruct and Llama‑3.2‑1B‑Instruct, Llama‑3.2‑3B‑Instruct, Llama‑3.1‑8B
  • Target models: Qwen2.5‑32B‑Instruct, Qwen2.5‑72B‑Instruct, Llama‑3.1‑70B‑Instruct, all decoded greedily with temperature 0.0
  • Baselines include zero‑shot, 1/3/5‑shot random and fixed exemplars, Self‑Consistency@3/5/7, Self‑Reflection, LLM‑as‑Judge (GPT‑4o‑mini)
  • Ablation compares failure‑mode‑based example selection against random, fixed, and semantic similarity retrieval (Table 3)
Numbers
  • GSM8K accuracy 93.6 vs 93.0 Consistency@5
  • MATH accuracy 59.2 vs 58.6 Consistency@5
  • Overall ID average 76.4 vs 75.8 CritICL‑dynamic
  • Overall OOD average 21.3 vs 21.1 CritICL‑dynamic
  • Overall average 49.8 vs 49.1 CritICL‑dynamic
  • Zero‑shot overall 36.6 vs CritICL‑static 49.8
Limitations

The paper does not establish a causal link between shared inductive biases and the transferability of failure modes, offering only a hypothesis.

CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methodsFound in the source text, word for word.

Picked because: Introduces CritICL, an inference-time framework that boosts LLM reasoning efficiency without costly repeated generation, offering a practical tool for self‑hosted LLM services.

Paper 2 of 5

Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners

Qianlong Lan, Vinothini Pandurangan, Anuj Kaul and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Prior evaluations of model security scanners only reported conditional precision/recall for cases where a scanner produced a usable judgment, ignoring how often scanners actually returned a decision; the obvious fix of reporting overall accuracy still hides the large portion of families with no definitive judgment.

Approach

The paper introduces a controlled benchmark that separates non‑N/A coverage, analysis completion, definitive security decisions, non‑security findings, and unsupported outcomes. It evaluates three static scanners-ModelScan, ModelAudit, and Fickling-on a synthetic corpus of Pickle and PyTorch artifacts. The methodology records family‑level decision coverage, conditional detection metrics, and cross‑scanner recovery for incomplete analyses. Latency and rename robustness are also measured. By reporting both judgment availability and conditional accuracy, the approach reveals coverage gaps and redundancy among tools.

Result

ModelAudit produced definitive security decisions for all 135 labeled families (100% coverage), Fickling for 110 families (81.5%), and ModelScan for 67 families (49.6%). Conditional on making a definitive judgment, ModelScan achieved 100% precision, recall, and F1. Fickling added no unique true‑positive families beyond the union of ModelAudit and ModelScan, but both ModelAudit and Fickling recovered detections for all 48 malicious families where ModelScan failed to complete analysis.

Why it matters

Security engineers and tool developers should care because the study shows that high conditional accuracy can coexist with low decision coverage, highlighting the need for multi‑scanner ensembles and coverage‑aware evaluation when protecting model supply chains.

Method details
  • Synthetic corpus of 170 artifacts organized into 145 specimen families (135 labeled families, 10 malformed).
  • Family composition: 70 malicious families, 65 benign families, 10 malformed families.
  • Scanner versions: ModelScan 0.8.6, ModelAudit 0.2.52, Fickling 0.1.12.
  • Scanner timeouts: ModelScan 60 s, ModelAudit 10 s, Fickling 5 s; two workers used.
  • Total executions: 510 scanner/artifact runs; latency measured over 170 executions per scanner.
  • Rename robustness experiment: 15 scanner‑level comparisons across five malicious Pickle families.
Numbers
  • ModelAudit definitive decision coverage 100% (135/135)
  • Fickling definitive decision coverage 81.5% (110/135)
  • ModelScan definitive decision coverage 49.6% (67/135)
  • ModelScan conditional precision 100%
  • ModelScan conditional recall 100%
  • ModelScan conditional F1 100%
  • ModelAudit benign‑family false‑positive rate 92.3%
  • ModelScan incomplete families 78 (48 malicious, 20 benign, 10 malformed)
Limitations

The benchmark is limited to synthetic Pickle and PyTorch artifacts and does not represent real‑world model‑scanner performance or adversarial evasion scenarios.

Conditional on making a definitive judgment, ModelScan achieved 100% precision, recall, and F1.Found in the source text, word for word.

Picked because: Provides concrete evaluation of AI model security scanners (ModelScan, ModelAudit, Fickling) with released benchmarks, directly useful for DevOps security pipelines.

Paper 3 of 5

Token-Level Advertising

Hanbing Liu, Bowei Zhang, Changyuan Yu and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.GT

Problem

Traditional advertising relies on predefined slots and cannot embed advertiser influence into the token generation process; simply allocating after generation does not capture token‑level participation and fails to integrate ads seamlessly.

Approach

LAMA lets advertisers submit local continuation value reports at each non‑terminal prefix; the platform verifies Bellman consistency, maintains an allocation posterior over advertisers, and samples an advertiser whose optimal next‑token policy generates the next token; after each token the belief is updated via Bayes and at the end a winning advertiser is sampled and charged; the mechanism satisfies Markov DSIC and IR while approximating KL‑regularized welfare.

Result

Across the three verticals LAMA attains the highest platform welfare (0.5205), revenue (0.8305), advertiser value (0.8568) and user quality (66.5239), outperforming the best allocate‑after policy baseline while preserving response quality.

Why it matters

Ad platforms and generative‑AI services should care because LAMA shows token‑level advertising can increase monetization without harming user experience, and mechanism‑design researchers gain a new sequential auction framework.

Method details
  • Reference language model is Qwen3‑14B
  • Report models are advertiser‑conditioned LoRA heads
  • Dataset is Webis Generated Native Ads 2024 with three verticals: Workout, Vacation, Car
  • Baselines include six heuristic IC combos plus MOSAIC
  • Training decomposes reports into local soft advantages and root values using supervised signals
  • Inference samples advertiser from posterior and updates ledger each token
Numbers
  • Welfare 0.5205 vs allocate‑after policy 0.5080
  • Revenue 0.8305 vs allocate‑after policy 0.7501
  • Advertiser Value 0.8568 vs allocate‑after policy 0.8253
  • Quality 66.5239 vs allocate‑after policy 65.5939
  • Welfare 0.5205 vs MOSAIC 0.4390
  • Revenue 0.8305 vs MOSAIC 0.5274
Limitations

The work is a proof‑of‑concept limited to a single‑winner setting and does not demonstrate large‑scale deployment or multi‑advertiser competition.

LAMA achieves the strongest overall performance among the compared methods, attaining the highest mean platform welfareFound in the source text, word for word.

Picked because: Presents LAMA, a token‑level advertising mechanism that embeds external influence into generation, demonstrating a deployable pattern for custom LLM tooling.

Paper 4 of 5

INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment

Yutong Zhang, Jianshuo Dong, Peng Xu and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

CoT monitoring only provides post‑hoc, coarse‑grained labels and relies on an external judge, so it cannot show when harmful intent emerges during generation and adding frequent judge calls would add prohibitive cost and latency.

Approach

INTENT‑AS‑A‑TOOL inserts a zero‑parameter intent tool into the model’s tool set and uses the probability of calling this tool on truncated prefixes as a fine‑grained, judge‑free signal of harmful intent. The model’s next‑tool distribution is scored after appending a tool‑call opener, producing a trajectory of intent probabilities. When the probability exceeds a threshold, an online reflection is inserted before decoding continues, allowing intervention at the point of intent emergence. This mechanism replaces costly external judge calls with lightweight prefix scoring and enables dynamic, intent‑guided defenses.

Result

Across the Qwen family, intent‑guided online intervention achieves higher case‑level defense success rates than the prompting baseline in most model‑scenario pairs, e.g., Qwen3‑32B blackmail 100.0% vs 80.0% and Qwen3‑32B murder 63.4% vs 98.9% for prompting. Timing ablation shows intent‑guided triggering yields the highest success (97.8% for Qwen3‑8B) with fewer interventions than random or fixed‑interval baselines.

Why it matters

Safety researchers and developers of autonomous LLM agents should care because INTENT‑AS‑A‑TOOL provides a low‑overhead, fine‑grained signal to intervene before harmful actions are executed.

Method details
  • Evaluated five open‑weight models: Qwen3‑8B, Qwen3‑32B, Qwen3‑235B‑A22B (MoE), Qwen3.5‑27B, Gemma‑4‑31B‑IT
  • Dataset: agentic‑misalignment benchmark, focusing on risky cases where at least one undefended rollout is judged harmful
  • Inference: thinking mode enabled, three rollouts per prompt, temperature unspecified, using vLLM with automatic prefix caching
  • Baseline: static prompting defense that appends scenario‑specific safety guidance to the system prompt
  • Ablations: timing ablation comparing intent‑guided, random, and fixed‑interval triggers; wording robustness (Table 10); tool‑set size sensitivity (Table 11)
Numbers
  • Case‑level defense success rate, Qwen3‑32B Blackmail, 100.0% (intent‑guided) vs 80.0% (prompting)
  • Case‑level defense success rate, Qwen3‑8B Leaking, 95.6% (intent‑guided) vs 75.0% (prompting)
  • Intent‑guided success, Qwen3‑8B, 97.8% vs random 70.2% (Table 3)
  • Mean intent‑tool probability, Qwen3‑32B Main description, 0.404 (Table 10)
  • Top‑1 rate, Qwen3‑8B original tool set, 0.2185 (Table 11)
  • Mean interventions per rollout, Qwen3‑8B intent‑guided, 1.69 (Table 3)
Limitations

The paper does not guarantee defense effectiveness when the intent tool fails to expose intent early or reliably, as seen with Gemma‑4‑31B‑IT and certain Qwen3.5‑27B scenarios.

Intent‑guided triggering performs best, suggesting that the intent signal identifies consequential decision points at which the model can be redirected.Found in the source text, word for word.

Picked because: Describes INTENT‑AS‑A‑TOOL, a method for detecting agentic misalignment via chain‑of‑thought monitoring, giving engineers a verification technique for LLM agents.

Paper 5 of 5

Decoupled I/O-Dominant Pipelines for Large-Scale Whole-Slide Image Embedding Extraction

Mayanka Chandrashekar, Xi Zhang, Ethan Seefried and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.DC

Problem

Patch-based WSI processing generated massive numbers of small files, causing severe I/O and orchestration overhead that dominated end-to-end performance. Simply adding more compute resources did not help because storage bandwidth, not compute, became the bottleneck.

Approach

The authors decouple the workflow into three independent stages: (1) MPI‑based patch generation and staging, (2) embarrassingly parallel SPMD embedding inference on GPUs, and (3) shard‑parallel vector‑database ingestion. Each stage is parallelized according to its dominant constraint-storage, compute, or write throughput-so they can scale independently. Patch generation uses deterministic cyclic partitioning of spatial coordinates across MPI ranks and writes patches to rank‑local directories. Embedding inference loads a shared patch index on each GPU task, processes a strided subset of patches, and writes results locally before a final file‑based merge. Ingestion writes each rank's embedding shard to a distributed vector database without coordination, relying on the shared filesystem for storage.

Result

Throughput increased with GPU count but efficiency dropped sharply, indicating storage‑limited scaling; GPU utilization stayed modest across all scales. I/O wait time grew as patch count rose, and the I/O fraction rose from 1.8% on a single GPU to 15.5% on 40 GPUs for H‑Optimus‑0. Vector‑database ingestion scaled from 22K to 188K rows/s (8.5× speedup) but efficiency peaked at intermediate concurrency and fell at higher node counts.

Why it matters

Researchers and engineers building large‑scale pathology pipelines should adopt the decoupled design to avoid storage bottlenecks and achieve higher throughput on HPC clusters.

Method details
  • Foundation models used: HIPT, H‑Optimus‑0, and Virchow2.
  • Patch generation implemented as an MPI program with cyclic coordinate partitioning and rank‑local file writes.
  • Embedding inference executed in SPMD fashion with one GPU per rank, using a local DataLoader and no collective communication.
  • Vector database ingestion performed shard‑parallel, each rank writing its own shard to a distributed database.
  • Strong‑scaling experiments evaluated up to 40 GPUs per model.
  • Baseline comparison is a monolithic end‑to‑end pipeline that does not decouple I/O from compute.
Numbers
  • Throughput 110.7 patches/s for h‑optimus‑0 on 1 GPU vs 855.8 patches/s on 40 GPUs
  • Efficiency 33.9% at 16 GPUs for h‑optimus‑0 (down from 100.0% at 1 GPU)
  • I/O % 15.5 at 40 GPUs for h‑optimus‑0 (up from 1.8% at 1 GPU)
  • Throughput 153.7 patches/s for hipt on 1 GPU vs 1104.1 patches/s on 40 GPUs
  • Efficiency 18.0% at 40 GPUs for hipt (down from 100.0% at 1 GPU)
  • Throughput increases from 22K to 188K rows/s (8.5× speedup) for vector database ingestion
Limitations

The paper does not evaluate how the decoupled pipeline affects downstream model accuracy or task performance, focusing solely on system‑level throughput.

This design isolates data movement from compute, enabling efficient patch delivery, scalable multi-node inference with minimal communication.Found in the source text, word for word.

Picked because: Shows a decoupled, I/O‑aware pipeline for large‑scale whole‑slide image embedding extraction, offering actionable infrastructure automation strategies for high‑throughput inference.