Yufan Wu, Yinghui He, Zhengyi Hu and 4 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Prior inference-time scaling methods improve reasoning but rely on repeated generation or external verification, which is costly; simply generating more samples does not address the inefficiency.
Approach
CritICL leverages structured failure modes observed in weaker models of the same family as guidance for stronger models. It builds a CritBank of failure mode labels and critiques from small-scale models. Two variants are introduced: CritICL-dynamic predicts input‑specific failure modes and retrieves relevant critiques, while CritICL-static uses an aggregated global failure mode profile to retrieve stable examples. During inference the selected critique examples are provided as in‑context demonstrations to the target model. This transfers weak‑model failure information to improve strong‑model reasoning without extra generations.
Result
CritICL-static achieves 93.6% on GSM8K and 59.2% on MATH, raising the overall in‑distribution average to 76.4% and the out‑of‑distribution average to 21.3%, which surpasses standard ICL and matches or exceeds test‑time scaling methods while using far fewer generations.
Why it matters
Researchers seeking efficient inference‑time reasoning improvements should consider CritICL, as it provides strong accuracy gains without the token overhead of repeated sampling.
Method details
Evaluated on GSM8K (7.4k train, 1.3k test) and MATH (7.5k train, 5k test) plus AMC23, AIME24, AIME25 OOD benchmarks
CritBank constructed from responses of Qwen2.5‑1.5B‑Instruct, Qwen2.5‑3B‑Instruct, Qwen2.5‑7B‑Instruct and Llama‑3.2‑1B‑Instruct, Llama‑3.2‑3B‑Instruct, Llama‑3.1‑8B
Target models: Qwen2.5‑32B‑Instruct, Qwen2.5‑72B‑Instruct, Llama‑3.1‑70B‑Instruct, all decoded greedily with temperature 0.0
Baselines include zero‑shot, 1/3/5‑shot random and fixed exemplars, Self‑Consistency@3/5/7, Self‑Reflection, LLM‑as‑Judge (GPT‑4o‑mini)
Ablation compares failure‑mode‑based example selection against random, fixed, and semantic similarity retrieval (Table 3)
Numbers
GSM8K accuracy 93.6 vs 93.0 Consistency@5
MATH accuracy 59.2 vs 58.6 Consistency@5
Overall ID average 76.4 vs 75.8 CritICL‑dynamic
Overall OOD average 21.3 vs 21.1 CritICL‑dynamic
Overall average 49.8 vs 49.1 CritICL‑dynamic
Zero‑shot overall 36.6 vs CritICL‑static 49.8
Limitations
The paper does not establish a causal link between shared inductive biases and the transferability of failure modes, offering only a hypothesis.
CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methodsFound in the source text, word for word.
Picked because: Introduces CritICL, an inference-time framework that boosts LLM reasoning efficiency without costly repeated generation, offering a practical tool for self‑hosted LLM services.
Qianlong Lan, Vinothini Pandurangan, Anuj Kaul and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CR
Problem
Prior evaluations of model security scanners only reported conditional precision/recall for cases where a scanner produced a usable judgment, ignoring how often scanners actually returned a decision; the obvious fix of reporting overall accuracy still hides the large portion of families with no definitive judgment.
Approach
The paper introduces a controlled benchmark that separates non‑N/A coverage, analysis completion, definitive security decisions, non‑security findings, and unsupported outcomes. It evaluates three static scanners-ModelScan, ModelAudit, and Fickling-on a synthetic corpus of Pickle and PyTorch artifacts. The methodology records family‑level decision coverage, conditional detection metrics, and cross‑scanner recovery for incomplete analyses. Latency and rename robustness are also measured. By reporting both judgment availability and conditional accuracy, the approach reveals coverage gaps and redundancy among tools.
Result
ModelAudit produced definitive security decisions for all 135 labeled families (100% coverage), Fickling for 110 families (81.5%), and ModelScan for 67 families (49.6%). Conditional on making a definitive judgment, ModelScan achieved 100% precision, recall, and F1. Fickling added no unique true‑positive families beyond the union of ModelAudit and ModelScan, but both ModelAudit and Fickling recovered detections for all 48 malicious families where ModelScan failed to complete analysis.
Why it matters
Security engineers and tool developers should care because the study shows that high conditional accuracy can coexist with low decision coverage, highlighting the need for multi‑scanner ensembles and coverage‑aware evaluation when protecting model supply chains.
Method details
Synthetic corpus of 170 artifacts organized into 145 specimen families (135 labeled families, 10 malformed).
The benchmark is limited to synthetic Pickle and PyTorch artifacts and does not represent real‑world model‑scanner performance or adversarial evasion scenarios.
Conditional on making a definitive judgment, ModelScan achieved 100% precision, recall, and F1.Found in the source text, word for word.
Picked because: Provides concrete evaluation of AI model security scanners (ModelScan, ModelAudit, Fickling) with released benchmarks, directly useful for DevOps security pipelines.
Hanbing Liu, Bowei Zhang, Changyuan Yu and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.GT
Problem
Traditional advertising relies on predefined slots and cannot embed advertiser influence into the token generation process; simply allocating after generation does not capture token‑level participation and fails to integrate ads seamlessly.
Approach
LAMA lets advertisers submit local continuation value reports at each non‑terminal prefix; the platform verifies Bellman consistency, maintains an allocation posterior over advertisers, and samples an advertiser whose optimal next‑token policy generates the next token; after each token the belief is updated via Bayes and at the end a winning advertiser is sampled and charged; the mechanism satisfies Markov DSIC and IR while approximating KL‑regularized welfare.
Result
Across the three verticals LAMA attains the highest platform welfare (0.5205), revenue (0.8305), advertiser value (0.8568) and user quality (66.5239), outperforming the best allocate‑after policy baseline while preserving response quality.
Why it matters
Ad platforms and generative‑AI services should care because LAMA shows token‑level advertising can increase monetization without harming user experience, and mechanism‑design researchers gain a new sequential auction framework.
Method details
Reference language model is Qwen3‑14B
Report models are advertiser‑conditioned LoRA heads
Dataset is Webis Generated Native Ads 2024 with three verticals: Workout, Vacation, Car
Baselines include six heuristic IC combos plus MOSAIC
Training decomposes reports into local soft advantages and root values using supervised signals
Inference samples advertiser from posterior and updates ledger each token
Numbers
Welfare 0.5205 vs allocate‑after policy 0.5080
Revenue 0.8305 vs allocate‑after policy 0.7501
Advertiser Value 0.8568 vs allocate‑after policy 0.8253
Quality 66.5239 vs allocate‑after policy 65.5939
Welfare 0.5205 vs MOSAIC 0.4390
Revenue 0.8305 vs MOSAIC 0.5274
Limitations
The work is a proof‑of‑concept limited to a single‑winner setting and does not demonstrate large‑scale deployment or multi‑advertiser competition.
LAMA achieves the strongest overall performance among the compared methods, attaining the highest mean platform welfareFound in the source text, word for word.
Picked because: Presents LAMA, a token‑level advertising mechanism that embeds external influence into generation, demonstrating a deployable pattern for custom LLM tooling.
Yutong Zhang, Jianshuo Dong, Peng Xu and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
CoT monitoring only provides post‑hoc, coarse‑grained labels and relies on an external judge, so it cannot show when harmful intent emerges during generation and adding frequent judge calls would add prohibitive cost and latency.
Approach
INTENT‑AS‑A‑TOOL inserts a zero‑parameter intent tool into the model’s tool set and uses the probability of calling this tool on truncated prefixes as a fine‑grained, judge‑free signal of harmful intent. The model’s next‑tool distribution is scored after appending a tool‑call opener, producing a trajectory of intent probabilities. When the probability exceeds a threshold, an online reflection is inserted before decoding continues, allowing intervention at the point of intent emergence. This mechanism replaces costly external judge calls with lightweight prefix scoring and enables dynamic, intent‑guided defenses.
Result
Across the Qwen family, intent‑guided online intervention achieves higher case‑level defense success rates than the prompting baseline in most model‑scenario pairs, e.g., Qwen3‑32B blackmail 100.0% vs 80.0% and Qwen3‑32B murder 63.4% vs 98.9% for prompting. Timing ablation shows intent‑guided triggering yields the highest success (97.8% for Qwen3‑8B) with fewer interventions than random or fixed‑interval baselines.
Why it matters
Safety researchers and developers of autonomous LLM agents should care because INTENT‑AS‑A‑TOOL provides a low‑overhead, fine‑grained signal to intervene before harmful actions are executed.
Method details
Evaluated five open‑weight models: Qwen3‑8B, Qwen3‑32B, Qwen3‑235B‑A22B (MoE), Qwen3.5‑27B, Gemma‑4‑31B‑IT
Dataset: agentic‑misalignment benchmark, focusing on risky cases where at least one undefended rollout is judged harmful
Inference: thinking mode enabled, three rollouts per prompt, temperature unspecified, using vLLM with automatic prefix caching
Baseline: static prompting defense that appends scenario‑specific safety guidance to the system prompt
Intent‑guided success, Qwen3‑8B, 97.8% vs random 70.2% (Table 3)
Mean intent‑tool probability, Qwen3‑32B Main description, 0.404 (Table 10)
Top‑1 rate, Qwen3‑8B original tool set, 0.2185 (Table 11)
Mean interventions per rollout, Qwen3‑8B intent‑guided, 1.69 (Table 3)
Limitations
The paper does not guarantee defense effectiveness when the intent tool fails to expose intent early or reliably, as seen with Gemma‑4‑31B‑IT and certain Qwen3.5‑27B scenarios.
Intent‑guided triggering performs best, suggesting that the intent signal identifies consequential decision points at which the model can be redirected.Found in the source text, word for word.
Picked because: Describes INTENT‑AS‑A‑TOOL, a method for detecting agentic misalignment via chain‑of‑thought monitoring, giving engineers a verification technique for LLM agents.
Mayanka Chandrashekar, Xi Zhang, Ethan Seefried and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.DC
Problem
Patch-based WSI processing generated massive numbers of small files, causing severe I/O and orchestration overhead that dominated end-to-end performance. Simply adding more compute resources did not help because storage bandwidth, not compute, became the bottleneck.
Approach
The authors decouple the workflow into three independent stages: (1) MPI‑based patch generation and staging, (2) embarrassingly parallel SPMD embedding inference on GPUs, and (3) shard‑parallel vector‑database ingestion. Each stage is parallelized according to its dominant constraint-storage, compute, or write throughput-so they can scale independently. Patch generation uses deterministic cyclic partitioning of spatial coordinates across MPI ranks and writes patches to rank‑local directories. Embedding inference loads a shared patch index on each GPU task, processes a strided subset of patches, and writes results locally before a final file‑based merge. Ingestion writes each rank's embedding shard to a distributed vector database without coordination, relying on the shared filesystem for storage.
Result
Throughput increased with GPU count but efficiency dropped sharply, indicating storage‑limited scaling; GPU utilization stayed modest across all scales. I/O wait time grew as patch count rose, and the I/O fraction rose from 1.8% on a single GPU to 15.5% on 40 GPUs for H‑Optimus‑0. Vector‑database ingestion scaled from 22K to 188K rows/s (8.5× speedup) but efficiency peaked at intermediate concurrency and fell at higher node counts.
Why it matters
Researchers and engineers building large‑scale pathology pipelines should adopt the decoupled design to avoid storage bottlenecks and achieve higher throughput on HPC clusters.
Method details
Foundation models used: HIPT, H‑Optimus‑0, and Virchow2.
Patch generation implemented as an MPI program with cyclic coordinate partitioning and rank‑local file writes.
Embedding inference executed in SPMD fashion with one GPU per rank, using a local DataLoader and no collective communication.
Vector database ingestion performed shard‑parallel, each rank writing its own shard to a distributed database.
Strong‑scaling experiments evaluated up to 40 GPUs per model.
Baseline comparison is a monolithic end‑to‑end pipeline that does not decouple I/O from compute.
Numbers
Throughput 110.7 patches/s for h‑optimus‑0 on 1 GPU vs 855.8 patches/s on 40 GPUs
Efficiency 33.9% at 16 GPUs for h‑optimus‑0 (down from 100.0% at 1 GPU)
I/O % 15.5 at 40 GPUs for h‑optimus‑0 (up from 1.8% at 1 GPU)
Throughput 153.7 patches/s for hipt on 1 GPU vs 1104.1 patches/s on 40 GPUs
Efficiency 18.0% at 40 GPUs for hipt (down from 100.0% at 1 GPU)
Throughput increases from 22K to 188K rows/s (8.5× speedup) for vector database ingestion
Limitations
The paper does not evaluate how the decoupled pipeline affects downstream model accuracy or task performance, focusing solely on system‑level throughput.
This design isolates data movement from compute, enabling efficient patch delivery, scalable multi-node inference with minimal communication.Found in the source text, word for word.
Picked because: Shows a decoupled, I/O‑aware pipeline for large‑scale whole‑slide image embedding extraction, offering actionable infrastructure automation strategies for high‑throughput inference.