arXiv digest

Saturday

September 19, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

dQwen3.5: Hybrid-Attention Diffusion Language Models

Anton Xue, Litu Rout, Aditya Akella and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Hybrid autoregressive backbones could not be directly adapted to diffusion language models because their RNN (Gated DeltaNet) layers are inherently causal and nontrivial to bidirectionalize, so simply making them bidirectional does not work.

Approach

The method bidirectionalizes only the attention layers of Qwen3.5 by disabling causal masks while leaving the Gated DeltaNet layers causal. It adds token shifting by prepending a BOS token and shifting the readout by one position to preserve AR next-token alignment. Special diffusion tokens (mask, padding, BOS) are repurposed from unused IDs. The hybrid blocks retain their original structure of three GDN layers followed by one attention layer. This enables efficient AR-to-DLM adaptation without redesigning the recurrent components.

Result

Hybrid backbones reach a given training loss in about half the tokens compared to the full‑attention control, and dQwen3.5-2B after 50B tokens outperforms CoDA (trained for 200B tokens) on 6 of 7 benchmarks; dQwen3.5-9B after 50B tokens attains the best DLM score on 4 of 7 benchmarks against Dream-7B, Dream-Coder-7B, and LLaDA-8B.

Why it matters

Researchers and engineers interested in efficiently converting hybrid autoregressive models to diffusion language models can adopt this approach to achieve strong downstream performance with far fewer training tokens.

Method details
  • Adapted Qwen3.5 models at 0.8B, 2B, 4B, and 9B parameter scales.
  • Hybrid architecture: each block contains three Gated DeltaNet layers and one attention layer; attention constitutes 25% of layers.
  • Bidirectionalized only the attention layers by disabling causal masks; GDN layers remain causal.
  • Training used 50B adaptation tokens for dQwen3.5 models.
  • Baselines: full‑attention Qwen3-1.7B control, CoDA (200B tokens), Dream-7B (580B tokens), Dream-Coder-7B (322B tokens), LLaDA-8B (2.3T tokens).
  • Evaluation used a common decoding scheme with a 1024‑token canvas and block size 32.
Numbers
  • dQwen3.5-2B after 50B adaptation tokens outperforms CoDA, trained for 200B tokens, on 6/7 benchmarks
  • dQwen3.5-9B after 50B tokens attains best DLM score on 4/7 benchmarks against Dream-7B (580B tokens), Dream-Coder-7B (322B tokens), and LLaDA-8B (2.3T tokens)
  • Hybrid reaches a given training loss in about half the tokens versus full‑attention control
  • dQwen3.5-9B total parameters 8.95B (trunk 6.92B, embed 1.02B, head 1.02B)
  • Qwen3-1.7B full‑attention control trunk 1.41B, embed 0.31B, total 1.72B
Limitations

The paper does not evaluate post‑training performance and treats post‑training as a separate problem.

dQwen3.5-2B after 50B adaptation tokens outperforms CoDA, trained for 200B tokens, on 6/7 benchmarks.Found in the source text, word for word.

Picked because: Provides a released hybrid-attention diffusion language model and adaptation code, giving engineers a concrete way to upgrade existing autoregressive LLMs for generative tasks.

Paper 2 of 5

On-Demand Attention: Language Models Know When to Recall

Haibo Feng, Ruiqi Liang, Hanyang Peng and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Full-attention decoding always reads the entire growing history at each step, making long-context inference inefficient, while a naive local-only policy cuts cost but severely degrades generation quality.

Approach

On-Demand Attention (ODA) uses a local-first decoding scheme where a lightweight recall head predicts the benefit of a global read from the model's decoding states. When the predicted benefit exceeds a threshold, ODA invokes full attention for that step; otherwise it stays local. Only the recall head is trained; all pretrained weights and the full KV cache remain unchanged. GPU-side conditional execution in vLLM implements the selective global reads, turning reduced reads into actual speedups.

Result

ODA restores most of the quality lost by the Local policy while cutting the proportion of full‑attention calls, and it yields substantial decoding speedups on long prompts. For Qwen3‑8B on RULER16K, ODA scores 91.07 versus 92.59 for Full and 25.43 for Local with only 47.32% Full calls. For Gemma‑4‑12B‑it on RULER16K, ODA scores 95.27 versus 96.61 for Full and 30.07 for Local with 57.8% Full calls. Warm decoding throughput on a 128,852‑token prompt rises from 75.54 TPS (Full) to 149.52 TPS (ODA).

Why it matters

Researchers and engineers building long‑context language‑model applications can use ODA to keep generation quality while reducing compute and latency, especially when deployed on GPUs with vLLM.

Method details
  • Qwen3-1.7B, Qwen3-8B, Qwen3.5-2B and Gemma-4-12B-it models are evaluated
  • Datasets: RULER16K (13 tasks, 100 examples each) and LongBench v1 (13 tasks, 2,550 examples)
  • Recall head has 28,325,889 parameters, uses RMSNorm, SiLU, two SwiGLU blocks and FP32 arithmetic
  • Training data: 196,608 long‑context SFT examples up to 16,384 tokens, BF16 frozen backbone, AdamW, global batch size 64, learning‑rate schedule as described
  • Inference: Local window of 2,048 tokens with four initial positions; Full‑call rates token‑weighted; vLLM 0.18 on a single A100‑SXM4‑80GB
  • Baselines: native Full attention and a fixed Local policy; ablations include varying full‑call penalty and history threshold
Numbers
  • RULER16K Qwen3‑8B Full score 92.59, Local 25.43, ODA 91.07 at 47.32% Full calls
  • RULER16K Gemma‑4‑12B‑it Full score 96.61, Local 30.07, ODA 95.27 at 57.8% Full calls
  • LongBench v1 Qwen3‑1.7B macro scores Full 37.9385, Local 23.6554, ODA 36.8185
  • Warm decode TPS at 128,852 tokens: Full 75.54, ODA 149.52
  • Recall head parameter count 28,325,889
  • Full‑call penalty for primary Qwen3‑1.7B run 0
Limitations

The paper does not provide LongBench ODA results for Gemma‑4‑12B‑it or Qwen3.5‑2B, and the speedup measurements are limited to the specific hardware and prompt lengths without confirming quality at extreme token counts.

selective recall recovers most of the performance lost under local attention while substantially reducing global readsFound in the source text, word for word.

Picked because: Introduces On-Demand Attention, a runtime optimization that lets production systems skip unnecessary full‑attention reads, reducing compute costs for long‑context LLM inference.

Paper 3 of 5

Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Prior weight‑space steering attempts failed because LayerNorm normalizes activations after the residual addition, attenuating small weight perturbations by about a factor of 1000, so the induced logit shifts were orders of magnitude too small to change decisions.

Approach

Deep Noir autonomously discovers steering parameters by (1) ranking layers with a combined logit‑lens differentiation and antagonist head strength score, (2) isolating the most influential heads via gradient magnitude voting, (3) computing a contrastive direction from mean hidden states of each class, (4) calibrating the intervention magnitude with a golden‑section search, and (5) selecting the configuration that maximizes accuracy while preferring fewer heads and lower magnitude. The pipeline is iterated to correct remaining errors. Activation‑space hooks are applied after attention output but before the next LayerNorm, bypassing the attenuation problem. Each component is validated by ablation.

Result

Deep Noir improves spam classification by up to 42.4 percentage points and sentiment classification by up to 13.1 points, with statistically significant gains (p<0.01) across all models; RepE fails to improve over baseline on sentiment.

Why it matters

Practitioners building safety‑critical or agent‑oriented LLM systems should adopt Deep Noir to automatically locate and apply effective activation steering without manual prompt engineering.

Method details
  • 1B, 2‑3B and 7‑9B model scales evaluated (Llama‑3.2‑1B, OLMo‑1B, Gemma‑3‑1B‑IT, Mistral‑7B, etc.)
  • Spam datasets: Enron, SMS Spam Collection, Phishing, SpamAssassin, Ultimate; sentiment dataset: SST‑2
  • Baseline prompting methods (standard, chain‑of‑thought, 5‑shot) and RepE without head masking are compared
  • Component ablations include removing head masking, fixing magnitude, and using a middle layer instead of ranked selection
  • Discovery uses 50 labeled probes per fold (100 probes for baseline comparison) and 5‑fold cross‑validation where applicable
Numbers
  • spam (1B) overall gain +16.7 pts vs. baseline, CI 4.7, 33/39 runs improved
  • sentiment (1B) overall gain +13.1 pts vs. baseline, CI 3.0, 15/15 runs improved
  • spam (7‑9B) Gemma‑2‑9B gain +42.4 pts vs. baseline
  • spam (7‑9B) Llama‑3.1‑8B gain +29.2 pts vs. baseline
  • spam (7‑9B) Mistral‑7B gain +22.0 pts vs. baseline
  • spam (7‑9B) OLMo‑7B gain +21.2 pts vs. baseline
Limitations

The method yields limited gains (≈3%) on reasoning tasks where the model cannot already distinguish correct from incorrect steps, and weight‑space edits remain ineffective due to LayerNorm attenuation.

our engine achieves +16.7 pts on spam at 1B (4.7, 39 runs)Found in the source text, word for word.

Picked because: Presents Deep Noir, an automated framework for discovering and applying activation‑steering hooks in transformer models, enabling practical LLM behavior control without manual tuning.

Paper 4 of 5

Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated Learning

Simon Süwer, Julian Klemm, Elisa Acitelli and 38 others · abstract · pdf

quote verifiedfigures checkedread: abstract onlycs.LG

Problem

Existing federated learning frameworks were evaluated and none fully satisfied the five literature‑derived requirements, leaving clinical studies limited to simulations; simply adopting an existing framework does not work because it lacks integrated data harmonization, discovery, disclosure control, versioned tools, and containerized workflow execution.

Approach

FL‑Net is a federated clinical research framework that combines modular data harmonization, data discovery, disclosure control, versioned FL‑Net‑Tools, and containerized federated workflow execution into a persistent network. The components are linked so that harmonized data can be reused across studies, while discovery services locate patients across datasets. Disclosure control mechanisms enforce privacy during cross‑site queries. Versioned tools ensure reproducibility and auditability of federated workflows. Containerization enables scalable execution with many simultaneous clients. The overall system satisfies all five identified requirements.

Result

FL‑Net was evaluated on harmonization, cross‑study patient discovery across MIMIC and US‑130, and reproducible audited federated workflows, demonstrating support for up to 50 concurrent clients and planning coverage of over 800,000 patients across 10 hospitals in 9 countries.

Why it matters

Researchers and clinicians developing multi‑center federated learning studies should care because FL‑Net provides an integrated, privacy‑preserving platform that can scale across many hospitals and large patient cohorts.

Method details
  • Analyzed 14 federated learning frameworks against five requirements
  • Integrates modular data harmonization
  • Provides cross‑study patient discovery across MIMIC and US‑130
  • Implements disclosure control for privacy preservation
  • Uses versioned FL‑Net‑Tools for reproducible audited workflows
  • Supports containerized federated workflow execution with up to 50 concurrent clients
Numbers
  • 5 requirements, derived from literature
  • 14 FL frameworks analyzed
  • 50 concurrent clients
  • 800,000 patients
  • 10 hospitals
  • 9 countries
Limitations

The paper does not explicitly state any limitations.

FL‑Net's end-to-end capabilities were evaluated through harmonization, cross-study patient discovery across MIMIC and US-130Found in the source text, word for word.

Picked because: Describes FL‑Net, a full‑stack federated learning platform with open‑source components for secure, self‑hosted multi‑center data mining, directly applicable to DevOps and infrastructure automation.

Paper 5 of 5

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

Zimu Han, Yiming Zeng, Jiyao Zhang and 9 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.RO

Problem

Supervised fine-tuning on static demonstrations provides limited coverage of out-of-distribution states and standard imitation objectives cannot distinguish progressing from less useful behavior; interactive post‑training that does address these issues typically requires repeated policy execution and human intervention on a physical robot.

Approach

HIL‑UMI introduces a policy‑guided Universal Manipulation Interface that queries the current policy on the same observation stream during handheld demonstrations without executing its predictions. An Energy Score compares the human action trajectory with the policy inference and triggers data collection when the discrepancy indicates an out‑of‑distribution region. A separate feedback loop uses low online advantage predictions to identify essential segments for refining a progress‑based advantage estimator. The updated estimator then guides advantage‑conditioned behavioral cloning using a balanced mixture of base demonstrations and newly collected policy data, preserving iterative, policy‑aware learning while decoupling data collection from robot deployment.

Result

Across four real‑world manipulation tasks HIL‑UMI consistently improves mean Task Progress Score over SFT and outperforms HG‑DAgger on Clean Up Table while requiring lower per‑frame collection time; the targeted collection incurs higher per‑frame latency but yields steadily rising TPS where SFT plateaus.

Why it matters

Robotics researchers and engineers working on vision‑language‑action models should care because HIL‑UMI enables robot‑free, policy‑aware human‑in‑the‑loop post‑training that improves performance without requiring on‑robot execution.

Method details
  • Base dataset contains 50 demonstrations for long‑horizon tasks (Fold Towel, Clean Up Table) and 80 demonstrations for precise tasks (Stack Cube, Stamp).
  • Per‑round data budget is 12,000 frames for long‑horizon tasks and 2,500 frames for precise tasks.
  • Policy and advantage models use an action horizon of 20, optimizer AdamW, cosine learning‑rate schedule, warm‑up of 500 steps, batch size 128, and 5,000 update steps per round.
  • Online collection uses 10 stochastic policy samples, Energy score weights 0.5, 0.25, 0.25, advantage‑evaluation interval 50, data mixture ratio 0.5, and positive‑advantage fraction 0.3.
  • Baselines compared are Supervised Fine‑Tuning (SFT) and HG‑DAgger; ablations remove the advantage model or vary online collection thresholds.
  • Advantage estimator is trained with a linear progress target and signed regression between two uniformly sampled frames from the same trajectory.
Numbers
  • 5.63× faster data collection compared to HG‑DAgger
  • Fold Towel: SFT 69.43 ms/frame, HIL‑UMI 91.92 ms/frame
  • Clean Up Table: SFT 41.70 ms/frame, HIL‑UMI 73.40 ms/frame
  • Stack Cube: SFT 77.04 ms/frame, HIL‑UMI 89.89 ms/frame
  • Stamp: SFT 64.30 ms/frame, HIL‑UMI 101.02 ms/frame
  • Base dataset size: 50 demonstrations for long‑horizon tasks
Limitations

The paper evaluates only four real‑world tasks and does not demonstrate scalability to more complex or higher‑dimensional manipulation scenarios.

HIL-UMI outperforms HG-DAgger on Clean Up Table with lower per-frame collection timeFound in the source text, word for word.

Picked because: Offers HIL‑UMI, a toolkit for human‑in‑the‑loop post‑training of vision‑language‑action models, delivering concrete pipelines to adapt large models to specific deployment environments.