arXiv digest

Sunday

October 4, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Cristian McGee, El Houcine Bergou, Aritra Dutta · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence while aggressive steps can destabilize it.

Approach

ZFO decouples direction selection from step-size by using a trusted first-order optimizer to compute the update direction and then performing only two zeroth-order function evaluations along that one-dimensional subspace. The three evaluations (gradient plus two probes) are used to build a local model of the objective along the direction, either a Taylor polynomial or a Padé rational approximation. The model is maximized over a bounded interval to obtain a curvature‑aware step length. If a Padé model is ill‑defined or has a pole inside the interval, ZFO falls back to the quadratic Taylor step. This yields an adaptive step‑selection mechanism that costs less than a full line search.

Result

Across all evaluated language models and reasoning benchmarks, ZFO variants frequently achieve higher final evaluation scores than the fixed‑step first‑order baseline AdamW, with the best ZFO variant improving by several points on each dataset.

Why it matters

Researchers and practitioners fine‑tuning large language models should consider ZFO to obtain more efficient and stable step‑size adaptation without costly line searches.

Method details
  • Qwen-2.5-Math-1.5B evaluated on GSM8K, MATH, SVAMP, AsDiv, OpenBookQA
  • Phi-2 evaluated on SVAMP, AsDiv, OpenBookQA
  • Gemma-2-2B evaluated on SVAMP
  • Llama-3.2-1B evaluated on GSM8K, AsDiv, OpenBookQA
  • Baselines compared: AdamW (first‑order) and MeZO (zeroth‑order) with three random seeds
Numbers
  • GSM8K Qwen-2.5-Math-1.5B: 81.00 (Taylor2) vs AdamW 78.24
  • MATH Qwen-2.5-Math-1.5B: 40.10 (Taylor2) vs AdamW 31.77
  • SVAMP Qwen-2.5-Math-1.5B: 90.78 (Taylor2) vs AdamW 88.11
  • OpenBookQA Qwen-2.5-Math-1.5B: 60.67 (Padé2) vs AdamW 26.07
Limitations

The paper only guarantees convergence to a neighborhood of a stationary point, not to an exact stationary point, and the Padé analysis is deferred to the appendix.

Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it.Found in the source text, word for word.

Picked because: Introduces the ZFO framework that separates direction selection from step-size, offering a lightweight, practical optimizer for fine‑tuning large language models that can be directly adopted in production pipelines.

Paper 2 of 5

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Standard decoding of looped Transformers discards intermediate states, wasting the weak predictions that earlier loops produce, and simply keeping the final output leaves a gap in leveraging these predictions.

Approach

LoopCD introduces a training‑free contrastive decoding that compares the final token prediction with an earlier recurrent pass. Two variants are offered: LoopCD‑Logits, which adds one extra output pass to obtain logits from the earlier state, and LoopCD‑Hidden, which operates on hidden states with no extra output pass. The method treats the earlier state as a weak prediction and the final state as a strong prediction, forming a contrastive signal that guides token selection. It can be applied at full recurrent depth or with reduced loops, preserving the same checkpoint and prompt settings. Adaptive and fixed strength rules control the contrastive weighting.

Result

LoopCD‑Logits raises Ouro‑2.6B‑Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD‑Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Across four model families, every multiple‑choice benchmark mean improves (e.g., Ouro‑1.4B mean +0.58). Halving the recurrent loops still matches or exceeds full‑depth unguided baselines, cutting forward FLOPs by 22.5% to 48.2%.

Why it matters

Researchers and engineers using looped Transformers can obtain higher decoding quality without extra training and can halve inference loops to save compute, making deployment more efficient.

Method details
  • Ouro‑2.6B‑Thinking model evaluated on AIME 2024
  • Huginn‑0125 model evaluated on HumanEval
  • Parcae‑370M and Parcae‑1.3B evaluated on seven multiple‑choice benchmarks
  • Looped‑Qwen3 built from frozen Qwen3‑4B with a four‑layer recurrent window
  • Baselines are unguided decoding using the same checkpoints, prompts, and recurrent depth
  • Ablations include fixed vs adaptive strength and Logits vs Hidden variants
Numbers
  • AIME 2024 pass@1 61.88% → 73.33% compared to unguided baseline
  • HumanEval pass@1 22.56% → 31.71% compared to unguided baseline
  • Multiple‑choice mean improvement +0.58 for Ouro‑1.4B
  • Multiple‑choice mean improvement +1.59 for Parcae‑1.3B adaptive
  • Forward FLOP reduction 22.5% to 48.2% when halving loops
Limitations

The paper does not evaluate LoopCD on non‑looped transformer architectures or on tasks beyond the reported benchmarks.

LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%.Found in the source text, word for word.

Picked because: Shows how to decode looped transformers with negligible overhead, delivering near‑free inference speedups that engineers can apply to reduce latency and cost in deployed models.

Paper 3 of 5

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang and 10 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Repeated use of the same external source is treated as repeated access rather than an opportunity to progressively improve understanding, so agents cannot build reusable source‑specific competence.

Approach

SourceLearn builds a persistent source model of entity representations and updates it via two mechanisms. Self‑Directed Source Learning inspects the current model, identifies gaps, and revisits the authoritative source to reconstruct missing knowledge. Task‑Guided Source Learning uses downstream task experience to expose local representational gaps and recurring needs, refining the model accordingly. Updates are only committed when supported by source evidence, and at inference the model activates up to k tokens of relevant regions alongside retrieved evidence.

Result

SourceLearn is best in 13 of 15 settings and second‑best in one more, improving over Hybrid RAG by an average of +14.3, +4.9, and +13.4 points for the three backends, with individual gains up to 22.6 points.

Why it matters

Researchers and engineers building LLM agents that repeatedly consult the same knowledge source should care because SourceLearn provides a systematic way to accumulate and reuse source‑specific competence, yielding consistent performance gains.

Method details
  • LLM backends: GPT-5.6-Luna, gpt-oss-120b, DeepSeek-V4.1-Flash.
  • Embedding model: text-embedding-3-large for all methods.
  • Benchmarks: MultiDoc2Dial, NarrativeQA, SWE-QA, APIBench, AppWorld.
  • Baselines compared: Hybrid RAG, RAPTOR, HippoRAG 2, AWM.
  • Ablations remove Self‑Directed, Task‑Guided, cross‑task representation, or failure‑guided local refinement.
Numbers
  • 13 of 15 settings best performance
  • +14.3 points over Hybrid RAG (GPT-5.6-Luna)
  • +4.9 points over Hybrid RAG (gpt-oss-120b)
  • +13.4 points over Hybrid RAG (DeepSeek-V4.1-Flash)
  • up to 22.6 points over Hybrid RAG
Limitations

The paper does not claim any limitations; none are explicitly stated.

SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG.Found in the source text, word for word.

Picked because: Presents a concrete system for source‑specific competence in LLM agents, including released code and evaluation, enabling reliable integration of external knowledge sources in self‑hosted deployments.

Paper 4 of 5

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CV

Problem

Prior on-policy self-distillation for MLLMs relied on privileged visual cues such as image crops, which required human-annotated grounding data or external teacher models and only helped tasks that benefit from visual zooming.

Approach

Where-OPD introduces on-policy self-distillation where a frozen teacher receives textual, spatially grounded hints that identify the locations of question-relevant objects in procedurally generated scenes, while the student sees only the raw image and question. The teacher scores the student’s sampled prefixes and provides a distillation loss, updating only the student parameters. Spatial guidance is automatically extracted from synthetic scenes, eliminating the need for human annotations or external teachers. The method is applied with a multiple‑choice protocol for smaller models and an open‑ended protocol for larger models. Training runs for a single epoch over the synthetic counting scenes.

Result

Where-OPD consistently outperforms all baselines, achieving the highest average accuracy across the 15 benchmarks for each MLLM. For Qwen3.5-4B the average rises from 73.86% to 78.31%, for Qwen3.5-9B from 71.81% to 77.57%, and for Qwen3-VL-4B from 65.44% to 72.82%. The gains also translate to a 3.23‑point improvement on the combined real‑world suite (CVBench, V*, ZoomBench, BLINK, HR‑Bench, MME‑RealWorld).

Why it matters

Researchers building multimodal large language models can use spatially grounded synthetic supervision to improve a broad range of visual reasoning abilities without extra annotation or external teachers.

Method details
  • Base models: Qwen3.5-4B, Qwen3.5-9B, Qwen3-VL-4B.
  • Synthetic dataset: procedurally generated counting scenes with object identities and coordinates.
  • Training: one epoch; Where-OPD results for Qwen3.5-4B and Qwen3.5-9B averaged over three runs with different seeds.
  • Baselines compared: Vision-OPD, OPD-V, Imagine-OPD, S2VOPD.
  • Ablations include OPSD variants with no privileged info, cropped‑image hint, answer‑only hint, spatial guidance with/without total count, and teacher EMA rate.
Numbers
  • Average accuracy Qwen3.5-4B base model 73.86% vs Where-OPD 78.31% (+4.45 pts)
  • Average accuracy Qwen3.5-9B base model 71.81% vs Where-OPD 77.57% (+5.76 pts)
  • Average accuracy Qwen3-VL-4B base model 65.44% vs Where-OPD 72.82% (+7.38 pts)
  • 3.23‑point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR‑Bench, and MME‑RealWorld
  • Standard deviation of Avg across runs 0.20 for Qwen3.5-4B and Qwen3.5-9B
Limitations

The paper does not discuss any explicit limitations.

Our method achieves the highest average across all evaluated benchmarks on each of the MLLMs.Found in the source text, word for word.

Picked because: Proposes Where‑OPD, an on‑policy self‑distillation pipeline for multimodal LLMs using synthetic scenes, with open artifacts that let engineers improve model reasoning without extra training data.

Paper 5 of 5

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

Yinheng Li, Justin Wagle · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Prior Jev-style pipelines scored each option in isolation, leading to lower accuracy (e.g., jev-local 74.9% on JevBench) and could not handle arbitrary numbers of candidates or multimodal inputs; the obvious fix of using isolated scoring or letter‑logit readouts fails because it lacks joint conditioning and is limited to 16‑26 options.

Approach

LLM2Jev extracts decisions from the next‑token probabilities of a causal LLM by wrapping each option in bracketed numeric identifiers, accumulating log‑probabilities for each suffix and applying a softmax over all candidates. It scores candidates jointly after the full prompt, caches the prefix KV activations and evaluates suffixes in parallel. For fine‑tuning it uses a tree‑factorized listwise loss together with KL‑divergence anchors that keep auxiliary predictions close to the base model. LoRA adapters are applied for parameter‑efficient adaptation, and option order is randomized during training to reduce position bias.

Result

The frozen Qwen3.5‑4B achieves 81.4% accuracy on JevBench, matching or exceeding community Jev models, with an ECE of 0.057, and supports 77‑way Banking77 at 69.0% accuracy. LoRA fine‑tuning raises JevBench accuracy to 84.0% while preserving multimodal performance (82.4% on MMBench). The 0.6B model improves only modestly to 56.7% accuracy and shows poor calibration (ECE 0.278).

Why it matters

Teams can deploy LLMs as decision engines without any fine‑tuning, saving compute, while fine‑tuning is only needed for smaller backbones or niche tasks.

Method details
  • Evaluated on Qwen3.5‑4B and Qwen3‑0.6B causal LLMs
  • Training‑free inference uses joint candidate scoring via accumulated log‑probs and softmax
  • Fine‑tuning mixes include Intent‑100, Intent‑1ep, Reason‑9k, Reason‑21k, Reason‑30k, Reason+Long‑12k
  • Baselines: SemIf, reflex 4B, open‑alternative‑jev, Winnow‑12B, GPT‑5.6 Sol
  • Ablations include full‑parameter vs LoRA fine‑tuning and varying KL anchor weights
  • Parallel suffix scoring batches candidates of equal length to reduce forward passes
Numbers
  • JevBench accuracy 81.4% for 4B training‑free vs 86.6% Winnow‑12B
  • ECE 0.057 for 4B training‑free vs 0.061 SemIf
  • Banking77 accuracy 69.0% for 4B training‑free (letter‑logit readouts unsupported)
  • MMBench zero‑shot accuracy 82.4% for Qwen3.5‑4B
  • 0.6B training‑free JevBench accuracy 56.7%
  • LoRA 4B JevBench accuracy 84.0%
Limitations

The paper does not prove that the approach generalizes to all possible decision tasks or to larger models beyond 4 B, and it only evaluates on the listed benchmarks.

Without task-specific adaptation, frozen Qwen3.5-4B correctly resolves 188 of the 231 public JevBench instances (81.4%; Table 2).Found in the source text, word for word.

Picked because: Demonstrates that standard LLMs already behave as Jev‑style decision models and provides a fine‑tuning recipe to turn them into robust, deployable classifiers, directly useful for building LLM‑driven agents.