Cristian McGee, El Houcine Bergou, Aritra Dutta · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence while aggressive steps can destabilize it.
Approach
ZFO decouples direction selection from step-size by using a trusted first-order optimizer to compute the update direction and then performing only two zeroth-order function evaluations along that one-dimensional subspace. The three evaluations (gradient plus two probes) are used to build a local model of the objective along the direction, either a Taylor polynomial or a Padé rational approximation. The model is maximized over a bounded interval to obtain a curvature‑aware step length. If a Padé model is ill‑defined or has a pole inside the interval, ZFO falls back to the quadratic Taylor step. This yields an adaptive step‑selection mechanism that costs less than a full line search.
Result
Across all evaluated language models and reasoning benchmarks, ZFO variants frequently achieve higher final evaluation scores than the fixed‑step first‑order baseline AdamW, with the best ZFO variant improving by several points on each dataset.
Why it matters
Researchers and practitioners fine‑tuning large language models should consider ZFO to obtain more efficient and stable step‑size adaptation without costly line searches.
Method details
Qwen-2.5-Math-1.5B evaluated on GSM8K, MATH, SVAMP, AsDiv, OpenBookQA
Phi-2 evaluated on SVAMP, AsDiv, OpenBookQA
Gemma-2-2B evaluated on SVAMP
Llama-3.2-1B evaluated on GSM8K, AsDiv, OpenBookQA
Baselines compared: AdamW (first‑order) and MeZO (zeroth‑order) with three random seeds
Numbers
GSM8K Qwen-2.5-Math-1.5B: 81.00 (Taylor2) vs AdamW 78.24
MATH Qwen-2.5-Math-1.5B: 40.10 (Taylor2) vs AdamW 31.77
SVAMP Qwen-2.5-Math-1.5B: 90.78 (Taylor2) vs AdamW 88.11
OpenBookQA Qwen-2.5-Math-1.5B: 60.67 (Padé2) vs AdamW 26.07
Limitations
The paper only guarantees convergence to a neighborhood of a stationary point, not to an exact stationary point, and the Padé analysis is deferred to the appendix.
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it.Found in the source text, word for word.
Picked because: Introduces the ZFO framework that separates direction selection from step-size, offering a lightweight, practical optimizer for fine‑tuning large language models that can be directly adopted in production pipelines.
Weihao Liu, Huangjie Zheng, Tianrong Chen and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
Standard decoding of looped Transformers discards intermediate states, wasting the weak predictions that earlier loops produce, and simply keeping the final output leaves a gap in leveraging these predictions.
Approach
LoopCD introduces a training‑free contrastive decoding that compares the final token prediction with an earlier recurrent pass. Two variants are offered: LoopCD‑Logits, which adds one extra output pass to obtain logits from the earlier state, and LoopCD‑Hidden, which operates on hidden states with no extra output pass. The method treats the earlier state as a weak prediction and the final state as a strong prediction, forming a contrastive signal that guides token selection. It can be applied at full recurrent depth or with reduced loops, preserving the same checkpoint and prompt settings. Adaptive and fixed strength rules control the contrastive weighting.
Result
LoopCD‑Logits raises Ouro‑2.6B‑Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD‑Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Across four model families, every multiple‑choice benchmark mean improves (e.g., Ouro‑1.4B mean +0.58). Halving the recurrent loops still matches or exceeds full‑depth unguided baselines, cutting forward FLOPs by 22.5% to 48.2%.
Why it matters
Researchers and engineers using looped Transformers can obtain higher decoding quality without extra training and can halve inference loops to save compute, making deployment more efficient.
Method details
Ouro‑2.6B‑Thinking model evaluated on AIME 2024
Huginn‑0125 model evaluated on HumanEval
Parcae‑370M and Parcae‑1.3B evaluated on seven multiple‑choice benchmarks
Looped‑Qwen3 built from frozen Qwen3‑4B with a four‑layer recurrent window
Baselines are unguided decoding using the same checkpoints, prompts, and recurrent depth
Ablations include fixed vs adaptive strength and Logits vs Hidden variants
Numbers
AIME 2024 pass@1 61.88% → 73.33% compared to unguided baseline
HumanEval pass@1 22.56% → 31.71% compared to unguided baseline
Multiple‑choice mean improvement +0.58 for Ouro‑1.4B
Multiple‑choice mean improvement +1.59 for Parcae‑1.3B adaptive
Forward FLOP reduction 22.5% to 48.2% when halving loops
Limitations
The paper does not evaluate LoopCD on non‑looped transformer architectures or on tasks beyond the reported benchmarks.
LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%.Found in the source text, word for word.
Picked because: Shows how to decode looped transformers with negligible overhead, delivering near‑free inference speedups that engineers can apply to reduce latency and cost in deployed models.
Lucheng Fu, Kejing Xia, Yiyang Wang and 10 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Repeated use of the same external source is treated as repeated access rather than an opportunity to progressively improve understanding, so agents cannot build reusable source‑specific competence.
Approach
SourceLearn builds a persistent source model of entity representations and updates it via two mechanisms. Self‑Directed Source Learning inspects the current model, identifies gaps, and revisits the authoritative source to reconstruct missing knowledge. Task‑Guided Source Learning uses downstream task experience to expose local representational gaps and recurring needs, refining the model accordingly. Updates are only committed when supported by source evidence, and at inference the model activates up to k tokens of relevant regions alongside retrieved evidence.
Result
SourceLearn is best in 13 of 15 settings and second‑best in one more, improving over Hybrid RAG by an average of +14.3, +4.9, and +13.4 points for the three backends, with individual gains up to 22.6 points.
Why it matters
Researchers and engineers building LLM agents that repeatedly consult the same knowledge source should care because SourceLearn provides a systematic way to accumulate and reuse source‑specific competence, yielding consistent performance gains.
Ablations remove Self‑Directed, Task‑Guided, cross‑task representation, or failure‑guided local refinement.
Numbers
13 of 15 settings best performance
+14.3 points over Hybrid RAG (GPT-5.6-Luna)
+4.9 points over Hybrid RAG (gpt-oss-120b)
+13.4 points over Hybrid RAG (DeepSeek-V4.1-Flash)
up to 22.6 points over Hybrid RAG
Limitations
The paper does not claim any limitations; none are explicitly stated.
SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG.Found in the source text, word for word.
Picked because: Presents a concrete system for source‑specific competence in LLM agents, including released code and evaluation, enabling reliable integration of external knowledge sources in self‑hosted deployments.
Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CV
Problem
Prior on-policy self-distillation for MLLMs relied on privileged visual cues such as image crops, which required human-annotated grounding data or external teacher models and only helped tasks that benefit from visual zooming.
Approach
Where-OPD introduces on-policy self-distillation where a frozen teacher receives textual, spatially grounded hints that identify the locations of question-relevant objects in procedurally generated scenes, while the student sees only the raw image and question. The teacher scores the student’s sampled prefixes and provides a distillation loss, updating only the student parameters. Spatial guidance is automatically extracted from synthetic scenes, eliminating the need for human annotations or external teachers. The method is applied with a multiple‑choice protocol for smaller models and an open‑ended protocol for larger models. Training runs for a single epoch over the synthetic counting scenes.
Result
Where-OPD consistently outperforms all baselines, achieving the highest average accuracy across the 15 benchmarks for each MLLM. For Qwen3.5-4B the average rises from 73.86% to 78.31%, for Qwen3.5-9B from 71.81% to 77.57%, and for Qwen3-VL-4B from 65.44% to 72.82%. The gains also translate to a 3.23‑point improvement on the combined real‑world suite (CVBench, V*, ZoomBench, BLINK, HR‑Bench, MME‑RealWorld).
Why it matters
Researchers building multimodal large language models can use spatially grounded synthetic supervision to improve a broad range of visual reasoning abilities without extra annotation or external teachers.
Method details
Base models: Qwen3.5-4B, Qwen3.5-9B, Qwen3-VL-4B.
Synthetic dataset: procedurally generated counting scenes with object identities and coordinates.
Training: one epoch; Where-OPD results for Qwen3.5-4B and Qwen3.5-9B averaged over three runs with different seeds.
Ablations include OPSD variants with no privileged info, cropped‑image hint, answer‑only hint, spatial guidance with/without total count, and teacher EMA rate.
Numbers
Average accuracy Qwen3.5-4B base model 73.86% vs Where-OPD 78.31% (+4.45 pts)
Average accuracy Qwen3.5-9B base model 71.81% vs Where-OPD 77.57% (+5.76 pts)
Average accuracy Qwen3-VL-4B base model 65.44% vs Where-OPD 72.82% (+7.38 pts)
3.23‑point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR‑Bench, and MME‑RealWorld
Standard deviation of Avg across runs 0.20 for Qwen3.5-4B and Qwen3.5-9B
Limitations
The paper does not discuss any explicit limitations.
Our method achieves the highest average across all evaluated benchmarks on each of the MLLMs.Found in the source text, word for word.
Picked because: Proposes Where‑OPD, an on‑policy self‑distillation pipeline for multimodal LLMs using synthetic scenes, with open artifacts that let engineers improve model reasoning without extra training data.
Prior Jev-style pipelines scored each option in isolation, leading to lower accuracy (e.g., jev-local 74.9% on JevBench) and could not handle arbitrary numbers of candidates or multimodal inputs; the obvious fix of using isolated scoring or letter‑logit readouts fails because it lacks joint conditioning and is limited to 16‑26 options.
Approach
LLM2Jev extracts decisions from the next‑token probabilities of a causal LLM by wrapping each option in bracketed numeric identifiers, accumulating log‑probabilities for each suffix and applying a softmax over all candidates. It scores candidates jointly after the full prompt, caches the prefix KV activations and evaluates suffixes in parallel. For fine‑tuning it uses a tree‑factorized listwise loss together with KL‑divergence anchors that keep auxiliary predictions close to the base model. LoRA adapters are applied for parameter‑efficient adaptation, and option order is randomized during training to reduce position bias.
Result
The frozen Qwen3.5‑4B achieves 81.4% accuracy on JevBench, matching or exceeding community Jev models, with an ECE of 0.057, and supports 77‑way Banking77 at 69.0% accuracy. LoRA fine‑tuning raises JevBench accuracy to 84.0% while preserving multimodal performance (82.4% on MMBench). The 0.6B model improves only modestly to 56.7% accuracy and shows poor calibration (ECE 0.278).
Why it matters
Teams can deploy LLMs as decision engines without any fine‑tuning, saving compute, while fine‑tuning is only needed for smaller backbones or niche tasks.
Method details
Evaluated on Qwen3.5‑4B and Qwen3‑0.6B causal LLMs
Training‑free inference uses joint candidate scoring via accumulated log‑probs and softmax
Fine‑tuning mixes include Intent‑100, Intent‑1ep, Reason‑9k, Reason‑21k, Reason‑30k, Reason+Long‑12k
Baselines: SemIf, reflex 4B, open‑alternative‑jev, Winnow‑12B, GPT‑5.6 Sol
Ablations include full‑parameter vs LoRA fine‑tuning and varying KL anchor weights
Parallel suffix scoring batches candidates of equal length to reduce forward passes
Numbers
JevBench accuracy 81.4% for 4B training‑free vs 86.6% Winnow‑12B
ECE 0.057 for 4B training‑free vs 0.061 SemIf
Banking77 accuracy 69.0% for 4B training‑free (letter‑logit readouts unsupported)
MMBench zero‑shot accuracy 82.4% for Qwen3.5‑4B
0.6B training‑free JevBench accuracy 56.7%
LoRA 4B JevBench accuracy 84.0%
Limitations
The paper does not prove that the approach generalizes to all possible decision tasks or to larger models beyond 4 B, and it only evaluates on the listed benchmarks.
Without task-specific adaptation, frozen Qwen3.5-4B correctly resolves 188 of the 231 public JevBench instances (81.4%; Table 2).Found in the source text, word for word.
Picked because: Demonstrates that standard LLMs already behave as Jev‑style decision models and provides a fine‑tuning recipe to turn them into robust, deployable classifiers, directly useful for building LLM‑driven agents.