Yifan Yang, Zhaoyan Wang, Zheng Gao and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
Zero-cost proxies rank architectures cheaply but their reliability varies across search spaces; under a unified protocol no single proxy robustly beats the trivial #Params/FLOPs baseline across both structure‑varying and size‑varying spaces, and the obvious fix of using a single proxy such as ZiCo fails to be space‑adaptive.
Approach
CoRA‑NAS is a two‑stage framework. Stage 1 (CoRA‑Rank) forms an equal‑weight log‑rank consensus of existing capacity and structure‑at‑initialization zero‑cost proxies and applies a target‑free consensus gate to adapt the proxy bank per space. Stage 2 (CoRA‑Refine) samples stratified anchor architectures from the prior, runs early validation curves at roughly 1% of full‑training cost, and learns a residual correction with an ExtraTrees model that refines the static ranking. The method never uses fully trained architecture‑accuracy labels to fit the ranker. A single configuration is used across all spaces, with only the architecture encoding being space‑specific.
Result
CoRA‑Refine attains mean Spearman correlations of 0.946 on NAS‑Bench‑201, 0.715 on NAS‑Bench‑101, 0.786 on TransNAS‑Bench‑101, and 0.894 on NATS‑SSS, making its worst‑space correlation (0.715) the highest among all compared methods. On NAS‑Bench‑201/CIFAR‑100 the architecture selected by CoRA‑Refine reaches 73.32% test accuracy, essentially matching the reported ground‑truth best of 73.37%. The refinement uses roughly 1% of the cost of fully training the candidate set.
Why it matters
Researchers and practitioners needing low‑cost, robust NAS across diverse search spaces can use CoRA‑NAS to obtain high‑quality rankings and strong architecture selections without full training, saving substantial compute.
Method details
Benchmarks: NAS‑Bench‑201, NAS‑Bench‑101, TransNAS‑Bench‑101, and NATS‑SSS.
Zero‑cost proxies reused include SNIP, GraSP, Synflow, GradSign, ZiCo, NASWOT/jacov, Zen‑score, TE‑NAS, MeCo, SWAP, and Dextr.
Stage 2 refinement employs an ExtraTrees regressor on early validation curves costing about 1% of full training.
Baseline comparisons include #Params, AZ‑NAS, random search, Synflow, LIBRA‑NAS, and other zero‑cost proxies.
Anchor‑count ablation shows Spearman correlation rising from 0.493 with 10 anchors to 0.563 with 1000 anchors.
Numbers
Spearman correlation 0.946 on NAS‑Bench‑201
Spearman correlation 0.715 on NAS‑Bench‑101
Spearman correlation 0.786 on TransNAS‑Bench‑101
Spearman correlation 0.894 on NATS‑SSS
Selected architecture accuracy 73.32% on NAS‑Bench‑201/CIFAR‑100 vs ground‑truth best 73.37%
Refinement cost approximately 1% of full‑training budget
Limitations
On a pure size space (NATS‑SSS) the unsupervised prior is beaten by #Params; the curve‑residual only matches the strongest capacity proxies and does not provide a clear ranking or selection win, and the architecture encoding must be defined per space.
CoRA‑Refine achieves mean Spearman correlations of 0.946, 0.715, 0.786, and 0.894, respectively.Found in the source text, word for word.
Picked because: CoRA-NAS delivers a two‑stage, zero‑cost proxy‑based neural architecture search framework with released code, letting engineers quickly evaluate and refine models without expensive training runs.
Previous first-order optimizers lacked an adaptive mechanism for controlling update magnitudes based on cosine similarity, and simply scaling gradients does not provide that adaptivity.
Approach
AdamX augments a standard first-order optimizer with a cosine similarity term that dynamically adjusts the magnitude of parameter updates. It computes the cosine similarity between gradients and uses it as a scaling factor. A variance rectification scheme is added to smooth optimization during early training stages. The method is designed to be scalable and model‑agnostic, allowing straightforward integration into existing pipelines. Training is evaluated by counting epochs needed to reach predefined performance thresholds under a fixed hyperparameter budget.
Result
The paper reports that AdamX achieves competitive convergence rates, requiring fewer epochs to reach target performance compared with existing baselines under the same hyperparameter budget.
Why it matters
Practitioners seeking faster convergence and smoother early‑training dynamics should consider AdamX as a drop‑in replacement for existing optimizers.
Method details
Evaluated on benchmark datasets
Tested across a range of neural network architectures
Performance measured by number of epochs to reach predefined thresholds
Hyperparameter budget kept fixed for all experiments
Limitations
The abstract does not specify any limitations of the proposed method.
AdamX achieves competitive convergence rates across a range of benchmark datasets and architectures.Found in the source text, word for word.
Picked because: AdamX introduces a cosine‑similarity‑driven optimizer and variance‑rectification scheme that can be dropped into existing training pipelines, offering a practical performance boost with minimal integration effort.
Yi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed and 1 others · abstract · pdf
quote verifiedfigures checkedread: full texteess.AS
Problem
SpeechLLMs lag behind text-only LLMs on complex reasoning and suffer an accuracy‑latency trade‑off; simply adding an external LLM to evaluate and rewrite model outputs does not help because of a mismatch between the external evaluator and the streaming model.
Approach
RetroThinker is a multi‑stage post‑training framework applied to the Moshi SpeechLLM. It first performs supervised fine‑tuning (SFT) on curated retrospective thinking data (Stage 1 rule‑based, Stage 2 sample‑based) using LoRA. Then it applies length‑based direct preference optimization (LengthDPO) to encourage shorter reasoning traces during early reasoning. The system streams user audio, generates CoT steps, self‑verifies them, and forward‑corrects errors without rolling back. Early reasoning is controlled by Question Completeness (QC) thresholds, and LengthDPO further reduces latency.
Result
The best configuration (EarlyReasoning + Retro + DPO at 95% QC) yields an 11 percentage‑point absolute accuracy improvement while maintaining latency comparable to Standard Reasoning, and it reduces latency by about 30% relative to Standard Reasoning + Retro with similar accuracy. RetroThinker adds roughly two seconds of latency before LengthDPO is applied.
Why it matters
Researchers and engineers building streaming speech language models for voice assistants should care because RetroThinker offers a way to boost reasoning accuracy without large latency penalties.
Method details
Model: Moshi SpeechLLM fine‑tuned with LoRA
Dataset: GSM8K with audio synthesized by Gemini‑Flash‑TTS
SFT learning rates: searched between 2e-5 and 2e-6
DPO learning rates: searched between 1e-7 and 1e-6, with length‑normalized probability distribution
Baselines: non‑retrospective EarlyReasoning [7] and Standard Reasoning
Ablations: Stage 1 only, Stage 1+2 across QC thresholds, and before/after LengthDPO CoT length analysis
Numbers
Accuracy gain, 11%, compared to Standard Reasoning
Latency reduction, ~30%, compared to Standard Reasoning + Retro
Latency overhead, ~2 seconds, before applying LengthDPO
Accuracy, 24%, EarlyReasoning + Retro + DPO at 95% QC
Evaluation is limited to a TTS version of GSM8K and relies on Whisper transcription plus an LLM judge, so results may not generalize to natural noisy speech or other tasks, and latency control via QC is weakened.
Our best setting, EarlyReasoning + Retro + DPO with the 95% QC threshold, achieves an 11 percentage-point absolute accuracy improvement while keeping nearly the same latency compared to Standard Reasoning.Found in the source text, word for word.
Picked because: RetroThinker adds retrospective thinking capabilities to speech‑LLMs, providing concrete tooling for building and verifying LLM agents that need to reason over prior utterances.
Current state-of-the-art membership inference attacks require training computationally expensive shadow models, making large‑scale privacy evaluation impractical. Using the conventional generalisation gap as a proxy does not capture all privacy leakage.
Approach
The authors compute inexpensive spectral metrics from trained weight matrices using the WeightWatcher framework, specifically stable rank and Log‑alpha‑Norm. These metrics are evaluated on a suite of MLP models trained with varied depths, widths, learning rates and weight decay. The resulting spectral values are correlated with privacy leakage measured by the LiRA membership inference attack. The strength of these correlations is compared against the conventional generalisation gap to assess predictive power. By showing stronger associations, the method proposes spectral analysis as a scalable privacy auditing tool.
Result
Stable rank shows the strongest positive correlation with LiRA AUC across both CIFAR‑10 and Volkert, while Log‑Norm exhibits a consistent negative correlation with LiRA [email protected] in the low false‑positive regime. Both relationships are stronger than those observed for the conventional generalisation gap, indicating that spectral metrics capture privacy leakage information not reflected by overfitting measures.
Why it matters
Privacy auditors and ML engineers can use cheap spectral analyses to flag models with high privacy risk without running expensive shadow‑model attacks.
Method details
Feedforward multi‑layer perceptrons (MLPs) with ReLU activations, no dropout or normalisation, and a 10‑class linear output layer
Training with SGD (momentum 0.9), batch size 32, learning rates {0.001,0.01}, weight decay {0, }, for 100 epochs
Spectral metrics computed via WeightWatcher: stable rank and Log‑Norm (Log‑alpha‑Norm)
Privacy risk measured with LiRA (AUC and [email protected]) using the SACRO‑ML implementation
Baseline comparison using the conventional generalisation gap
Numbers
total models, 44, across datasets
learning rates, {0.001,0.01}, values explored
weight decay values, {0, }, values explored
batch size, 32, fixed across runs
epochs, 100, training length
CIFAR‑10 images, 60,000, dataset size
Limitations
The study is limited to MLPs on two datasets and a single attack (LiRA), and it does not establish a causal relationship between spectral properties and privacy leakage.
stable rank exhibits the strongest association with LiRA AUCFound in the source text, word for word.
Picked because: Predicting Privacy Leakage from Weight Spectral Density proposes a cheap spectral metric to estimate membership‑inference risk, enabling engineers to audit model privacy without costly shadow‑model training.
Adithiyan Rajan Indira Saravanan, Kathleen C. Fraser · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Retrieval-augmented generation (RAG) degrades the safety of LLM responses to harmful queries, and simply applying existing safety guardrails does not prevent this degradation.
Approach
The authors introduce RAG‑Safety‑Bench, a benchmark that isolates safety effects by removing retriever quality confounds and defining four evaluation conditions: non‑RAG, RAG with an oracle document containing the answer, RAG with on‑topic but non‑answer documents, and RAG with random safe documents. By directly supplying context documents, the benchmark cleanly separates the impact of document relevance from retrieval performance. Each condition uses the same deterministic prompt template and decoding settings, allowing direct comparison of safety outcomes across conditions. The benchmark also includes a taxonomy of 20 harmful subcategories derived from the MLCommons AILuminate hazard taxonomy.
Result
Across the five models the benchmark shows an inverse relationship between benign and unsafe capability, strong evidence that baseline safety guardrails do not guarantee safety in RAG, and model‑specific cases where even benign documents cause unsafe generation.
Why it matters
Researchers and practitioners building RAG systems for corporate or knowledge‑base integration should be aware that safety guardrails may not suffice and need dedicated evaluation of retrieval‑induced safety risks.
Method details
Five open‑source LLMs evaluated: Gemma‑3‑12B‑It, Llama‑3.1‑8B‑Instruct, Ministral‑3‑8B‑Instruct, Qwen‑2.5‑7B‑Instruct, Phi‑4‑14B.
Each model generated up to 1024 new tokens per unsafe query.
Experiments run on academic research cluster with NVIDIA H100 GPUs; total compute ~74 GPU‑hours.
Benchmark constructed from Wikipedia using a semi‑automated pipeline and Claude Sonnet‑4.5 for question generation.
Balanced subset contains 346 examples; expanded to 1,384 example‑condition rows per model, yielding 6,920 total responses.
Responses generated deterministically with temperature and top‑p disabled, fixed seed 42.
Numbers
6,920 total responses across five models
74 GPU‑hours total compute
95% of cases returned a valid true/false label
Cohen’s Kappa between 0.90 and 0.92 for human vs. automated scoring
Oracle unsafe condition: 200 yes (57.8%) and 146 partial (42.2%)
On‑topic safe and random control conditions: 346 no (100%) each
Limitations
The study only evaluates open‑source models and does not test commercial LLM APIs.
baseline safety guardrails do not lead to downstream safety guarantees in the RAG caseFound in the source text, word for word.
Picked because: RAG‑Safety‑Bench supplies an open‑source benchmark suite for evaluating the safety of retrieval‑augmented LLMs, giving practitioners a concrete method to verify and harden LLM‑based agents.