arXiv digest

Monday

October 5, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

Hui Chen, Xuan Qi, James Xu Zhao and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.NE

Problem

Prior LLM‑guided evolutionary methods like AlphaEvolve optimize performance over a fixed number of iterations, ignoring the cost of LLM calls. The obvious fix of simply limiting iterations does not address the need to maximize gain per unit cost, because token usage and model pricing vary across calls.

Approach

FrugalEvo introduces a cost‑aware loop that separates strategy exploration and code implementation between a high‑cost and a low‑cost LLM. It begins with a cold‑start phase where a low‑cost LLM refines an initial program to create a strong incumbent. The context builder then assembles problem description, incumbent code, and search memory into a prompt. The strategy explorer (high‑cost LLM) proposes multiple design strategies in one API call, and the solution generator (low‑cost LLM) implements and iteratively refines each strategy, ranking them by evaluation score. Search memory updates the context for the next round, and cache‑efficient prompt construction places shared content first to maximize prefix reuse. The process repeats until the predefined cost budget is exhausted.

Result

FrugalEvo matches or surpasses all listed baselines on final solution quality and achieves higher BA‑AUC on nine of the ten mathematical and systems tasks. It attains the highest mean BA‑AUC on all five mathematical tasks under both model configurations and sets a new state‑of‑the‑art on Circle Packing with a total radius of 2.635996 for $1.68 and 2.635990 for $0.55, outperforming multi‑agent methods that cost about $50.

Why it matters

Researchers and practitioners developing cost‑sensitive LLM‑driven optimization pipelines should consider FrugalEvo to reduce API expenses while maintaining or improving solution quality.

Method details
  • Uses GPT‑5.6 Terra for strategy generation and GPT‑5.6 Luna for implementation (high‑cost and low‑cost models respectively)
  • Uses GLM‑5.3 for strategy generation and GLM‑5.3 Flash for implementation in a second configuration
  • Cost budget per run is $2 for mathematical tasks and $1 for systems/algorithmic tasks in the GPT configuration; $1 and $0.5 respectively in the GLM configuration
  • Baselines include OpenEvolve, ShinkaEvolve, AdaEvolve, EvoX and AlphaEvolve for mathematical tasks
  • Ablations test model collaboration, removal of cold‑start, and removal of sequential feedback
  • Evaluated on 20 real‑world tasks: 5 math, 5 systems (ADRS benchmark) and 10 algorithmic tasks from ALE‑Bench‑Lite
Numbers
  • Circle Packing total radius 2.635996 achieved with GPT‑5.6 Terra/Luna for $1.68, surpassing baselines
  • Circle Packing total radius 2.635990 achieved with GLM‑5.3/Flash for $0.55, matching or surpassing baselines
  • Higher BA‑AUC on 9 out of 10 tasks compared to OpenEvolve, ShinkaEvolve, AdaEvolve, EvoX
  • Matches or surpasses state‑of‑the‑art baselines on 10 algorithmic optimization tasks from ALE‑Bench‑Lite
  • CORAL and SwarmResearch cost approximately $50 on average to reach comparable Circle Packing scores
Limitations

The paper does not discuss any limitations of the proposed approach.

FrugalEvo achieves the highest mean BA-AUC on all five tasks under both model configurations.Found in the source text, word for word.

Picked because: Presents FrugalEvo, a cost‑aware LLM‑guided program evolution framework with concrete cost metrics and released code, enabling engineers to optimize compute expense when using LLMs for code generation.

Paper 2 of 5

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

Yu Li, Guangfeng Cai, Long-Fei Li and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing trajectory-level and step-level credit assignment methods do not explicitly trace read‑write dependencies, so training signals are often assigned to irrelevant commands, weakening learning. The obvious fix of uniformly applying the group‑relative advantage to all steps (as in GRPO) does not address this because it still spreads credit to unrelated operations.

Approach

DepGPO first builds a command dependency graph from the execution trace of each trajectory. It then traces backward from the resources inspected by the task verifier to identify relevant write commands and their supporting read commands. Credits for these commands are summed per step, normalized to step factors, and used to redistribute the group‑relative advantage across steps. Tokens generated in each step receive the per‑step advantage, replacing the uniform assignment used by GRPO. This mechanism integrates read‑write dependency information into policy‑gradient updates.

Result

DepGPO attains the highest pass@1 in every configuration, outperforming the strongest competing RL baseline by 3.22 to 10.03 percentage points and beating GRPO by 4.34 to 14.53 points. The method also shows more stable reward, entropy and gradient dynamics compared with GRPO and GiGPO, and its step‑factor variance increases during training, indicating selective credit allocation.

Why it matters

Researchers building terminal agents and credit‑assignment RL methods should care because DepGPO demonstrates that dependency‑aware credit redistribution yields substantial performance gains and training stability.

Method details
  • Evaluated with Qwen3.5‑9B and Qwen3.6‑27B models
  • Training data are SETA and TMAX datasets
  • Benchmarks are Terminal‑Bench 2.0 and 2.1
  • Baselines include GRPO, DAPO, DPPO, GiGPO and GraphGPO
  • Ablations test binary write contribution, verifier filtering, supporting‑read credit and direct writes only
  • Training batches contain four tasks with eight trajectories per task (32 trajectories total)
Numbers
  • DepGPO 25.24% pass@1 vs DPPO 15.36% on Qwen3.5‑9B SETA TB2.0
  • DepGPO 24.64% vs DPPO 15.81% on Qwen3.5‑9B SETA TB2.1
  • DepGPO 20.45% vs DPPO 17.23% on Qwen3.5‑9B TMAX TB2.0
  • DepGPO 22.17% vs DPPO 18.95% on Qwen3.5‑9B TMAX TB2.1
  • DepGPO 37.75% vs DPPO 28.24% on Qwen3.6‑27B SETA TB2.0
  • DepGPO 36.70% vs DPPO 26.67% on Qwen3.6‑27B TMAX TB2.1
Limitations

The paper does not discuss any limitations or applicability beyond terminal tasks.

DepGPO achieves the best pass@1 across all eight settings.Found in the source text, word for word.

Picked because: Introduces a dependency‑aware policy optimization method for terminal‑based agents, offering practical techniques and artifacts for building reliable LLM agents that execute shell commands.

Paper 3 of 5

NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents

Lijie Ding, Changwoo Do · abstract · pdf

quote verified2 figures not in sourceread: full textcs.AI

Problem

Prior to this work there was no executable, non‑arguable environment for neutron instrument design, so LLM agents could only be judged by a language model which can be contested; the obvious fix of using an LLM judge fails because the grading can be argued with.

Approach

NeutronGym provides an executable environment where agents use 22 tools that wrap McStasScript to place components, run simulations and read monitor statistics; McStas compiles each family once and then executes parameter assignments directly, giving a fast 33 ms median runtime. A level‑resolved reward ladder grades syntax, runtime, structure and scientific output without any LLM judge. Reinforcement learning is applied to the reward signal, training Qwen3‑8B to improve its design decisions. The same pipeline is evaluated on held‑out families and on the curated McStasBench benchmark.

Result

RL training raises Qwen3‑8B pass rate from 11.3% to 76.7% on 300 held‑out instances; a second training seed achieves 69.0%. Removing the ladder collapses performance to 16.7%. A hand‑coded physics rule plus search reaches 81.0%, slightly above the RL model, while frontier zero‑shot models achieve 98 to 99% on the first 100 instances.

Why it matters

Researchers building LLM agents for scientific design should care because NeutronGym offers a trainable, physics‑graded benchmark that demonstrates large gains from RL without a disputable LLM judge.

Method details
  • Model Qwen3‑8B (8 B parameters) trained with RL on reward ladder
  • Model Qwen3‑32B (32 B parameters) used as untrained baseline
  • Benchmark McStasBench contains 16 scored tasks with one episode per task
  • Held‑out evaluation uses 300 instances per family
  • Baseline optimizers include coordinate descent, Nelder‑Mead and random search
  • Ablation removing the reward ladder reduces pass rate to 16.7%
Numbers
  • 76.7% pass rate of 300 held‑out instances (trained Qwen3‑8B) vs 11.3% untrained
  • 69.0% pass rate with a second training seed
  • 16.7% pass rate when the reward ladder is removed
  • 81.0% pass rate for hand‑coded physics rule plus search
  • 98 to 99% pass rate for frontier zero‑shot models on first 100 instances
  • collapse by 60 points when ladder is removed
Limitations

The paper does not establish that the RL‑trained model can consistently match frontier zero‑shot performance or that the training recipe is stable across random seeds.

From reward alone the trained model reaches what a classical optimizer reaches, at the agent’s simulation budgetFound in the source text, word for word.
These figures do not appear in the source text: 32 B, 8 B. Treat them as unverified.Number check failed.

Picked because: Provides NeutronGym, an executable environment for neutron instrument design that lets LLM agents be tested and verified in a realistic, self‑hostable simulation pipeline.

Paper 4 of 5

MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

Sean Culatana, Shang-En Huang, Kang Li · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Dense-retrieval services needed separate quantizers for each embedding dimension and code rate, which gave the best quality but required multiple code streams and quantizer states, inflating resident memory.

Approach

MRVQ fits a post‑hoc residual vector quantizer on frozen embeddings. Residual stages create nested rates, while coordinate prefixes create nested dimensions, allowing a single maximum‑rate code to be truncated either by dropping stages or by dropping embedding coordinates. At serving time the stored code is selected without re‑encoding, and queries are scored asymmetrically as in PQ/OPQ. The method thus provides one resident artifact that serves every (dimension, rate) pair evaluated.

Result

MRVQ uses 17.8‑22.0× less memory than three separately trained QINCo2 indices and 1.89‑2.02× less than a lean shared‑model steelman, while per‑rate QINCo2 is only 0.026‑0.107 nDCG@10 better on FiQA. At matched code size MRVQ outperforms PQ, OPQ, and AdANNS‑OPQ.

Why it matters

System engineers who need to support elastic retrieval under tight resident‑memory budgets should consider MRVQ as a low‑memory deployment option.

Method details
  • MRVQ uses an -stage residual vector quantizer with codewords per stage, each stage index occupying one byte.
  • Resident state size is bytes (fp32 codebooks plus one‑byte code per vector) for the evaluated .
  • Evaluated on FiQA (≈? documents, queries) and NFCorpus (≈? documents, queries) from BEIR.
  • Four embedding families: MPNet, Mxbai, Nomic, and BGE.
  • Baselines include per‑rate QINCo2 (separate and shared), PQ, OPQ, AdANNS‑OPQ, RaBitQ, and Extended‑RaBitQ.
  • A low‑build‑cost PCA‑scalar design is also evaluated as an alternative to iterative quantizer training.
Numbers
  • 17.8-22.0x less memory than three separately trained QINCo2 indices
  • 1.89-2.02x less memory than a lean shared-model steelman
  • 0.026-0.107 nDCG@10 better per-rate QINCo2 on FiQA
  • 12.88 MiB MRVQ RAM on FiQA MPNet/Nomic/BGE
  • 228.9 MiB QINCo2 separate RAM on FiQA MPNet/Nomic/BGE
  • 25.17 MiB QINCo2 shared RAM on FiQA MPNet/Nomic/BGE
Limitations

The study covers only two modest BEIR corpora and does not measure end‑to‑end latency, build energy, or recall‑throughput, so it does not establish a universal quality winner.

MRVQ is therefore a low-memory operating point for elastic retrieval, not a universal quality winner.Found in the source text, word for word.

Picked because: Offers MRVQ, a post‑hoc residual vector quantizer supporting dynamic dimension and bitrate changes, with released implementation useful for building scalable, self‑hosted retrieval services.

Paper 5 of 5

Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection

Shuo Yang, Lihao Fang, Yi Zhang and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CV

Problem

Prior caching methods pre‑select a tensor to reuse, but after step distillation the sampling steps are farther apart, so reusing the same tensor across a larger gap introduces more error and degrades quality.

Approach

AutoTarget measures the reuse error of each candidate tensor on a few uncached runs for a given model, solver and reuse schedule, then selects the candidate with the lowest error if it is sufficiently below the runner‑up. It calibrates this ranking before deployment and then reuses only the chosen tensor for new prompts. The method also analyzes how local errors accumulate and identifies equivalent cache targets under Euler sampling. It requires no learned predictor and stores only the selected solver‑facing tensor.

Result

AutoTarget on FLUX.1‑schnell (4 steps) runs in 1.120 s with 27.641 dB PSNR, 0.6743 SSIM, ImageReward 0.886 and CLIP 28.376 while storing only 0.500 MiB of cache. On PixArt‑LCM (4 steps) it runs in 0.124 s with 23.15 dB PSNR, 0.7518 SSIM, ImageReward 0.413 and CLIP 26.070 and retains 0.125 MiB. On HunyuanVideo it runs in 150.4 s achieving overall speedup 7.55×, cache‑only speedup 1.539× and VBench scores 70.9/80.9/78.9.

Why it matters

Deployers of few‑step diffusion transformers can use AutoTarget to pick the optimal cache tensor, gaining latency and memory savings without noticeable quality loss.

Method details
  • Evaluated on FLUX.1‑schnell, PixArt‑LCM and a distilled 20‑step HunyuanVideo Diffusion Transformer
  • Baselines compared: TeaCache, DiCache, TaylorSeer, HiCache, DisCa
  • Quality metrics: ImageReward, CLIP, VBench, PSNR, SSIM with paired references (50‑step FLUX.1‑dev, four‑step PixArt‑LCM)
  • Ablation: Figure 9 compares four cache targets on PixArt‑LCM; calibration set size varied from 1 to 32 prompts
  • Retained‑cache measurement includes stored tensors and any required history
Numbers
  • Overall speedup vs 50‑step FLUX.1‑dev, 16.00, AutoTarget on FLUX.1‑schnell
  • Cache‑only speedup vs uncached four‑step FLUX.1‑schnell, 12.41, AutoTarget on FLUX.1‑schnell
  • PSNR (dB), 27.641, AutoTarget on FLUX.1‑schnell
  • SSIM, 0.6743, AutoTarget on FLUX.1‑schnell
  • Retained cache (MiB), 0.500, AutoTarget on FLUX.1‑schnell
Limitations

Calibration must be redone when model, solver, schedule or precision changes and the ranking may not transfer to other prompt distributions.

AutoTarget reduces DiT evaluations and retained cache storage. Generation quality remains close to the corresponding uncached run.Found in the source text, word for word.

Picked because: Describes a solver‑aware caching strategy for diffusion transformers that cuts inference cost, accompanied by code, directly applicable to infrastructure optimization for generative models.