arXiv digest

Friday

August 21, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

Adam Fisch, Shubhendu Trivedi, Fantine Huot and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Prior routing systems required exhaustive value estimation for each specialist, which is prohibitively costly, and the naive fix of always using the cheap estimator yields noisy routing while always using the expensive estimator wastes compute.

Approach

The paper formulates routing as a Pandora's Box problem and derives closed‑form value‑of‑information (VoI) criteria under a Gaussian signal model. A centralized policy, Pandora's Router, decides per query whether to refine a specialist's cheap value estimate with the costly estimator. A decentralized variant, Pandora's Bidder, lets specialists independently decide to invest in self‑assessment before accepting a price. The cheap estimator is a KNN over prompt embeddings; the costly estimator is a fine‑tuned small transformer that consumes private information (e.g., reasoning trace or retrieval results). VoI thresholds determine when the extra inspection cost is justified.

Result

Pandora's Router consistently achieves the lowest regret + inspection cost across cost levels, matching the quality of exhaustive estimation while invoking the expensive estimator far less often; for example on the MATH benchmark at inspection cost 1.0e‑02 it attains 0.110 versus 0.117 for the cheap‑only baseline.

Why it matters

Systems that combine heterogeneous models-different sizes, retrieval‑augmented or variable‑length reasoning-can use these VoI‑based routing policies to reduce compute while preserving answer quality.

Method details
  • Cheap estimator: KNN on pretrained prompt embeddings using cosine similarity.
  • Costly estimator: fine‑tuned small transformer (e.g., Gemini 3.1 Flash‑Lite) predicting reward from prompt plus private info.
  • Math domain uses Gemma3‑4B (low‑cost) and Gemini‑3.1‑Flash‑Lite (costly) on 16,512 problems from MATH, Omni‑Math, AIME, HMMT.
  • RAG domain routes among a no‑retrieval model, a Wikipedia RAG model, and a PubMed RAG model, each incurring a cost.
  • EmbedLLM benchmark includes >100 open‑weight models with costs proportional to parameter count, evaluated on MMLU and GSM8K.
  • Training of value estimators required 1‑3 h (up to 6 h for Math) on 64 Google TPUs using up to 436.69 GiB memory.
Numbers
  • Regret+cost on MATH at cost 1.0e-02, 0.110, lower than f-only 0.117
  • Regret+cost on RAG at cost 1.0e-04, 0.371, lower than f-only 0.381
  • Regret+cost on EmbedLLM at cost 1.0e-05, 0.371, lower than f-only 0.393
Limitations

The paper notes that in the decentralized setting noisy competing estimates can increase a specialist's utility at the expense of others, and it focuses mainly on Gaussian signal models despite exploring a KNN alternative.

Pandora’s router matches the routing quality of exhaustive estimation, while querying the expensive estimator far less often.Found in the source text, word for word.

Picked because: Introduces a practical routing system for heterogeneous AI models that balances inference cost and quality, with released code useful for building LLM‑agent pipelines.

Paper 2 of 5

Design and Empirical Evaluation of a Network-Centric, On-Premises Architecture for Earth Observation Data Access

João Pinelo, João Gonçalves, Denis Willett and 7 others · abstract · pdf

quote verifiedfigures checkedread: abstract onlycs.DC

Problem

Earth observation programmes generate data volumes that exceed the transfer and storage capacity of most institutional networks, making on‑premises access infeasible. Adding more network bandwidth alone does not solve the issue because once a threshold is passed the endpoint memory topology, not the network, becomes the limiting factor.

Approach

The authors build a network‑centric architecture composed of a MinIO object storage cluster deployed on a 100 GbE fabric, a PostGIS metadata catalogue, and an OGC API‑EDR access layer. They characterise the fabric under sustained parallel load and evaluate object‑storage throughput using EO‑representative workloads. Performance is compared against throttled baselines on identical hardware, isolating network bandwidth as the sole variable. Multi‑site replication benchmarks with partner institutions are run to assess the federation primitive. The study identifies the bandwidth threshold where memory topology, rather than network capacity, limits throughput.

Result

The evaluation shows that network bandwidth dominates storage throughput for bulk EO data access up to a threshold, and beyond that threshold endpoint memory topology governs usable bandwidth. For the hardware class tested the threshold lies above 10 Gbps per server, indicating that below this point network capacity alone determines performance.

Why it matters

Institutions that must host EO data on‑premises and network planners should care because the work quantifies when investing in faster network fabric yields diminishing returns without concurrent memory provisioning.

Method details
  • MinIO object storage cluster deployed on 100 GbE fabric
  • PostGIS used as metadata catalogue
  • OGC API‑EDR provides the access layer
  • Performance compared against throttled baselines on identical hardware
  • Multi‑site replication benchmarks characterize the federation primitive
  • Threshold observed above 10 Gbps per server where memory topology limits throughput
Numbers
  • fabric bandwidth, 100 GbE, network fabric used in deployment
  • threshold, >10 Gbps per server, point where memory topology becomes limiting
Limitations

The paper does not establish results for hardware classes other than the one evaluated.

Network bandwidth is the dominant constraint on storage throughput for bulk EO data access up to a threshold; beyond it, endpoint memory topology rather than capacity governs how much bandwidth a system can use.Found in the source text, word for word.

Picked because: Presents a fully evaluated on‑premises, network‑centric architecture for large Earth‑observation datasets, offering concrete deployment guidelines for self‑hosted infrastructure.

Paper 3 of 5

Which Eviction Policy Should an LLM Cache Use? A Systematic Study Across Workloads, Capacities, and Encoders

Yash Kulkarni, Shubham Harkare, Arvind Suresh Yogesh Babu · abstract · pdf

quote verifiedfigures checkedread: full textcs.DB

Problem

Prior semantic caches lacked a systematic comparison of eviction policies and geometry-aware policies were expected to help, but under exact lookup with insert-on-miss a new entry cannot have a resident neighbor within the hit radius, so the redundancy signal needed for those policies is missing.

Approach

The authors built the CLEVER framework, which chains an index benchmark, a cost‑based adaptive router, an eviction layer hosting seven interchangeable policies, and an LLM‑based audit of hit quality. They evaluate FIFO, LRU, LFU, ARC, GDSF, a streaming adaptation of SISO, and a semantic‑redundancy policy across three ordered, deduplicated query corpora, three cache capacities, and two embedding encoders. All policies run on the same FAISS HNSW index under exact lookup and insert‑on‑miss semantics. The audit judges answer‑substitutability of hits using an LLM. Results are reported both as raw hit rates and quality‑adjusted rates.

Result

Across eighteen settings no policy beats LFU by more than 0.041 percentage points, while FIFO and streaming SISO fall behind LFU by up to 8.67 and 8.55 points at the smallest capacity. An audit shows only 2.1 to 3.9% of sampled LMSYS and QQP hits are answer‑substitutable, collapsing raw hit rates of 51 to 60% to quality‑adjusted rates of 1.1 to 2.2%. Cross‑encoder replication confirms that thresholds do not transfer between encoders.

Why it matters

LLM serving engineers should adopt LFU as the default eviction policy and first validate answer validity and encoder‑specific thresholds before fine‑tuning sub‑point policy differences.

Method details
  • Policies evaluated: FIFO, LRU, LFU, ARC, GDSF, streaming SISO, semantic redundancy.
  • Datasets: LMSYS‑Chat‑1M first utterance, Quora Question Pairs, MOSS instruction prompts.
  • Encoders: all‑MiniLM‑L6‑v2 (384‑dim) and gte‑base (768‑dim).
  • Cache capacities: 10%, 20%, 30% of the total unique query volume.
  • Index structures benchmarked: FAISS Flat, HNSW, IVF, LSH.
  • Seeds for eviction randomness: 42, 123, 456; routing uses five seeds.
Numbers
  • 0.041 percentage points, difference vs LFU
  • 8.67 points, FIFO trailing LFU at tight capacity
  • 8.55 points, streaming SISO trailing LFU at tight capacity
  • 2.1 to 3.9%, answer‑substitutable hits
  • 51 to 60%, raw hit rates
  • 1.1 to 2.2%, quality‑adjusted hit rates
Limitations

The study uses ordered, deduplicated corpora rather than real production request traces and relies on insert‑on‑miss semantics, so results may not generalize to workloads with repetition or different admission policies.

No evaluated policy improves on LFU by more than 0.041 percentage points in any of the eighteen settings.Found in the source text, word for word.

Picked because: Provides a systematic, empirical comparison of LLM cache eviction policies across workloads and encoders, giving engineers actionable recommendations for cache management.

Paper 4 of 5

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference

Christos Koutsiaris · abstract · pdf

quote verified2 figures not in sourceread: full textcs.IR

Problem

The paper reports that about 48% of short‑convolution channels are inert and cannot be removed; attempting to prune them by narrowing the convolution tensors fails because the runtime checks tensor dimensions and rejects the narrowed shape.

Approach

Daedalus‑150M uses an 18‑block stack where six blocks are full attention and twelve are short convolutions with a two‑timestep state, keeping cache size constant regardless of context length. Each block also contains a feed‑forward sub‑layer with inner dimension 2048. The attention blocks use grouped‑query attention (4 KV heads for 12 query heads) to reduce per‑token cost. The model is quantised to 4‑bit weights for CPU inference. The architecture interleaves the blocks in the pattern C C C C A C C A C A C A C A C C A C to balance compute and memory bandwidth.

Result

The hybrid model achieves a five‑task mean of 47.31, clearing the 42.20 bar by 5.11 points and beating all listed baselines. It wins the pre‑registered quality metric by 0.81 % over the dense twin, produces a 6.3 % smaller 4‑bit file, and decodes 1.76× faster at 2048 context (2.08× against an external model). Validation bits‑per‑byte is 0.8685.

Why it matters

Developers building CPU‑only inference for small language models should consider this convolution‑attention hybrid to gain latency and bandwidth efficiency without sacrificing accuracy.

Method details
  • 18‑block stack with pattern C C C C A C C A C A C A C A C C A C (6 attention, 12 convolution)
  • Feed‑forward inner dimension 2048, vocabulary 49,152, context length 2048
  • Trained from scratch on 59.9 B tokens using Muon for matrix weights and AdamW for embeddings, bf16 precision on a single RTX 5090
  • Baseline models: GPT‑2 124M, Pythia‑160M, OPT‑125M, GPT‑neo‑125M, MobileLLM‑125M
  • A dense all‑attention twin with 24 layers and ~161 M parameters was trained as a direct comparison
  • Quantising to 4‑bit yields a 6.3 % smaller file and adds roughly 6 % perplexity cost
Numbers
  • Five‑task mean 47.31 vs bar 42.20
  • Quality metric win 0.81 % over dense
  • File size reduction 6.3 % compared to dense
  • Decoding speed 1.76× faster at 2048 context (2.08× vs external)
  • Validation bits‑per‑byte 0.8685
  • Inert convolution channels 47.9 % of short‑convolution channels
Limitations

The paper does not demonstrate that the hybrid advantage scales to larger models or other hardware, and it cannot reclaim the inert convolution channels.

The model scores 47.31 on a five‑task benchmark against a bar of 42.20Found in the source text, word for word.
These figures do not appear in the source text: 161 M, 59.9 B. Treat them as unverified.Number check failed.

Picked because: Describes Daedalus‑150M, a CPU‑optimized convolution‑attention hybrid model with released artifacts, enabling efficient on‑device inference for small LLMs.

Paper 5 of 5

MidTool: Mid-training Data Synthesis for Agentic Tool Use

Fengqing Jiang, Yite Wang, Boyi Liu and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Prior work focused on reasoning or software-engineering tool use and left general agentic tool use underexplored. Simply applying post‑training fine‑tuning does not yield strong multi‑turn tool‑use performance.

Approach

The authors build a MidTool pipeline that mixes large‑scale web, PDF, and code data with synthesized supervision from real‑world tool APIs, MCP skills, and document‑grounded workflows. The mixture, called MidTool‑Mix, is used for mid‑training Qwen3‑4B‑Base and Qwen3‑8B‑Base models. After mid‑training, the models undergo supervised fine‑tuning (SFT) and reinforcement learning (RL). The pipeline normalizes all trajectories into a plain chat‑style template without special control tokens. Evaluation compares the mid‑trained models against baselines that skip the mid‑training stage.

Result

MidTool‑Mix substantially raises BFCL overall scores, e.g., from 39.73% (4B SFT) to 54.18% (4B MidTool+SFT+RL) and from 47.62% (8B SFT) to 55.12% (8B MidTool+SFT+RL). On -Bench, Pass@4 improves from 20.50% to 38.49% for the 4B model and from 28.06% to 39.57% for the 8B model. Gains are especially large on multi‑turn subsets and harder interaction horizons. The improvements persist across both SFT‑only and SFT+RL settings.

Why it matters

Researchers and engineers building agentic LLMs for general tool use should consider dedicated mid‑training with curated tool‑use data to achieve stronger multi‑turn performance.

Method details
  • Model sizes: Qwen3‑4B‑Base and Qwen3‑8B‑Base.
  • MidTool‑Mix composition: 20.3B tokens and 11.22M samples.
  • Source balance: 42% web, 26% code, 23% PDF, 9% native agentic trajectories.
  • Training stages: mid‑training on MidTool‑Mix, then SFT, then RL.
  • Baselines: no mid‑training and Dolmino‑20BT generic mid‑training mixture.
  • Ablations: raw sources only, context‑grounded augmentation only, native agentic trajectories only.
Numbers
  • BFCL overall 4B MidTool+SFT+RL 54.18% vs SFT baseline 39.73%
  • BFCL overall 8B MidTool+SFT+RL 55.12% vs SFT baseline 47.62%
  • -Bench Pass@4 overall 4B MidTool+SFT+RL 38.49% vs SFT baseline 20.50%
  • -Bench Pass@4 overall 8B MidTool+SFT+RL 39.57% vs SFT baseline 28.06%
  • MidTool‑Mix size 20.3B tokens
  • MidTool‑Mix sample count 11.22M
Limitations

The paper only provides a pilot study on visual tool use and audits benchmark overlap only at the surface‑level n‑gram level.

Compared with baselines, MidTool-Mix consistently improves downstream performance under both SFT and RL on BFCL, tau2-Bench, and MCP Universe.Found in the source text, word for word.

Picked because: Shows how targeted mid‑training data synthesis can improve LLM tool‑use capabilities, delivering a reproducible method for enhancing agentic functionality.