Adam Fisch, Shubhendu Trivedi, Fantine Huot and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Prior routing systems required exhaustive value estimation for each specialist, which is prohibitively costly, and the naive fix of always using the cheap estimator yields noisy routing while always using the expensive estimator wastes compute.
Approach
The paper formulates routing as a Pandora's Box problem and derives closed‑form value‑of‑information (VoI) criteria under a Gaussian signal model. A centralized policy, Pandora's Router, decides per query whether to refine a specialist's cheap value estimate with the costly estimator. A decentralized variant, Pandora's Bidder, lets specialists independently decide to invest in self‑assessment before accepting a price. The cheap estimator is a KNN over prompt embeddings; the costly estimator is a fine‑tuned small transformer that consumes private information (e.g., reasoning trace or retrieval results). VoI thresholds determine when the extra inspection cost is justified.
Result
Pandora's Router consistently achieves the lowest regret + inspection cost across cost levels, matching the quality of exhaustive estimation while invoking the expensive estimator far less often; for example on the MATH benchmark at inspection cost 1.0e‑02 it attains 0.110 versus 0.117 for the cheap‑only baseline.
Why it matters
Systems that combine heterogeneous models-different sizes, retrieval‑augmented or variable‑length reasoning-can use these VoI‑based routing policies to reduce compute while preserving answer quality.
Method details
Cheap estimator: KNN on pretrained prompt embeddings using cosine similarity.
Costly estimator: fine‑tuned small transformer (e.g., Gemini 3.1 Flash‑Lite) predicting reward from prompt plus private info.
Math domain uses Gemma3‑4B (low‑cost) and Gemini‑3.1‑Flash‑Lite (costly) on 16,512 problems from MATH, Omni‑Math, AIME, HMMT.
RAG domain routes among a no‑retrieval model, a Wikipedia RAG model, and a PubMed RAG model, each incurring a cost.
EmbedLLM benchmark includes >100 open‑weight models with costs proportional to parameter count, evaluated on MMLU and GSM8K.
Training of value estimators required 1‑3 h (up to 6 h for Math) on 64 Google TPUs using up to 436.69 GiB memory.
Numbers
Regret+cost on MATH at cost 1.0e-02, 0.110, lower than f-only 0.117
Regret+cost on RAG at cost 1.0e-04, 0.371, lower than f-only 0.381
Regret+cost on EmbedLLM at cost 1.0e-05, 0.371, lower than f-only 0.393
Limitations
The paper notes that in the decentralized setting noisy competing estimates can increase a specialist's utility at the expense of others, and it focuses mainly on Gaussian signal models despite exploring a KNN alternative.
Pandora’s router matches the routing quality of exhaustive estimation, while querying the expensive estimator far less often.Found in the source text, word for word.
Picked because: Introduces a practical routing system for heterogeneous AI models that balances inference cost and quality, with released code useful for building LLM‑agent pipelines.
Earth observation programmes generate data volumes that exceed the transfer and storage capacity of most institutional networks, making on‑premises access infeasible. Adding more network bandwidth alone does not solve the issue because once a threshold is passed the endpoint memory topology, not the network, becomes the limiting factor.
Approach
The authors build a network‑centric architecture composed of a MinIO object storage cluster deployed on a 100 GbE fabric, a PostGIS metadata catalogue, and an OGC API‑EDR access layer. They characterise the fabric under sustained parallel load and evaluate object‑storage throughput using EO‑representative workloads. Performance is compared against throttled baselines on identical hardware, isolating network bandwidth as the sole variable. Multi‑site replication benchmarks with partner institutions are run to assess the federation primitive. The study identifies the bandwidth threshold where memory topology, rather than network capacity, limits throughput.
Result
The evaluation shows that network bandwidth dominates storage throughput for bulk EO data access up to a threshold, and beyond that threshold endpoint memory topology governs usable bandwidth. For the hardware class tested the threshold lies above 10 Gbps per server, indicating that below this point network capacity alone determines performance.
Why it matters
Institutions that must host EO data on‑premises and network planners should care because the work quantifies when investing in faster network fabric yields diminishing returns without concurrent memory provisioning.
Method details
MinIO object storage cluster deployed on 100 GbE fabric
PostGIS used as metadata catalogue
OGC API‑EDR provides the access layer
Performance compared against throttled baselines on identical hardware
Multi‑site replication benchmarks characterize the federation primitive
Threshold observed above 10 Gbps per server where memory topology limits throughput
Numbers
fabric bandwidth, 100 GbE, network fabric used in deployment
threshold, >10 Gbps per server, point where memory topology becomes limiting
Limitations
The paper does not establish results for hardware classes other than the one evaluated.
Network bandwidth is the dominant constraint on storage throughput for bulk EO data access up to a threshold; beyond it, endpoint memory topology rather than capacity governs how much bandwidth a system can use.Found in the source text, word for word.
Picked because: Presents a fully evaluated on‑premises, network‑centric architecture for large Earth‑observation datasets, offering concrete deployment guidelines for self‑hosted infrastructure.
Prior semantic caches lacked a systematic comparison of eviction policies and geometry-aware policies were expected to help, but under exact lookup with insert-on-miss a new entry cannot have a resident neighbor within the hit radius, so the redundancy signal needed for those policies is missing.
Approach
The authors built the CLEVER framework, which chains an index benchmark, a cost‑based adaptive router, an eviction layer hosting seven interchangeable policies, and an LLM‑based audit of hit quality. They evaluate FIFO, LRU, LFU, ARC, GDSF, a streaming adaptation of SISO, and a semantic‑redundancy policy across three ordered, deduplicated query corpora, three cache capacities, and two embedding encoders. All policies run on the same FAISS HNSW index under exact lookup and insert‑on‑miss semantics. The audit judges answer‑substitutability of hits using an LLM. Results are reported both as raw hit rates and quality‑adjusted rates.
Result
Across eighteen settings no policy beats LFU by more than 0.041 percentage points, while FIFO and streaming SISO fall behind LFU by up to 8.67 and 8.55 points at the smallest capacity. An audit shows only 2.1 to 3.9% of sampled LMSYS and QQP hits are answer‑substitutable, collapsing raw hit rates of 51 to 60% to quality‑adjusted rates of 1.1 to 2.2%. Cross‑encoder replication confirms that thresholds do not transfer between encoders.
Why it matters
LLM serving engineers should adopt LFU as the default eviction policy and first validate answer validity and encoder‑specific thresholds before fine‑tuning sub‑point policy differences.
Datasets: LMSYS‑Chat‑1M first utterance, Quora Question Pairs, MOSS instruction prompts.
Encoders: all‑MiniLM‑L6‑v2 (384‑dim) and gte‑base (768‑dim).
Cache capacities: 10%, 20%, 30% of the total unique query volume.
Index structures benchmarked: FAISS Flat, HNSW, IVF, LSH.
Seeds for eviction randomness: 42, 123, 456; routing uses five seeds.
Numbers
0.041 percentage points, difference vs LFU
8.67 points, FIFO trailing LFU at tight capacity
8.55 points, streaming SISO trailing LFU at tight capacity
2.1 to 3.9%, answer‑substitutable hits
51 to 60%, raw hit rates
1.1 to 2.2%, quality‑adjusted hit rates
Limitations
The study uses ordered, deduplicated corpora rather than real production request traces and relies on insert‑on‑miss semantics, so results may not generalize to workloads with repetition or different admission policies.
No evaluated policy improves on LFU by more than 0.041 percentage points in any of the eighteen settings.Found in the source text, word for word.
Picked because: Provides a systematic, empirical comparison of LLM cache eviction policies across workloads and encoders, giving engineers actionable recommendations for cache management.
quote verified2 figures not in sourceread: full textcs.IR
Problem
The paper reports that about 48% of short‑convolution channels are inert and cannot be removed; attempting to prune them by narrowing the convolution tensors fails because the runtime checks tensor dimensions and rejects the narrowed shape.
Approach
Daedalus‑150M uses an 18‑block stack where six blocks are full attention and twelve are short convolutions with a two‑timestep state, keeping cache size constant regardless of context length. Each block also contains a feed‑forward sub‑layer with inner dimension 2048. The attention blocks use grouped‑query attention (4 KV heads for 12 query heads) to reduce per‑token cost. The model is quantised to 4‑bit weights for CPU inference. The architecture interleaves the blocks in the pattern C C C C A C C A C A C A C A C C A C to balance compute and memory bandwidth.
Result
The hybrid model achieves a five‑task mean of 47.31, clearing the 42.20 bar by 5.11 points and beating all listed baselines. It wins the pre‑registered quality metric by 0.81 % over the dense twin, produces a 6.3 % smaller 4‑bit file, and decodes 1.76× faster at 2048 context (2.08× against an external model). Validation bits‑per‑byte is 0.8685.
Why it matters
Developers building CPU‑only inference for small language models should consider this convolution‑attention hybrid to gain latency and bandwidth efficiency without sacrificing accuracy.
Method details
18‑block stack with pattern C C C C A C C A C A C A C A C C A C (6 attention, 12 convolution)
A dense all‑attention twin with 24 layers and ~161 M parameters was trained as a direct comparison
Quantising to 4‑bit yields a 6.3 % smaller file and adds roughly 6 % perplexity cost
Numbers
Five‑task mean 47.31 vs bar 42.20
Quality metric win 0.81 % over dense
File size reduction 6.3 % compared to dense
Decoding speed 1.76× faster at 2048 context (2.08× vs external)
Validation bits‑per‑byte 0.8685
Inert convolution channels 47.9 % of short‑convolution channels
Limitations
The paper does not demonstrate that the hybrid advantage scales to larger models or other hardware, and it cannot reclaim the inert convolution channels.
The model scores 47.31 on a five‑task benchmark against a bar of 42.20Found in the source text, word for word.
These figures do not appear in the source text: 161 M, 59.9 B. Treat them as unverified.Number check failed.
Picked because: Describes Daedalus‑150M, a CPU‑optimized convolution‑attention hybrid model with released artifacts, enabling efficient on‑device inference for small LLMs.
Fengqing Jiang, Yite Wang, Boyi Liu and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Prior work focused on reasoning or software-engineering tool use and left general agentic tool use underexplored. Simply applying post‑training fine‑tuning does not yield strong multi‑turn tool‑use performance.
Approach
The authors build a MidTool pipeline that mixes large‑scale web, PDF, and code data with synthesized supervision from real‑world tool APIs, MCP skills, and document‑grounded workflows. The mixture, called MidTool‑Mix, is used for mid‑training Qwen3‑4B‑Base and Qwen3‑8B‑Base models. After mid‑training, the models undergo supervised fine‑tuning (SFT) and reinforcement learning (RL). The pipeline normalizes all trajectories into a plain chat‑style template without special control tokens. Evaluation compares the mid‑trained models against baselines that skip the mid‑training stage.
Result
MidTool‑Mix substantially raises BFCL overall scores, e.g., from 39.73% (4B SFT) to 54.18% (4B MidTool+SFT+RL) and from 47.62% (8B SFT) to 55.12% (8B MidTool+SFT+RL). On -Bench, Pass@4 improves from 20.50% to 38.49% for the 4B model and from 28.06% to 39.57% for the 8B model. Gains are especially large on multi‑turn subsets and harder interaction horizons. The improvements persist across both SFT‑only and SFT+RL settings.
Why it matters
Researchers and engineers building agentic LLMs for general tool use should consider dedicated mid‑training with curated tool‑use data to achieve stronger multi‑turn performance.
Method details
Model sizes: Qwen3‑4B‑Base and Qwen3‑8B‑Base.
MidTool‑Mix composition: 20.3B tokens and 11.22M samples.
BFCL overall 4B MidTool+SFT+RL 54.18% vs SFT baseline 39.73%
BFCL overall 8B MidTool+SFT+RL 55.12% vs SFT baseline 47.62%
-Bench Pass@4 overall 4B MidTool+SFT+RL 38.49% vs SFT baseline 20.50%
-Bench Pass@4 overall 8B MidTool+SFT+RL 39.57% vs SFT baseline 28.06%
MidTool‑Mix size 20.3B tokens
MidTool‑Mix sample count 11.22M
Limitations
The paper only provides a pilot study on visual tool use and audits benchmark overlap only at the surface‑level n‑gram level.
Compared with baselines, MidTool-Mix consistently improves downstream performance under both SFT and RL on BFCL, tau2-Bench, and MCP Universe.Found in the source text, word for word.
Picked because: Shows how targeted mid‑training data synthesis can improve LLM tool‑use capabilities, delivering a reproducible method for enhancing agentic functionality.