Yujie Zhang, Huiying Lan, Ehsan Aghapour and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.DC
Problem
Traditional pipelining across heterogeneous SoC units improves throughput but fails to meet the latency demands of modern neural networks with extensive operator parallelism. Conversely, pure operator parallel execution reduces latency but harms throughput, forcing a trade‑off that prior methods cannot resolve.
Approach
Para‑Pipe introduces a hierarchical mapping framework that combines intra‑stage and inter‑stage operator parallelism within a pipelined architecture. It uses a Graph Partitioner to split the model into subgraphs, a Pipeline Configuration Generator to define stage assignments, and coarse‑grained and fine‑grained ILP‑based mapping passes to allocate operators to processors. The framework fine‑tunes parallelism levels within each stage and across stages, reducing inter‑processor communication overhead and improving energy efficiency.
Result
On the Amlogic SoC, throughput‑optimized Para‑Pipe configurations achieve an average energy‑efficiency improvement of 11.0% over purely pipelined strategies and 23.3% over non‑pipelined parallel execution. ILP‑based fine‑grained mapping solves most subgraphs within about 5 minutes, while the largest PETR subgraph requires roughly 6 hours.
Why it matters
Edge AI engineers targeting heterogeneous SoCs should consider hierarchical operator parallelism to achieve balanced latency, throughput, and energy efficiency for complex models.
Amlogic SoC: ARM big.LITTLE (quad‑core Cortex‑A73 + dual‑core Cortex‑A53) and ARM G52 MP4 GPU.
BST SoC: deep‑learning accelerator (NPU) and two DSPs.
Inference setup: stream of 50 frames per mapping strategy, measuring average frames‑per‑second throughput, per‑frame latency, active power via USB power meter.
Baselines: Layer‑switched algorithm, HEFT & CPOP DAG mapping algorithms.
energy efficiency improvement, 11.0%, over purely pipelined strategies
energy efficiency improvement, 23.3%, relative to non‑pipelined parallel execution
ILP solution time, ~5 minutes, for most subgraphs
ILP solution time, ~6 hours, for the largest PETR subgraph
subgraphs per model, 11‑45, across six benchmarks
Limitations
The evaluation on the BST SoC relies on simulations rather than real hardware, and the study focuses only on inference, not training workloads.
throughput-optimized configurations under Para-Pipe on Amlogic SoC show an average energy efficiency improvement of 11.0% over purely pipelined strategies and 23.3% relative to non-pipelined parallel execution.Found in the source text, word for word.
Picked because: Provides a released framework for exploiting hierarchical operator parallelism on heterogeneous SoCs, giving engineers concrete tools to accelerate ML pipelines in production stacks.
Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild · abstract · pdf
quote unverified1 figure not in sourceread: full textcs.CR
Problem
LLM agents cannot hold a multi‑thousand‑host authentication graph in their context window and free‑form generation cannot guarantee containment actions respect topology, making them unreliable at enterprise scale.
Approach
Sentinel‑RL separates topological from semantic reasoning. A heterogeneous graph attention encoder embeds the live authentication subgraph into a fixed‑dimensional state. A Proximal Policy Optimization (PPO) policy maps this state to a constrained set of investigative actions. An LLM agent consumes the policy recommendations and generates analyst‑readable narratives, gated by an internal critic. The architecture is organized into four planes (data, strategic, telemetry, orchestration) to keep components modular.
Result
Ingestion of a 24 M‑edge subgraph completes in 14.2 minutes, a sliding‑window alert triggers in ≤2.5 seconds across 50 trials, PPO converges to a mean episodic return of 8.74 with held‑out precision 0.91 and recall 0.87, and the end‑to‑end containment loop has a median latency of 6.3 seconds.
Why it matters
Security operations teams seeking scalable, graph‑aware automation should consider Sentinel‑RL because it delivers fast, precise detection while offloading heavy topological reasoning from LLMs.
Method details
HetGAT graph encoder processes two‑hop neighborhoods of the flagged host
PPO policy trained for 200 iterations on a custom MDP
Training uses five independent random seeds
Dataset: LANL Comprehensive Multi‑Source Cyber‑Security Events and Indiana University Quartz HPC cluster
Baseline comparisons include LMDetect, Bowman et al., Euler, PIKACHU
Inference runs on a single 32‑core node with 128 GB RAM, no GPU
Numbers
ingestion time 14.2 minutes vs canonical MERGE pipeline (≈24× faster)
alert latency ≤2.5 seconds across 50 trials
mean episodic return 8.74 after 200 iterations
precision 0.91 ±0.02 on held‑out LANL red‑team events
recall 0.87 ±0.03 on held‑out LANL red‑team events
median end‑to‑end cycle 6.3 seconds
Limitations
The paper does not demonstrate performance on GPU‑accelerated hardware or on datasets beyond LANL and the HPC cluster.
The model could not quote the source for its claim, so nothing above has been checked against the paper. Read the abstract before trusting it.Citation check failed.
These figures do not appear in the source text: 24 M. Treat them as unverified.Number check failed.
Picked because: Introduces SENTINEL‑RL, an LLM‑agent system that offloads topological reasoning for security operations, offering a practical, verifiable agent tooling artifact.
Gustavo Claudio Karl Couto, Eric Aislan Antonelo, Gabriel George Zipperer · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
Prior end-to-end policies trained in simulation performed poorly on the miniature Ackermann vehicle because the real system suffers from limited camera field of view and a visual appearance gap. Simply widening the simulated camera or using only real data does not solve these issues.
Approach
The authors build a low‑cost platform that includes a physical Ackermann vehicle, a printed urban track, a data‑collection pipeline, a map‑registration tool, and a Webots digital twin. They implement a command‑conditioned behavior‑cloning policy that receives an on‑board image and a high‑level navigation command and outputs steering and speed. To bridge the sim‑to‑real gap they generate synthetic driving data in the twin and translate its images with a four‑level U‑Net. A higher‑capacity CNN variant is trained on the mixed synthetic‑plus‑real dataset, while a compact CNN baseline is trained on real data only. The system is evaluated both on the real vehicle and in the digital twin, with ablations on camera field of view and network capacity.
Result
On the physical track the compact policy achieved a mean cross‑track error of 6.1 cm, close to the 4.7 cm error of human demonstrations. In the digital twin, widening the camera field of view reduced the mean cross‑track error from 35.6 cm to 3.3 cm. The higher‑capacity policy trained on mixed synthetic‑plus‑real data was the only configuration that completed all four track routes in closed loop.
Why it matters
Researchers working on sim‑to‑real autonomous driving and low‑cost robotics should care because the platform provides a reproducible testbed and demonstrates that camera field of view and synthetic data augmentation are critical for transferring policies to real vehicles.
Method details
Compact CNN baseline: convolutional encoder with global average pooling, one‑hot command concatenated, decoded by a two‑layer MLP.
Larger CNN variant: spatial encoder, flattened features concatenated with command, decoded by a three‑layer MLP (about M parameters).
Real expert dataset collected by human teleoperation across four route families.
Synthetic dataset generated in Webots twin with actuator noise and off‑route perturbations.
Sim‑to‑real image translator: four‑level U‑Net trained on paired simulator-real images with photometric domain randomization.
Training uses behavior cloning with weighted loss, AdamW optimizer, and checkpoint selection by closed‑loop rollout in the digital twin.
Numbers
mean cross‑track error, 6.1 cm, real vehicle vs human 4.7 cm
mean cross‑track error, 35.6 cm, narrow FOV (58°) vs widened FOV (120°) 3.3 cm
camera field of view, 58°, narrow lens
camera field of view, 120°, widened lens
Limitations
The paper reports a small number of physical runs, manual trajectory registration, and preliminary ablations that lack a compact‑network/mixed‑data condition and confound capacity with input resolution.
reaching a mean cross-track error of 6.1 cm with respect to the reference route, close to the 4.7 cm observed in human demonstrations.Found in the source text, word for word.
Picked because: Offers an open, low‑cost hardware and software platform for end‑to‑end autonomous driving, enabling self‑hosted experimentation and full‑stack integration.
Zhiyuan Fan, Tinghao Yu, Yuanjun Cai and 9 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Existing synthesized environments become insufficiently challenging for frontier models, providing limited learning signals, and prior co‑evolution methods rely on on‑policy rollouts which restrict generalization and continuous signal provision as models improve.
Approach
The paper introduces environment evolution, which incrementally raises environment difficulty off‑policy and schedules generations during training. It derives three difficulty directions-scenario novelty, skill rarity, and execution length-from the multi‑turn learning objective. A loop‑engineered multi‑agent harness implements evolution via Plan Refinement and Environment Refinement loops. An Evolution‑Lineage Scheduler advances generations only after a pass‑rate threshold is met, ensuring informative rollout budgets. This pipeline replaces on‑policy dependence with off‑policy difficulty signals.
Result
Quantitative rollout experiments with Hy4 preview, Claude Opus 5, and GPT‑5.6 Sol show that environment evolution consistently yields more difficult environments, and long‑horizon RL training on Qwen3.6‑27B and Qwen3.6‑35B‑A3B improves Terminal‑Bench 2.1 performance by 14.4 and 18.0 percentage points respectively.
Why it matters
Researchers building terminal agents and RL pipelines should care because environment evolution provides continuous off‑policy difficulty signals and demonstrably boosts benchmark performance without relying on on‑policy rollouts.
Method details
Qwen3.6-27B is a 27B‑parameter dense model used for RL training
Qwen3.6-35B-A3B is a mixture‑of‑experts model with 35B total parameters and 3B activated
Training uses GRPO with a staleness bound of 5 for 200 steps
Evaluation uses the Claude Code harness with a 256K context window and auto‑compact at 16K
Benchmark is Terminal‑Bench 2.1 Verified, reporting averages over five runs
Rollout difficulty is estimated over 8 independent rollouts using Claude Opus 5, GPT‑5.6 Sol, and Hy4 preview
Numbers
Performance improvement 14.4 percentage points vs baseline for Qwen3.6‑27B
Performance improvement 18.0 percentage points vs baseline for Qwen3.6‑35B‑A3B
Tokens per turn increase from approximately 951 to 1,221 for Qwen3.6‑27B
Tokens per turn increase from approximately 947 to 1,103 for Qwen3.6‑35B‑A3B
Seed pool contains 500 environments selected after filtering
Rollout difficulty estimated over 8 independent rollouts
Limitations
The paper does not establish effectiveness beyond terminal agents or for software‑engineering tasks, noting future work is needed in those areas.
environment evolution improves Terminal-Bench 2.1 performance by 14.4 and 18.0 percentage points, respectively.Found in the source text, word for word.
Picked because: Presents a method for automatically evolving terminal environments for agents, delivering automation scripts that can be integrated into DevOps pipelines for scalable training data generation.
Repeated‑query audits of LLM brand recommendations lacked a standardized protocol for choosing iteration counts, stability metrics, and reliability thresholds. Simply increasing the number of prompt repetitions does not address the conditional, non‑Gaussian dependencies of autoregressive generation.
Approach
The Dice Roll Method formalizes repeated‑query auditing by decomposing total response variance into sampling, prompt‑phrasing, run‑to‑run, and model‑version components. It fits a negative‑binomial generalized linear mixed model (GLMM) with iterations as repeated measures nested within prompt, model, and language. Effect size is quantified with Cliff’s delta and confidence intervals via bias‑corrected accelerated bootstrap. Dependence‑preserving moving‑block bootstrap and simulation‑based power analysis (simr) are used, and a generalizability‑theory D‑study provides iteration‑count guidance. Drift diagnostics (Kolmogorov-Smirnov, Population Stability Index) are applied to pinned model snapshots.
Result
The D‑study yields three iteration tiers: exploratory (n = 5, G = 0.58), confirmatory (n = 10, G = 0.74), and rigorous (n = 15, G = 0.81). An external validation on three independent corpora reproduced the D‑study reliability prediction in 37 of 39 cells with no failures, while the fixed tiers did not transfer, supporting a pilot‑then‑solve approach.
Why it matters
Researchers and reviewers auditing LLM brand recommendations should adopt the Dice Roll Method to obtain statistically principled reliability estimates and appropriate iteration counts.
Datasets: five source brand‑recommendation studies, ~190,000 observations, 270+ brands, 6 languages, iteration counts 5 to 40.
Inference setup: temperature 0.3, max tokens 1,024, system prompt "You are a helpful assistant with broad knowledge of businesses and technology."
Baseline comparison: independent‑samples t‑test retained only as a reference point.
Statistical stack: negative‑binomial GLMM, Cliff’s delta, moving‑block and parametric bootstrap, simr power simulation, generalizability‑theory decomposition.
Numbers
exploratory tier: n = 5, G = 0.58
confirmatory tier: n = 10, G = 0.74
rigorous tier: n = 15, G = 0.81
external validation reproduced reliability in 37 of 39 cells
small effect (Cliff’s δ = 0.15) power = 0.78 at 10 iterations
medium effect (Cliff’s δ = 0.33) power = 0.97 at 5 iterations
Limitations
The fixed iteration tiers do not transfer across all contexts, so the protocol requires a pilot‑then‑solve reading rather than universal tier application.
Three tiers of iteration guidance emerge from the D-study: exploratory (n = 5, G = 0.58), confirmatory (n = 10, G = 0.74), and rigorous (n = 15, G = 0.81).Found in the source text, word for word.
Picked because: Defines the Dice Roll Method, a standardized, reproducible protocol with tooling for repeated‑query LLM auditing, directly applicable to verification workflows.