Nguyen Phuc Tran, Brigitte Jaumard, Oscar Delgado · abstract · pdf
quote verifiedfigures checkedread: full textcs.IT
Problem
Previous O-RAN edge management either placed cloud-native functions without considering migration energy or forced each DU to use a single CU-UP, limiting placement flexibility. Adding more CU-UPs without joint optimization does not reduce energy because migration and wake-up costs are ignored.
Approach
The paper formulates the Energy-Aware Joint Placement and Migration (EJPM) problem as a mixed-integer linear program (MILP) solved per control interval. To achieve scalability it proposes a deterministic k-means-based heuristic (H_EJPM) that runs once per interval and consists of four phases: logical pairing of slice‑flow groups, delay‑aware physical mapping to servers, energy‑efficient migration and consolidation, and post‑processing feasibility repair. Phase‑1 clusters slice‑flow groups with k‑means++ and matches clusters to CU‑UPs via the Hungarian algorithm. Phase‑2 maps the logical assignment onto physical servers respecting delay constraints. Phase‑3 evaluates migration moves and consolidates under a strict energy‑improvement threshold. Phase‑4 repairs any remaining feasibility violations.
Result
Joint Multi-CU placement reduces modeled energy consumption by 5.7% compared to the Single-CU baseline, and the heuristic achieves results within approximately 9.7% of the MILP optimum. The energy advantage is larger during peak traffic (2.11 kWh or 8.3%) than during low traffic (0.99 kWh or 4.39%). Joint Multi-CU also triggers fewer migrations than Joint Single-CU.
Why it matters
Network operators deploying O-RAN edge clouds can lower energy consumption by allowing slice‑aware multi‑CU‑UP placement and by using the proposed heuristic for fast near‑optimal decisions.
Method details
MILP formulation (EJPM) solved per control interval using Gurobi Optimizer.
Deterministic heuristic H_EJPM with four phases: logical pairing, delay-aware physical mapping, migration & consolidation, feasibility repair.
Phase‑1 uses k‑means++ clustering with a fixed number of restarts, iterations, and a pseudo‑random seed.
Hungarian algorithm matches clusters to eligible CU‑UPs.
Experiments use a 24‑hour deterministic Montreal traffic trace generated following [5].
Numbers
energy reduction, 5.7%, Multi-CU vs Single-CU baseline
heuristic gap, 9.7%, heuristic vs MILP
low‑traffic hourly energy difference, 0.99 kWh (4.39%), Multi-CU vs Single-CU
peak‑traffic hourly energy difference, 2.11 kWh (8.3%), Multi-CU vs Single-CU
joint Multi-CU consumes 5.7% less energy than joint Single-CU
Limitations
The model does not consider RU‑DU fronthaul latency, UE radio‑access delay, core transport, or stochastic demand; it only optimizes a deterministic per‑interval snapshot.
the theoretical Multi-CU relaxation reduces modeled energy consumption by 5.7% relative to the Single-CU baseline.Found in the source text, word for word.
Picked because: Presents concrete algorithms and evaluation for energy‑aware placement and migration of cloud‑native functions in O‑RAN edge clouds, with released code, directly applicable to DevOps and infrastructure automation.
Fei Tang, Huawen Shen, Zhiqiong Lu and 7 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Existing web-agent datasets contain only a few thousand trajectories from a fixed, narrow set of sites, limiting agent exposure. The obvious fix of simply gathering more data fails because prior synthesis pipelines remain bound to predefined site lists and cannot scale to the open web.
Approach
BrowserForge generates web interaction data at scale by running many browser sandboxes in parallel over the open web. It consists of three components: an open‑web sourcing stage that discovers hundreds of thousands of real websites, a sandbox cluster manager that schedules hundreds of concurrent browsers, and a Proposer‑Solver dual‑agent loop that creates executable tasks and collects verified trajectories. After collection, a rule‑plus‑model cleaning pipeline filters failed runs and rewrites reasoning into a unified chain‑of‑thought format. The final corpus is used to fine‑tune a compact multimodal model that operates solely from screenshots.
Result
Fine‑tuning on the BrowserForge corpus raises Online‑Mind2Web success rate from 25.66% to 33.33% for the 4B model and from 29.3% to 38.0% for the 9B model, and improves Multimodal‑Mind2Web step accuracy (Pass@1 from 36.45% to 44.88% and Pass@4 from 45.41% to 54.74%). Controlled analyses show the gain stems from the open‑web data and the cleaning pipeline.
Why it matters
Researchers and engineers building web agents should care because BrowserForge shows that large, diverse open‑web data can substantially improve agent performance without changing model architecture or training recipes.
Method details
Model sizes: Qwen3.5‑4B and Qwen3.5‑9B fine‑tuned as BrowserForge‑4B and BrowserForge‑9B.
Dataset: 203,238 trajectories, each from a distinct website.
Training setup: freeze visual encoder and adapter layers, full‑parameter LM fine‑tuning on the BrowserForge corpus using LLaMA‑Factory, with a maximum of pixels per image.
Ablations: cleaning pipeline components (rule filter, model judge, unified CoT) and data‑source comparison (open‑source trajectories vs BrowserForge).
Numbers
Online‑Mind2Web SR 33.3% (BrowserForge‑4B) vs 25.7% baseline (+7.6)
Online‑Mind2Web SR 38.0% (BrowserForge‑9B) vs 29.3% baseline (+9.3)
Pass@1 average step accuracy 44.88% (BrowserForge) vs 36.45% (open‑source trajectories)
Pass@4 average step accuracy 54.74% (BrowserForge) vs 45.41% (open‑source trajectories)
Corpus size 203,238 trajectories from distinct websites
Training trajectories used per model 20K
Limitations
The paper evaluates only on Online‑Mind2Web and Multimodal‑Mind2Web and does not demonstrate generalization to other web‑interaction tasks or domains.
Fine-tuning a compact multimodal model on this corpus raises its success rate on the live Online-Mind2Web from 25.66% to 33.33%.Found in the source text, word for word.
Picked because: Introduces BrowserForge, an open‑source system that parallelizes browser sandboxes to generate large‑scale web interaction data, offering a practical stack for building and scaling web‑based LLM agents.
Esakkivel Esakkiraja, Denis Akhiyarov, Vikas Yadav and 4 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Prior agents using the default Stirrup harness suffered large failures due to mismatches in tool interfaces, environment conventions, and missing operational knowledge, and simply optimizing prompts (e.g., GEPA) does not fix these because the harness architecture and tool configurations remain unchanged.
Approach
StarHarness keeps the model weights fixed and evolves the surrounding harness by stratifying tasks according to baseline failure behavior, separating proposer-visible search tasks from hidden selection tasks, and reserving held‑out tasks for generalization evaluation. It runs an exploration stage (tree‑search) followed by a hill‑climbing stage to propose patches to harness components such as prompts, tool interfaces, skill modules, MCP providers, sub‑agent structure, and loop configuration. Accepted patches are applied to create an environment‑specific harness which is then evaluated on the full benchmark. The method leverages a compact evolution pool and evaluates transfer without re‑evolution across model families.
Result
StarHarness improves full‑benchmark scores by 20‑35 percentage points over the default harness after 4‑12 accepted changes per environment, outperforms GEPA prompt optimization by 13.8‑22.3 points, and reduces estimated inference cost per task by up to 53 %. The evolved harness also transfers to other GPT and Qwen models, yielding large gains without re‑evolution.
Why it matters
Developers of tool‑rich enterprise AI agents should care because they can achieve large performance gains and cost reductions by evolving the harness instead of scaling model size.
Method details
Evolution uses GPT‑5.4 (medium reasoning) as both the agent under test and the proposer
Benchmarks are ITBench SRE (40 Kubernetes root‑cause scenarios), EnterpriseOps‑Gym ITSM (103 tasks), and AutomationBench Finance (100 tasks)
Baseline is the unmodified Stirrup agent framework; comparisons also include Pi, Codex, and GEPA prompt optimization on Pi
Transfer experiments evaluate the frozen evolved harness on GPT‑5.4‑mini, GPT‑5.4, GPT‑5.5, Qwen3.5‑27B, and Qwen3.6‑27B models
A total of 21 patches were accepted across the three environments (4 ITBench, 12 EnterpriseOps‑Gym, 5 AutomationBench)
Relative to GEPA (Pi): +13.8 pp on ITBench, +22.3 pp on EnterpriseOps‑Gym, +17.6 pp on AutomationBench
Inference cost reduction: 17% on ITBench, 53% on EnterpriseOps‑Gym, 29% on AutomationBench
Qwen3.5‑27B transfer on ITBench: 25.6% baseline → 70.0% evolved (+44.4 pp)
Limitations
The paper does not isolate the causal contribution of individual patches and cannot claim that harness evolution will work for all enterprise tasks.
StarHarness therefore offers a practical way to reduce persistent model-environment mismatch in tool-rich enterprise tasks.Found in the source text, word for word.
Picked because: Describes StarHarness, a framework that automatically evolves task‑specific agent harnesses (prompts, tool interfaces, loop configs) while keeping model weights fixed, and provides released tooling for enterprise LLM deployments.
Zhijie Zheng, Yu Li, Chen Qian and 5 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Existing guardrails evaluate only completed trajectories, leaving pre‑execution monitoring of step‑level actions underexplored, and simply applying trajectory‑level guards does not catch risky tool use before it happens.
Approach
The paper introduces StepGuard, a step‑level guard model that audits completed agent trajectories and checks tool actions before execution. StepGuard is trained with StepGen, an automatic data engine that creates paired safe and unsafe trajectories sharing the same context but differing at the risky step. To avoid over‑defense and under‑defense, Balance‑GRPO dynamically balances learning between safe and unsafe actions based on observed accuracy. The overall system combines a 4B LLM backbone, synthetic step‑level supervision, and a balanced reinforcement‑learning‑style fine‑tuning stage.
Result
StepGuard attains the highest average accuracy among open‑weight guard models, matching GPT‑5.4 performance, with 83.0 accuracy / 83.3 F1 on trajectory‑level and 84.8 accuracy / 84.1 F1 on step‑level benchmarks. When deployed, it cuts mean attack success rate by 77.3% relative to no‑guard while utility drops only 2.8 points, achieving ASR 1.2 and utility 90.7 on AgentDojo and ASR 9.3 and utility 66.7 on AgentDyn.
Why it matters
Researchers and engineers building LLM‑based agents should adopt StepGuard to obtain step‑level safety monitoring without large utility loss, especially in tool‑use environments.
Method details
StepGuard uses a 4B model backbone.
StepGen generates safe and unsafe trajectories with identical context but different actions at the risky step.
Balance‑GRPO adjusts the loss weighting between safe and unsafe classes based on their accuracy gaps.
Static evaluation benchmarks include ATBench, R‑Judge, ASSE Security, TS‑Bench‑Dojo, and TS‑Bench‑Harm.
Guarded‑agent evaluation uses AgentDojo, AgentDyn, and AgentHarm with a Qwen3.6‑35B‑A3B agent backbone.
Ablations test StepGen components (prefix supervision, benign tool reuse) and compare Balance‑GRPO against GRPO, safe up‑sampling, and fixed class weights.
Numbers
Trajectory‑level accuracy 83.0, F1 83.3 (StepGuard)
Step‑level accuracy 84.8, F1 84.1 (StepGuard)
Mean attack success rate reduction 77.3% relative to no‑guard
Utility drop 2.8 percentage points
AgentDojo ASR 1.2, utility 90.7
AgentDyn ASR 9.3, utility 66.7
Limitations
The method still struggles on highly adversarial harmful‑agent settings such as AgentHarm, where no evaluated guard achieves a clearly favorable safety‑utility trade‑off.
StepGuard achieves 83.0 accuracy and 83.3 F1 on trajectory-level evaluation, and 84.8 accuracy and 84.1 F1 on step-level evaluation.Found in the source text, word for word.
Picked because: Offers StepGuard, a step‑level guardrail model with scalable supervision and safety‑utility balancing, accompanied by artifacts for integrating pre‑execution safety checks into LLM agents.
Maitreyee Das Urmi, Jessica Pourleyli, Fabio Santos and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CR
Problem
Prior LLM code generation for security-sensitive tasks suffered high refusal rates and many security weaknesses, and simply adding security guidance to prompts did not consistently lower overall weakness prevalence.
Approach
The study designs a series of five prompt variants (Prompt 0.0 to 4.0) that incrementally add structural constraints, persona definition, output formatting, and security guidance, then generates Python code with two LLMs and evaluates the outputs with static analysis tools. Prompt 0.0 is a minimal baseline; Prompt 1.0 adds explicit output constraints; Prompt 2.0 adds high‑level security framework references; Prompt 3.0 removes framework citations while keeping security intent; Prompt 4.0 adds an adversarial context. The generated code is filtered for executability and analyzed with Bandit and CodeQL to measure compliance and security weakness severity and CWE distribution.
Result
Structured prompting dramatically lowered refusal rates, with GPT‑4o invalid outputs falling from 338 of 424 to between 37 and 52, and it shifted the severity profile of detected weaknesses: high‑severity findings dropped from 20.8% to 13.6% while low‑severity findings rose from 32% to 43.5%. LLaMA showed weaker and less consistent shifts.
Why it matters
Prompt engineers and developers using LLMs for code generation should note that structural prompt design can reduce refusals but cannot be relied on alone to ensure secure code.
Method details
GPT‑4o (proprietary high‑capacity model) used via API
LLaMA 3.1‑8B open‑weight model run locally
424 security‑sensitive Python tasks as the dataset
Code generation and analysis performed on macOS M1 CPU with 8 GB RAM, Python 3.11.5, Clang 2.23.3
Static analysis performed with Bandit (LOW, MEDIUM, HIGH severity) and CodeQL
Prompt variants 0.0 to 4.0 serve as baselines and incremental ablations
Numbers
invalid outputs, 338/424 → 37‑52, GPT‑4o
high‑severity findings, 20.8% → 13.6%
low‑severity findings, 32% → 43.5%
tasks, 424 security‑sensitive Python problems
Limitations
The paper shows prompt structure improves compliance but does not establish it as a reliable substitute for robust security controls.
Structured prompting substantially reduces refusals (e.g., GPT-4o invalid outputs drop from 338 of 424 to 37-52)Found in the source text, word for word.
Picked because: Provides an empirical study of how structured prompts affect security weaknesses in LLM‑generated Python code, along with a benchmark suite that engineers can use to evaluate and harden their code‑generation pipelines.