arXiv digest

Thursday

September 17, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Replication-Aware Placement of Functions and Data in the Edge-Cloud Continuum

Dario d'Abate, Matteo Cenzato, Matteo Briscini and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.DC

Problem

Prior work did not jointly schedule stateless functions and place replicated data under heterogeneous consistency, leading to high client latency; simply placing all data centrally (cloud‑only) ignores locality and consistency constraints and thus fails to reduce latency.

Approach

The paper formulates the joint function‑scheduling and data‑placement problem as a Binary Linear Programming (BLP) model that captures SR and ER consistency, then introduces a topology‑aware greedy heuristic (TA) that approximates the BLP solution efficiently by exploiting the hierarchical tree structure of the edge‑cloud continuum.

Result

TA achieves placement quality close to the BLP optimum, with latency gaps shrinking to near zero as scale increases and storage gaps collapsing to zero, while consistently outperforming the CO and CD naive baselines across different consistency mixes.

Why it matters

Edge‑cloud platform designers and researchers should care because the TA heuristic enables near‑optimal function and data placement with low computational cost, supporting periodic reconfiguration in heterogeneous environments.

Method details
  • BLP scales cubically with the number of infrastructure nodes
  • TA is a greedy heuristic that respects the tree topology
  • Synthetic workloads use a balanced -ary tree topology
  • Implementation uses C++ and IBM ILOG CPLEX for the BLP
  • Baselines are Cloud‑only (CO) and Cloud‑data (CD) policies
  • Experiments run on a 16‑core AMD Ryzen 9 9950X server with 64 GB RAM
Numbers
  • processor cores, 16, AMD Ryzen 9 9950X
  • threads, 32, AMD Ryzen 9 9950X
  • RAM, 64 GB, server
  • BLP timeout, 600 s, enforced
  • virtual memory cap, 50 GB, enforced
  • read ratio, 0.8, functions
  • replicas factor, 3, evaluated
  • seeds, 10, per instance
Limitations

The evaluation is limited to synthetic scenarios and does not demonstrate results on real‑world workloads; the BLP becomes intractable for larger problem sizes.

TA approaches the BLP optimum across the explored rangeFound in the source text, word for word.

Picked because: Presents a concrete Binary Linear Programming model and implementation for jointly scheduling functions and placing data across edge‑cloud, directly applicable to DevOps and infrastructure automation.

Paper 2 of 5

Ermes: a Stateful Serverless Platform for the Edge-to-Cloud Continuum

Matteo Cenzato, Dario d'Abate, Arianna Dragoni and 4 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.DC

Problem

Stateless FaaS forces functions to fetch state from remote cloud stores, reintroducing latency that edge computing aims to eliminate, and existing fixes that simply move computation close to the client do not address remote state access.

Approach

Ermes integrates state management into the FaaS model by organizing state into collections and using a distributed coordination algorithm that jointly maps collections and functions onto edge, fog, and cloud nodes. It supports fine‑grained replication and per‑collection consistency levels (sequential or eventual). Functions declare collection properties rather than identities, and the platform resolves these at invocation time. The system places state close to where functions execute and migrates replicas as demand shifts, aiming to minimize client‑perceived latency.

Result

Placement turns remote accesses into local ones within a few epochs, provisioning replicas without perturbing ongoing writes; the remaining latency is near‑local under eventual consistency and bounded by the single‑leader floor under sequential consistency. Ermes sustains low and stable latency as concurrency, state cardinality, and data‑intensity grow, degrading only under strict ordering or edge memory saturation. Compared to cloud‑only and cloud‑data baselines, Ermes outperforms both by a wide margin across all scenarios.

Why it matters

Edge and serverless platform engineers should care because Ermes demonstrates a practical way to reduce latency by co‑locating state with computation across the edge‑to‑cloud continuum.

Method details
  • Implemented in Go, leveraging its lightweight concurrency for asynchronous message‑driven operation.
  • State stored in memory using Redis.
  • Execution uses Wasmtime with just‑in‑time compilation and a hybrid caching mechanism.
  • Supports two consistency policies per collection: sequential consistency and eventual consistency.
  • Placement algorithm jointly maps collections and functions across a three‑tier hierarchy (cloud, mid, edge).
  • Provides fine‑grained replication with per‑collection consistency and access policies.
Numbers
  • RTT delay, 5 ms to 20 ms, per link in emulated network
  • Number of collections, up to a thousand, in scalability study
  • Node hierarchy, 1 Cloud Node, 2 Mid Nodes, 5 Edge Nodes, in deployment
  • CPU resources, 8 vCPUs (cloud), 4 vCPUs (mid), 2 vCPUs (edge), in deployment
  • Memory resources, 8 GB (cloud), 4 GB (mid), 2 GB (edge), in deployment
  • Baseline count, 2 cloud‑based baselines (CO and CD), in comparison
Limitations

Placement currently considers only storage capacity, not compute resources, and consistency policies are static; fault‑tolerance mechanisms are not fully explored.

Ermes outperforms both the CD and CO baselines by a wide margin throughoutFound in the source text, word for word.

Picked because: Introduces Ermes, a released stateful serverless platform for the edge‑to‑cloud continuum, offering engineers a self‑hosted runtime they can deploy and extend.

Paper 3 of 5

Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

João Meneses dos Santos, Arlindo L. Oliveira · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Language agents were brittle in interactive environments, failing to track long‑horizon state, execute valid actions, and recover from mistakes; simply adding memory or reflection alone did not fix the instability of execution.

Approach

The method augments the SwiftSage dual‑process controller with two modular extensions. An Adaptive Memory Module (AMM) writes salience‑gated episodic records and retrieves them on trigger events to augment Swift, Sage, and Critic prompts. A Self‑Reflection Module (SRM) inserts a fast Gate‑1 validation before actions reach the environment, detects stagnation, and optionally calls a bounded Critic to inject corrective actions. Both modules are feature‑flagged so they can be ablated independently. The full system runs the same Swift/Sage arbitration and action‑buffer execution as the baseline, but with AMM and SRM inserted at precise runtime interfaces.

Result

Across four configurations the full system attains the highest mean final score of 64.62, a success rate of 43.17%, and the lowest Steps@Succ of 19.33, while SRM alone provides most of the gain (final score 64.33, success 41.33%). AMM alone yields modest improvements (final score 53.83, success 26.94%).

Why it matters

Researchers building interactive language agents should consider modular, causally placed memory and reflection components, as they can substantially improve reliability without increasing deliberation cost.

Method details
  • Local model runtime uses Qwen2.5‑7B‑Instruct‑1M for Sage and Critic
  • Evaluation dataset is ScienceWorld tasks 0‑29 with up to 10 natural‑language variations per task (≈271 episodes per configuration)
  • Four configurations compared: baseline SwiftSage, baseline+AMM, baseline+SRM, and full system
  • AMM is evaluated with a populated memory agent to reuse prior episodes
  • SRM Critic budget is limited to three calls per episode in the main experiments
  • Ablations include disabling Swift‑level T1 memory injection and increasing Critic budget from three to six calls
Numbers
  • Final score, 64.62, full system vs baseline 51.67
  • Success rate, 43.17%, full system vs baseline 23.99%
  • Steps@Succ, 19.33, full system vs baseline 24.48
  • Final score, 64.33, baseline+SRM vs baseline 51.67
  • Final score, 53.83, baseline+AMM vs baseline 51.67
  • Success rate improvement, 25.1%, full system vs baseline
Limitations

The paper notes limited memory representation, hand‑crafted reflection rules, evaluation confined to ScienceWorld, and only correlational prompt‑level analysis.

the full system achieves the best mean final score (64.62), success rate (43.17%), and successful-step efficiency (19.33 steps)Found in the source text, word for word.

Picked because: Extends the SwiftSage dual‑process LLM agent with modular memory and self‑reflection components, providing reusable code for building more reliable autonomous agents.

Paper 4 of 5

Affora: A Design System for Agent-Friendly Interfaces

Jin Gao · abstract · pdf

quote verifiedfigures checkedread: full textcs.HC

Problem

Computer-use agents fail because interfaces often omit a machine-readable semantic substrate-controls, stable names, state, and choices are not explicitly represented. The missing substrate prevents agents from reliably identifying operable elements regardless of visual styling. Simply restyling or adding visual cues does not fix the underlying semantic deficit.

Approach

Affora is a design system that separates visual expression from the semantic substrate, making task-relevant structure explicit and machine‑readable while preserving visual freedom. It provides reusable UI components that embed required semantic information by default and a set of design rules applied at component, layout, flow, and site scales. Designers can vary colour, typography, shape, and composition as long as the underlying interaction meaning is retained. The system is delivered as plug‑and‑play components and guidelines, enabling AI‑assisted coding agents to reuse them without inferring semantics. Evaluation uses an agent‑evaluation harness with OpenAI ChatGPT 5.6 Luna and Terra models to run tasks on WebArena Magento, WebShop, and other applications.

Result

Affora achieved full completion on both the primary and harder test sets (60/60, 100% and 33/33, 100%). The instruction‑file condition also reached full completion on the primary set (60/60, 100%) and 97% on the harder set (32/33). Baseline performance was 77% (46/60) on the primary set and 85% (28/33) on the harder set, while ARIA and structured data yielded 75% (45/60) and 88% (29/33) respectively. The WebMCP‑style condition achieved 100% (55/55) on its applicable action set.

Why it matters

Interface designers and developers of computer‑use agents should care because Affora improves agent reliability without constraining visual design, enabling more robust automation of existing web interfaces.

Method details
  • Models: OpenAI ChatGPT 5.6 Luna and ChatGPT 5.6 Terra, low reasoning tier.
  • Datasets: WebArena Magento storefront, WebShop, and four independently authored web applications.
  • Baselines: ARIA, structured‑data augmentation, instruction‑file (agent.md), and WebMCP‑style action layer.
  • Evaluation harness: ReAct‑style observe-reason-act cycles with bounded retry and task‑scoped memory.
  • Comparison: measured successful episodes over recorded attempts for each intervention.
Numbers
  • Primary set baseline: 46/60 (77%) vs Affora: 60/60 (100%)
  • Harder set baseline: 28/33 (85%) vs Affora: 33/33 (100%)
  • Instruction file primary set: 60/60 (100%) vs baseline 46/60 (77%)
  • WebMCP‑style applicable set: 55/55 (100%)
Limitations

The paper does not establish that one mechanism is universally superior and the layout experiment does not validate every possible composition.

Affora and the instruction-file condition both reach full completion, an improvement of approximately 23 percentage points over baselineFound in the source text, word for word.

Picked because: Delivers Affora, a design system and component library that makes user interfaces agent‑friendly, with concrete guidelines and artifacts for engineering agent‑human interaction.

Paper 5 of 5

Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN

Seyed Bagher Hashemi Natanzi, Bo Tang · abstract · pdf

quote verifiedfigures checkedread: full textcs.NI

Problem

Two independent agents, one protecting latency SLA and one maximizing utilization, cause recurring opposing excursions of the shared PRB partition, and existing conflict‑mitigation assumes a static set of applications and cannot handle run‑time emergent behavior.

Approach

AURA interposes a lightweight arbiter between agents and the network. Each agent submits a proposal containing its intended quota change and priority class. The arbiter admits a proposal only if (i) the resulting shared state satisfies feasibility invariants, (ii) no variable touched was modified within a dwell‑time window, and (iii) the change lies outside a deadband. Accepted proposals update a shared quota table that the OAI gNB MAC downlink pre‑processor enforces as per‑slice PRB quotas. Rejected proposals are returned for re‑observation, ensuring only safe actions affect the radio resources.

Result

AURA reduces shared‑state PRB excursions from 8.4 to 0.4 PRB amplitude and cuts cross‑slice throughput starvation from 40‑55% to 0.3%, while the protected slice's latency compliance remains unchanged; slice‑1 latency‑violation rate is slightly higher (92.9% vs 84.5% Direct) but the system‑wide cost is explicit.

Why it matters

Network operators and O‑RAN developers should care because AURA provides a provably convergent arbitration layer that prevents destabilizing interactions between autonomous slicing agents without needing to modify agent internals.

Method details
  • Agents run as rApps in the non‑real‑time RIC and observe telemetry every seconds with a control period of seconds.
  • Arbiter checks three conditions: feasibility invariants, per‑variable dwell time, and deadband, with SLA‑restoring actions taking priority.
  • Enforcement is performed by a quota module in the OAI NR MAC downlink pre‑processor that reloads a file‑based quota table.
  • Evaluation compares four regimes: Static, Single (only SLA agent), Direct (both agents, no arbiter), and AURA (both agents, arbitrated).
  • Experiments run on a containerized OAI 5G stack on an AMD EPYC 9354 server with 64 threads and 377 GiB RAM.
Numbers
  • PRB amplitude reduction, 0.4 vs 8.4, Direct regime
  • Throughput starvation, 0.3% vs 40‑55%, Direct regime
  • Slice‑1 latency violation, 92.9% vs 84.5%, Direct regime
  • Latency gap to Static, 4.2 percentage points, AURA settled point
Limitations

The paper does not demonstrate scalability to many agents or multi‑cell deployments, and it leaves the protected slice's latency improvement unaddressed.

AURA reduces recurring shared-state excursions by more than an order of magnitude (from 8.4 to 0.4 PRB amplitude) and virtually eliminates cross-slice throughput starvation (from 40-55% to 0.3%)Found in the source text, word for word.

Picked because: Demonstrates a live O‑RAN deployment where autonomous AI agents are arbitrated for stability, offering practical insights and tooling for safety‑critical agentic infrastructure.