arXiv digest

Friday

August 28, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Prior skill‑evolution work left the insights that guide skill development scattered across optimization histories, which prevents systematic reuse across iterations. The obvious fix of simply reusing raw traces without a persistent store fails because it lacks long‑term historical awareness and repeats failed proposals.

Approach

WikiSkill introduces a three‑layer knowledge architecture: an immutable Raw Layer of execution traces, a persistent Wiki Layer that compiles and retains patterns across iterations, and a Skills Layer of active procedural instructions. In each evolutionary loop the Inference Agent runs rollouts while accessing only the Skills Layer, the Wiki Maintainer aggregates sampled traces into the Wiki, and the Skill Proposer (using the ReAct mechanism) suggests skill updates. A gating mechanism validates proposals and either commits them to the Skills Layer or rolls them back while keeping the Wiki unchanged. The Wiki accumulates pattern files, an evolution log, and a skill‑impact tracker that inform future proposals, enabling continuous knowledge refinement.

Result

WikiSkill achieves the highest average performance across all five models, improving over the strongest competing method by 3.3 points for Qwen‑3.5‑4B, 5.1 points for Qwen‑3.5‑9B, 10.0 points for Qwen‑3.6‑27B, 5.8 points for Gemma‑4‑31B, and 12.0 points for Gemini‑3.5‑Flash. It also raises individual task scores, e.g., LiveMath for Gemini‑3.5‑Flash from 33.0% to 72.6% and SpreadSheet from 50.5% to 76.6%. The method consistently outperforms no‑skill baselines in most model‑dataset pairs and matches or exceeds the best competing skill‑evolution method on average.

Why it matters

Researchers and engineers building interactive AI agents should care because WikiSkill demonstrates that a persistent knowledge base can substantially boost skill evolution across model scales and tasks.

Method details
  • Evaluated on five benchmarks: LiveMath, SealQA, SpreadSheet, OfficeQA, and ALFWorld
  • Used closed‑weight Gemini‑3.5‑Flash and open‑weight Qwen‑3.5‑4B, Qwen‑3.5‑9B, Qwen‑3.6‑27B, Gemma‑4‑31B‑It
  • Compared against Trace2Skill, EvoSkill, SkillOpt, and a no‑skill baseline
  • All skill‑evolution methods start with an empty skill set and run three independent evolution runs per method
  • Ablation study shows that persistent Wiki accumulation is critical for effective skill evolution
Numbers
  • LiveMath Qwen‑3.5‑4B WikiSkill 49.7 vs No skill 29.1
  • SealQA Qwen‑3.5‑9B WikiSkill 43.1 vs No skill 26.3
  • SpreadSheet Qwen‑3.6‑27B WikiSkill 81.7 vs No skill 40.8
  • OfficeQA Gemma‑4‑31B WikiSkill 44.2 vs No skill 43.3
  • ALFWorld Gemini‑3.5‑Flash WikiSkill 85.9 vs No skill 85.9 (no change)
Limitations

The paper does not explicitly discuss any limitations of the approach.

WikiSkill yields consistent improvements across models and datasetsFound in the source text, word for word.

Picked because: WikiSkill presents a concrete system that extracts and persists agent experience as reusable skills, enabling engineers to build and evolve LLM‑based tooling pipelines.

Paper 2 of 5

SWE-Prime: Fewer Trajectories, Better Performance

Dewu Zheng, Ruizhe Ye, Yanlin Wang and 7 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Prior work fine‑tunes on all successful trajectories, but successful trajectories can still contain ineffective, redundant, or risky steps, so using them directly introduces noisy supervision and undesirable problem‑solving behaviors.

Approach

SWE‑Prime is a two‑stage data selection pipeline. Stage 1 screens whole trajectories using three criteria: process quality, result quality, and data representativeness, and keeps a high‑quality, representative subset. Stage 2 splits each retained trajectory into semantic segments, scores each segment on contribution, learnability, and risk, and masks low‑scoring segments from the loss while keeping them in context. During supervised fine‑tuning the loss is computed only on assistant tokens in the selected segments. The pipeline thus filters supervision at both trajectory and segment granularity before SFT.

Result

Training on the 10 % trajectory subset selected by SWE‑Prime yields relative performance gains of up to 12.2 % on SWE‑Bench Pro and 24.2 % on SWE‑Bench Verified compared with training on the full resolved dataset, and improves process‑quality metrics such as redundancy, observe‑before‑edit rate, and tool success.

Why it matters

Researchers and engineers building coding agents should consider quality‑aware trajectory and segment selection to improve SFT efficiency and final performance.

Method details
  • Base models evaluated: GLM‑4.7‑Flash, Qwen3‑30B‑A3B‑Instruct, and Qwen3‑Coder‑30B‑A3B‑Instruct.
  • Training data: SWE‑rebench OpenHands Trajectories dataset with 67,074 trajectories, of which 32,161 are successful.
  • Benchmarks: SWE‑Bench Verified (500 tasks) and SWE‑Bench Pro public split (731 tasks).
  • Baseline comparisons: Raw Model, Resolved‑Trajectory SFT (full resolved dataset), and SWE‑Prime.
  • Hyperparameter ablations: trajectory retention ratios of 5 %, 10 %, 20 %, 30 % and segment score thresholds of 5‑9, leading to the chosen 10 % retention and threshold 7.
Numbers
  • Resolved Rate, 50.2%, at 10% trajectory retention (peak compared to 46.2% at 5%).
  • Resolved Rate, 53.2%, with segment score threshold 7 (compared to 46.8% at threshold 5).
  • Relative performance gain, 12.2%, on SWE‑Bench Pro versus full resolved dataset.
  • Relative performance gain, 24.2%, on SWE‑Bench Verified versus full resolved dataset.
  • Redundancy reduction, up to 0.188, compared to Resolved‑Trajectory SFT.
Limitations

The paper does not demonstrate that the selected 10 % retention and threshold 7 configuration generalizes beyond the three evaluated models and two benchmarks.

training on the 10% trajectory subset selected by SWE-Prime outperforms training on the full resolved dataset, yielding relative performance gains of up to 12.2% and 24.2%, respectively.Found in the source text, word for word.

Picked because: SWE‑Prime proposes a method to filter and improve LLM‑generated software‑fix trajectories, offering actionable guidance for deploying LLM assistants in real‑world dev‑ops workflows.

Paper 3 of 5

Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit

Yisen Xi · abstract · pdf

quote verifiedfigures checkedread: full textcs.SE

Problem

Prior single‑domain LLM agents could not let the persona evolve freely while keeping execution traceable without costly reconstruction; the obvious fix of adding typed change objects, an external gate, and a stable audit anchor re‑introduces tight coupling and defeats the goal of cheap, independent drift and audit.

Approach

Persona‑Execution Separation (PES) splits the agent into two trust domains: a permissive domain for the mutable persona and a restrictive domain for audited execution. A governed contract bridge connects them, enforcing an approval matrix, data‑loss‑prevention (DLP) scanning, ACL checks and audit writes on every crossing. The persona remains singly‑homed and may drift; execution is faceless and fully logged. Identity continuity is maintained via conversation‑run binding. The bridge is fail‑closed to prevent unauthorized leakage.

Result

The PES implementation showed zero execution‑side re‑validation in all five perturbation rounds (R = 0/5 = 0.00) and no detectable persona fingerprint in the A/B evaluation, with only a single extra retrieval in seven runs. Cross‑model tests confirmed V1 zero re‑validation on all five models, while V2 passed on four models and qwen3.8‑max failed to converge, rendering its V2 result unmeasurable.

Why it matters

Organizations that need regulated digital‑employee platforms with frequent persona updates and strict execution audit should consider PES to decouple persona drift from audited actions.

Method details
  • Five perturbation rounds (L1‑L5) were run on five model configurations: deepseek‑v4‑flash, deepseek‑v4‑pro, qwen3.8‑max, kimi‑k3, glm‑5.3.
  • A development/pilot case recorded five decisions over one month, each with a rejected alternative.
  • Baseline single‑domain probe used deepseek‑v4‑flash to test persona‑injection paths.
  • Cross‑model run ledger: deepseek‑v4‑flash 13 attempts, 13 completed; kimi‑k3 10 attempts, 4 completed, 6 excluded (provider capacity); glm‑5.3 5 attempts, 3 completed, 2 excluded (gateway 404, rate‑limit).
  • V1 measured zero execution‑side re‑validation (R = 0/5 = 0.00) across all five configurations.
  • V2 A/B test showed no persona fingerprint; an extra retrieval occurred once in seven runs (1/7).
Numbers
  • V1 re‑validation rate: 0/5 = 0.00
  • Extra retrieval in V2: 1/7 runs
  • deepseek‑v4‑flash attempts/completed: 13/13
  • kimi‑k3 attempts/completed: 10/4 (6 excluded)
  • glm‑5.3 attempts/completed: 5/3 (2 excluded)
  • qwen3.8‑max attempts/completed: 8/0 (4 transport errors, 4 round‑cap aborts)
Limitations

The paper does not provide a paired quantitative comparison showing that PES improves auditability over single‑domain designs, so the measured win of PES over baseline is not established.

V1’s zero replicated on five configurations (v4-flash, v4-pro, qwen3.8-max, kimi-k3, glm-5.3).Found in the source text, word for word.

Picked because: Persona‑Execution Separation defines an architecture that isolates LLM persona from audited execution state, giving engineers a practical pattern for safe, verifiable agent deployment.

Paper 4 of 5

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

Kairong Luo, Jiarui Cui, Yaorui Yin and 8 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Curriculum ordering conflicts with conventional terminal learning rate decay because examples presented near the end receive very small parameter updates, so simply applying a decay does not improve data efficiency.

Approach

The paper introduces Curriculum Model Averaging (CMA) which combines an ascending data curriculum with a constant learning rate continuation in the late stage and averages multiple late checkpoints. Phase 2 builds a curriculum that orders data within each component from low to high quality. After reaching step 218,000 the training resumes with the base learning rate fixed, and six checkpoints spaced about 100 optimizer steps apart are equally averaged. This design decouples curriculum progression from learning rate decay while preserving the benefits of high‑quality data exposure.

Result

The resulting Puro‑2B model approaches Qwen2.5‑1.5B performance while costing under $6.9K to train, and a derived cost scaling law indicates that about $4.4K is enough to match Qwen2‑1.5B performance.

Why it matters

Researchers and small labs seeking cost‑effective, reproducible pretraining can adopt this recipe to build competitive 2B‑scale models without large data‑center resources.

Method details
  • 2B‑parameter base model trained from scratch on 1.4 trillion tokens
  • FP8 mixed‑precision training on consumer‑grade RTX 5090 GPUs
  • Compute cost of the best model is less than $6.9K
  • Late‑stage constant‑LR continuation resumes from step 218,000
  • Six checkpoints are averaged, each separated by roughly 100 optimizer steps
  • Curriculum constructed with 376 intervals (buckets) across components
Numbers
  • compute cost, < $6.9K, compared to Qwen2.5‑1.5B performance
  • cost to match Qwen2‑1.5B, $4.4K, compared to $5,090 baseline
  • training tokens, 1.4 trillion, total pretraining data
  • curriculum buckets, 376, used to order data
  • resume step, 218,000, start of constant‑LR continuation
  • averaged checkpoints, 6, used for final model
Limitations

The paper does not establish performance beyond its specific evaluation protocol or on downstream tasks not covered in the study.

Our best model is trained at a compute cost of less than $6.9K and approaches Qwen2.5-1.5B performance under our evaluation protocol.Found in the source text, word for word.

Picked because: Puro‑2B releases a low‑cost, RTX 5090‑based recipe for pre‑training a 1.5 B‑parameter LLM, providing a reproducible workflow for self‑hosted model training.

Paper 5 of 5

BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks

Jeong-Yoon Kim · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Existing industrial telemetry benchmarks do not specify how to compile read-only logs into executable multi-turn agent tasks. Manually constructing such benchmarks does not scale to the volume of data in sites like BTS.

Approach

BTS-AgentBench normalizes BTS metadata and raw histories into a read-only tool store, then compiles static tasks with tool-derived gold answers and evidence. Retained tasks are lifted into typed, bounded operator-facing episodes via deterministic episode compilation. The pipeline includes deterministic contract transformation and an audit trail for replay. A controller-aware acceptance rule hardens the release by rejecting rows that fail the deterministic audit suite. The entire process is deterministic and replayable, enabling exact row-level reproducibility.

Result

On the released test split, GPT‑5.5 solved 79 out of 89 rows (88.8%), Gemini 3.1 Pro solved 71/89 (79.8%) and Claude Opus 4.7 solved 58/89 (65.2%). Deterministic diagnostic scores show GPT‑5.5 achieving a final score of 0.978, evidence coverage 0.955, phase compliance 0.949 and task score 0.965, with 86/89 rows passing the protocol check. The deterministic controller solved none of the test rows, confirming the hardening effect.

Why it matters

Researchers building industrial AI assistants should care because BTS‑AgentBench provides a reproducible, deterministic pipeline to turn real telemetry logs into multi‑turn agent evaluation without manual annotation.

Method details
  • Dataset: BTS telemetry with 532 rows (train/dev/test split 356/87/89).
  • Dataset: XAI4HEAT telemetry converted to 204 episodes, with a 41‑row held‑out test split.
  • Frontier models evaluated: GPT‑5.5 (OpenAI 2026), Gemini 3.1 Pro (Google 2026), Claude Opus 4.7 (Anthropic 2026).
  • Baseline: deterministic construction‑exclusion controller that solves 0/532 rows.
  • Two independent raw‑to‑episode builds reproduced all 11 logical tool‑store exports and the exact 356/87/89 split.
Numbers
  • Overall success rate, 79/89 (88.8%), GPT‑5.5 on released test rows
  • Overall success rate, 71/89 (79.8%), Gemini 3.1 Pro on released test rows
  • Overall success rate, 58/89 (65.2%), Claude Opus 4.7 on released test rows
  • Controller solved, 0/89 (0.0%), deterministic controller on test split
  • Final diagnostic score, 0.978, GPT‑5.5
  • Total released rows, 532, BTS‑AgentBench benchmark
Limitations

The benchmark only evaluates read‑only telemetry search, aggregation, comparison, ranking, timestamp reporting and quality‑aware reporting; it does not cover write‑side control, safety‑critical actuation, or long‑horizon troubleshooting.

GPT-5.5 accomplishes 79/89, Gemini 3.1 Pro 71/89, and Claude Opus 4.7 58/89.Found in the source text, word for word.

Picked because: BTS‑AgentBench delivers an end‑to‑end pipeline that turns read‑only telemetry logs into deterministic agent benchmarks, a ready‑to‑use tool for infrastructure automation and evaluation.