Paper 1 of 5
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
Problem
Prior skill‑evolution work left the insights that guide skill development scattered across optimization histories, which prevents systematic reuse across iterations. The obvious fix of simply reusing raw traces without a persistent store fails because it lacks long‑term historical awareness and repeats failed proposals.
Approach
WikiSkill introduces a three‑layer knowledge architecture: an immutable Raw Layer of execution traces, a persistent Wiki Layer that compiles and retains patterns across iterations, and a Skills Layer of active procedural instructions. In each evolutionary loop the Inference Agent runs rollouts while accessing only the Skills Layer, the Wiki Maintainer aggregates sampled traces into the Wiki, and the Skill Proposer (using the ReAct mechanism) suggests skill updates. A gating mechanism validates proposals and either commits them to the Skills Layer or rolls them back while keeping the Wiki unchanged. The Wiki accumulates pattern files, an evolution log, and a skill‑impact tracker that inform future proposals, enabling continuous knowledge refinement.
Result
WikiSkill achieves the highest average performance across all five models, improving over the strongest competing method by 3.3 points for Qwen‑3.5‑4B, 5.1 points for Qwen‑3.5‑9B, 10.0 points for Qwen‑3.6‑27B, 5.8 points for Gemma‑4‑31B, and 12.0 points for Gemini‑3.5‑Flash. It also raises individual task scores, e.g., LiveMath for Gemini‑3.5‑Flash from 33.0% to 72.6% and SpreadSheet from 50.5% to 76.6%. The method consistently outperforms no‑skill baselines in most model‑dataset pairs and matches or exceeds the best competing skill‑evolution method on average.
Why it matters
Researchers and engineers building interactive AI agents should care because WikiSkill demonstrates that a persistent knowledge base can substantially boost skill evolution across model scales and tasks.
Method details
- Evaluated on five benchmarks: LiveMath, SealQA, SpreadSheet, OfficeQA, and ALFWorld
- Used closed‑weight Gemini‑3.5‑Flash and open‑weight Qwen‑3.5‑4B, Qwen‑3.5‑9B, Qwen‑3.6‑27B, Gemma‑4‑31B‑It
- Compared against Trace2Skill, EvoSkill, SkillOpt, and a no‑skill baseline
- All skill‑evolution methods start with an empty skill set and run three independent evolution runs per method
- Ablation study shows that persistent Wiki accumulation is critical for effective skill evolution
Numbers
- LiveMath Qwen‑3.5‑4B WikiSkill 49.7 vs No skill 29.1
- SealQA Qwen‑3.5‑9B WikiSkill 43.1 vs No skill 26.3
- SpreadSheet Qwen‑3.6‑27B WikiSkill 81.7 vs No skill 40.8
- OfficeQA Gemma‑4‑31B WikiSkill 44.2 vs No skill 43.3
- ALFWorld Gemini‑3.5‑Flash WikiSkill 85.9 vs No skill 85.9 (no change)
Limitations
The paper does not explicitly discuss any limitations of the approach.
WikiSkill yields consistent improvements across models and datasetsFound in the source text, word for word.
Picked because: WikiSkill presents a concrete system that extracts and persists agent experience as reusable skills, enabling engineers to build and evolve LLM‑based tooling pipelines.