AIJul 6

MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution

arXiv:2607.0529727.8
Predicted impact top 3% in AI · last 90 daysOriginality Highly original
AI Analysis

This work addresses the limitation of non-recursive self-improvement in LLM agents by enabling recursive evolution of both task and meta-skills, which is a novel contribution for the agentic AI community.

MetaSkill-Evolve introduces a recursive two-timescale framework where task skills evolve on a fast loop and meta-skills (which govern the improvement pipeline) evolve on a slower loop, enabling self-improvement of LLM agents without additional models. It achieves +23.54, +16.09, and +1.92 point improvements over a frozen backbone on OfficeQA, SealQA, and ALFWorld benchmarks.

Recent LLM agents tackle increasingly long-horizon, open-ended tasks, and external skills, reusable procedural knowledge supplied to the agent, further extend this capability. However, a fixed, hand-authored skill is rarely optimal, and cannot adapt to the diversity of tasks an agent encounters. Self-improving agents address this by rewriting their own skill files from execution traces, yielding meaningful gains on challenging benchmarks. Yet such self-evolution remains non-recursive: it improves only the task skill (what the agent does) while the improvement procedure (how it improves) is authored once and held fixed. We introduce MetaSkill-Evolve, a two-timescale framework that makes agentic skill improvement recursive: every branch carries both a task skill $s$ and a branch-local meta-skill $m=(ψ,σ,α,π,\varepsilon)$ whose five components parameterise the Analyzer, Retriever, Allocator, Proposer, and Evolver agents of the improvement pipeline. Task skills evolve on a fast loop while the meta-skill evolves on a slower one under the same pipeline applied to itself, with no additional model or objective. With all five pipeline agents sharing a single frozen backbone, MetaSkill-Evolve outperforms no-skill, static-skill, and single-level evolution baselines on three agentic benchmarks (OfficeQA, SealQA, ALFWorld), improving held-out test accuracy over the raw backbone by +23.54, +16.09, and +1.92 points respectively.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes