HyperAIHyperAI

Command Palette

Search for a command to run...

LLM
Agent

EVISKILL: GROUNDING SKILL EVOLUTION IN RE-PLAYABLE EVIDENCE

Yan Zhou Yili Wang Yiwei Dai Qinggang Zhang Xin Wang

Abstract

Continual skill evolution enables LLM agents to accumulate and refine reusable procedural knowledge from interaction experience without updating model parameters. Its effectiveness depends on determining not only what to change, but also why a change is justified and when it should become persistent guidance. However, existing experience-driven methods can lose the behavioral evidence and task contexts supporting edits. Moreover, a global validation outcome provides an incomplete judgment of its constituent changes: locally supported corrections may be discarded with a rejected revision, while evidence may require further experience to inform useful updates. To this end, we introduce EVISKILL, an evidence-driven framework that organizes execution observations into Replayable Evidence Cards and synthesizes edits with explicit links to their supporting contexts. Targeted replay verifies these edits through re-execution and provides feedback for correction. Across epochs, EVISKILL preserves evidence and provisionally retains supported edits for further refinement, while global validation governs their incorporation into the final skill. Experiments on three interactive benchmarks across six LLM backbones demonstrate the effectiveness of this approach. Our code and implementation details are available at https://github.com/Zhouyaner/eviskill.

One-sentence Summary

Researchers at Jilin University propose EVISKILL, an evidence-driven framework for continual LLM skill evolution that stores execution observations as Replayable Evidence Cards, synthesizes edits explicitly linked to supporting contexts, verifies those edits through targeted replay, and provisionally retains supported edits for further refinement before global validation incorporates them into final skills, with effectiveness demonstrated on three interactive benchmarks across six LLM backbones.

Key Contributions

  • EVISKILL is an evidence-driven skill evolution framework that organizes execution observations into Replayable Evidence Cards and links skill edits to their supporting behavioral contexts.
  • Replay-guided edit verification tests whether edits produce their intended behavioral effects through re-execution, while cross-epoch evidence propagation retains and refines supported edits even when the overall revision is rejected.
  • Experiments on three interactive benchmarks across six LLM backbones show consistent improvements over strong skill evolution baselines, with analyses indicating that evidence-grounded replay filters unsupported revisions and cross-epoch propagation enables prior evidence to support continued refinement.

Introduction

LLM agents increasingly handle interactive tasks such as tool use, web navigation, and long-horizon decision making, where reusable external skills provide procedural guidance without changing model parameters. Prior experience-driven skill evolution methods extract updates directly from execution trajectories, but execution experience is local and context-dependent, so edits can overfit observed cases or capture incidental behavior. Rejected revisions may also contain useful edits that are discarded. The authors introduce EVISKILL, an evidence-driven framework that organizes continual skill evolution into evidence construction, behavioral verification, and cross-epoch refinement. EVISKILL grounds candidate edits in replayable execution evidence, re-executes referenced behaviors to verify intended effects, and preserves uncertain or unresolved evidence for reconsideration across future epochs. Experiments on three interactive benchmarks across multiple LLM backbones show consistent improvements over strong skill evolution baselines.

Method

The authors introduce EVISKILL, a framework that organizes skill evolution around Replayable Evidence Cards. These cards record proposed skill corrections alongside their supporting execution evidence, enabling a structured pipeline for continuous improvement. The overall architecture consists of three interconnected phases: Evidence-Grounded Edit Synthesis, Replay-Guided Edit Verification, and Cross-Epoch Evidence Propagation.

As shown in the figure below:

In the first phase, Evidence-Grounded Edit Synthesis, the system distinguishes between the Validated Skill, which is updated only upon successful global validation, and the Working Skill used for ongoing interactions. At epoch eee, the Working Skill SeWS_e^WSeW​ combines the latest Validated Skill SeVS_e^VSeV​ with a Provisional Edit Ledger PeP_ePe​, which contains replay-supported edits retained for further refinement:

SeW=SeV⊕PeS_e^W = S_e^V \oplus P_eSeW​=SeV​⊕Pe​

where ⊕\oplus⊕ applies the edits sequentially. The Action Agent collects trajectories using SeWS_e^WSeW​, providing the foundational evidence for skill revision. An LLM evidence extractor analyzes the Working Skill and its trajectories to propose corrections, recording each proposal as a Replayable Evidence Card. A card CiC_iCi​ is formally defined as:

Ci=⟨i,Gi,di⟩C_i = \langle i, \mathcal{G}_i, d_i \rangleCi​=⟨i,Gi​,di​⟩

where iii is a persistent identifier, Gi\mathcal{G}_iGi​ represents supporting trigger ranges from trajectories, and did_idi​ is the proposed correction. The framework extracts Trajectory Evidence Cards from current interactions and Contrastive Evidence Cards from adjacent-epoch trajectory comparisons to identify persistent deficiencies. These cards are combined into an Evidence Pool, grouped into Evidence Windows based on compatible contexts, and processed by an LLM editor to synthesize a candidate edit set Ue\mathcal{U}_eUe​.

The second phase, Replay-Guided Edit Verification, ensures that synthesized edits produce their intended behavioral effects. For each candidate edit u∈Ueu \in \mathcal{U}_eu∈Ue​, the system retrieves its supporting cards and replays the associated trigger ranges. By reproducing the prefix of the source trajectory, the framework reconstructs the execution state and re-executes the segment under the modified skill SeW⊕uS_e^W \oplus uSeW​⊕u. An LLM evaluator assesses the source and replayed segments, returning a decision to accept, reflect, or reject the edit. If the decision is to reflect, the editor generates a revised edit u′u^\primeu′ based on the feedback, which then undergoes another round of replay and evaluation. Successfully verified edits are consolidated into an ordered collection, yielding a candidate revision S~e\widetilde{S}_eSe​ for global validation.

The final phase, Cross-Epoch Evidence Propagation, governs how skills and evidence are updated across epochs. Global validation compares the candidate revision S~e\widetilde{S}_eSe​ against the current Validated Skill SeVS_e^VSeV​ on a validation set. If S~e\widetilde{S}_eSe​ demonstrates superior performance, it replaces SeVS_e^VSeV​, the Provisional Edit Ledger is cleared, and the incorporated edits become the new Tracked Edits for future cross-epoch comparisons.

If the candidate revision is globally rejected, the framework employs post-rejection replay to salvage effective edits. Each consolidated edit is replayed independently under the original Validated Skill SeV⊕vS_e^V \oplus vSeV​⊕v to determine if it remains effective without the provisional edits. Edits that pass this isolated evaluation are retained in the updated Provisional Edit Ledger for the next epoch, ensuring that locally beneficial corrections are not discarded due to global incompatibilities. Furthermore, the lifecycle of the Evidence Cards is dynamically managed. Cards supporting accepted edits are archived, while those supporting rejected edits may be marked as stale or protected for future cross-epoch comparisons. After EEE evolution epochs, the final Validated Skill is frozen for inference, ensuring that only thoroughly validated and replay-supported skills are deployed.

Experiment

The preliminary study identifies two weaknesses in existing skill evolution: edits based on experience often fail to produce intended execution effects, and globally rejected revisions may still contain useful constituent edits. The main experiments evaluate EVISKILL across App-World, ScienceWorld, and ALFWorld with multiple LLM backbones, showing that grounding skill updates in execution evidence yields broad gains and that evolving a skill does not guarantee improvement over its initial version. Ablation and mechanism analyses confirm that replay-guided edit verification improves behavioral corrections, while cross-epoch evidence propagation preserves supported edits and unresolved evidence, with many edits from rejected revisions later incorporated into accepted skills and evidence cards reused across epochs.

Execution-grounded skill evolution yields broad task-success gains across interactive benchmarks, improving over the no-skill baseline in every setting and ranking among the top methods in most comparisons. Several skill-evolution baselines regress on ScienceWorld, falling below both the no-skill and initial-skill accuracy, while the execution-grounded approach avoids this pattern. The results indicate that verifying updates through observed behavior helps preserve useful skills while incorporating new experience. Execution-grounded skill updates achieve the best accuracy in most model-dataset settings and improve over no skill in all reported settings. On ScienceWorld, several evolution baselines underperform the initial skill, with some dropping below the no-skill baseline, whereas execution-grounded updates improve over the initial skill in nearly all settings.

The experiments evaluate execution-grounded skill evolution on interactive benchmarks, testing whether updates grounded in observed behavior improve task success. The approach consistently improves over the no-skill baseline across all settings and ranks among the top methods, while several skill-evolution baselines regress on ScienceWorld and fall below both no-skill and initial-skill accuracy. Execution-grounded updates also achieve the best accuracy in most model-dataset settings and improve over the initial skill in nearly all ScienceWorld settings. These results suggest that grounding skill updates in execution helps retain useful skills while integrating new experience.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp