UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories.
ProofPaper ↗
Key points
- Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions.
- Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve.
- Specifically, the actor learns from environment rewards, while contrastive action feedback guides skill proposal learning.
- Since proposal-level feedback may suppress an otherwise appropriate edit operation when the proposed skill content scores poorly, we further apply skill-edit support regularization to preserve exploration.
Sources (2)
- [1]UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving PolicyHugging Face Daily Papers · Oct 7, 12:00 AM
In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories.
Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions.
- [2]UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving PolicyarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:41 PM · same content
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
- Oct 6, 2026RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
- Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
- Sep 29, 2026Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
- Sep 28, 2026Notes on NVIDIA Nemotron
- Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0