ResearchResearch paperReinforcement Learning · Large Language Models · Robotics & Embodied AI2 sources · Oct 7, 2026

UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy

In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories.

Key points

  • Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions.
  • Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve.
  • Specifically, the actor learns from environment rewards, while contrastive action feedback guides skill proposal learning.
  • Since proposal-level feedback may suppress an otherwise appropriate edit operation when the proposed skill content scores poorly, we further apply skill-edit support regularization to preserve exploration.

Sources (2)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
  2. Oct 6, 2026RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
  3. Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more
  4. Sep 29, 2026Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
  5. Sep 28, 2026Notes on NVIDIA Nemotron
  6. Aug 10, 2026huggingface/transformers v5.15.0: Release: v5.15.0

Related