AION
Research paperReinforcement Learning1 source · Oct 6, 2026

A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning

We introduce Hebero (Heterogeneous Benchmark for Robot Learning), a GPU-parallel Isaac Lab benchmark that enables efficient joint training and evaluation of a single policy across all 40 heterogeneous tasks.

Key points

  • GPU-parallel simulation provides abundant robot interaction, but existing benchmarks rarely combine this scale with heterogeneous manipulation tasks and standardized multi-task RL evaluation.
  • Scaling experiments show that increasing parallel replicas per task improves success under a fixed wall-clock budget.
  • To support learning with sparse rewards and limited demonstrations, we propose Demonstration-Guided Policy Optimization (DGPO), which reuses demonstrations for dense tracking rewards and asymmetric value learning.
  • Within DGPO framework, we introduce IW-ABC, which uses a lightweight per-task learning progress signal to coordinate adaptive behavior cloning (ABC), relaxing demonstration guidance with task progress, and importance weighting (IW), emphasizing lagging tasks in PPO updates.

Sources (1)

  • [1]A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning
    Hugging Face Daily Papers · Oct 6, 12:00 AM
    We introduce Hebero (Heterogeneous Benchmark for Robot Learning), a GPU-parallel Isaac Lab benchmark that enables efficient joint training and evaluation of a single policy across all 40 heterogeneous tasks.
    GPU-parallel simulation provides abundant robot interaction, but existing benchmarks rarely combine this scale with heterogeneous manipulation tasks and standardized multi-task RL evaluation.

Extractive summary: sentences quoted from the sources.