AION
Technique

PPO

Also known as: proximal policy optimization

21stories this week
21last 30 days
21all time

Timeline

  1. Oct 8, 2026 · Research paper · 2 sources
    ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
    We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.
  2. Oct 8, 2026 · Research paper · 1 source
    Real-Time Motion Planning with Dynamic Hazards: Classical vs. Learning-Based Methods
    We study real-time motion planning in dynamic hazard fields through a controlled comparison between classical planning and learning-based methods.
  3. Oct 8, 2026 · Research paper · 1 source
    Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games
    Deep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution.
  4. Oct 8, 2026 · Research paper · 1 source
    FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment
    Vision-based reinforcement learning for robotic manipulation is sample-inefficient because RGB-D observations are high-dimensional and noisy.
  5. Oct 7, 2026 · Research paper · 1 source
    RFPO: Rectified Flow Policy Optimization for Embodied Control
    Flow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost.
  6. Oct 7, 2026 · Research paper · 1 source
    A Closed-Loop Non-Asymptotic Convergence Analysis of PPO with Learned Critics and Clipping
    We develop a non-asymptotic analysis of PPO-Clip as a closed-loop actor--critic system.
  7. Oct 7, 2026 · Research paper · 1 source
    Deadline-Aware Multi-Agent Reinforcement Learning for TSN-Based Vehicular Edge Networks
    To address these limitations, we propose a multi-agent reinforcement learning (MARL) approach for queue-level scheduling in TSN-enabled VEC.
  8. Oct 7, 2026 · Research paper · 1 source
    COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning
    Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories.
  9. Oct 7, 2026 · Research paper · 1 source
    Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense
    Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces.
  10. Oct 6, 2026 · Research paper · 1 source
    RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms
    LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn.
  11. Oct 6, 2026 · Research paper · 1 source
    Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection
    Reinforcement Learning (RL) has enabled legged robots to perform a range of skills in single-task settings.
  12. Oct 6, 2026 · Research paper · 1 source
    FlashNeRD: Performance-First Contact-Rich Neural Robot Dynamics
    Compared with analytical physics, learned dynamics models promise robot simulation that is faster, inherently differentiable, and easily adaptable to real data.
  13. Oct 6, 2026 · Research paper · 1 source
    Convex-Concave Reinforcement Learning
    Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today.
  14. Oct 6, 2026 · Research paper · 1 source
    Learning to Report Unsafe Tasks in a Multi-Agent Game
    When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task.
  15. Oct 6, 2026 · Research paper · 1 source
    VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents
    We introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic.
  16. Oct 6, 2026 · Research paper · 1 source
    Energy-Aware Path Following: Comparative Analysis of Reinforcement Learning and NMPC for Electric Vehicles
    Path-following control strategies typically follow the bi-objective optimization dilemma: minimizing deviations from a reference path while maintaining smooth speed profiles.
  17. Oct 6, 2026 · Research paper · 1 source
    Learning in Dreams, Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBA
    We learn a structured, multi-agent world model of a complete ten-hero MOBA (206 units, every hero acting every tick, games of up to 6,000 ticks), train a policy only inside it with 1,400-tick free-running imagined episodes, and measure that policy in the real game against the opponent the game ships with.
  18. Oct 6, 2026 · Research paper · 1 source
    Revisiting Numerical Forecasting Models for Language-Based Trajectory Prediction
    Language-based trajectory predictors represent coordinates as discrete tokens and learn auxiliary tasks such as destination and group reasoning.
  19. Oct 6, 2026 · Research paper · 1 source
    Learned Adaptive Multiresolution Diffusion Imaging
    We introduce Learned Adaptive Multiresolution Diffusion Imaging (Learned AMDI), which preserves the AMDI fixed-tree propagator and hierarchy constraints while replacing the post-propagation selector with a shared local policy trained by proximal policy optimization.
  20. Oct 6, 2026 · Research paper · 1 source
    Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models
    Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication.
  21. Oct 6, 2026 · Research paper · 1 source
    A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning
    We introduce Hebero (Heterogeneous Benchmark for Robot Learning), a GPU-parallel Isaac Lab benchmark that enables efficient joint training and evaluation of a single policy across all 40 heterogeneous tasks.

Often appears with