PPO
Also known as: proximal policy optimization
21stories this week
21last 30 days
21all time
Timeline
- Oct 8, 2026 · Research paper · 2 sourcesViSkill: Reinforcing VLM Agents with Evolving Visual-Native SkillsWe propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents.
- Oct 8, 2026 · Research paper · 1 sourceReal-Time Motion Planning with Dynamic Hazards: Classical vs. Learning-Based MethodsWe study real-time motion planning in dynamic hazard fields through a controlled comparison between classical planning and learning-based methods.
- Oct 8, 2026 · Research paper · 1 sourceConstrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy GamesDeep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution.
- Oct 8, 2026 · Research paper · 1 sourceFOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation AlignmentVision-based reinforcement learning for robotic manipulation is sample-inefficient because RGB-D observations are high-dimensional and noisy.
- Oct 7, 2026 · Research paper · 1 sourceRFPO: Rectified Flow Policy Optimization for Embodied ControlFlow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost.
- Oct 7, 2026 · Research paper · 1 sourceA Closed-Loop Non-Asymptotic Convergence Analysis of PPO with Learned Critics and ClippingWe develop a non-asymptotic analysis of PPO-Clip as a closed-loop actor--critic system.
- Oct 7, 2026 · Research paper · 1 sourceDeadline-Aware Multi-Agent Reinforcement Learning for TSN-Based Vehicular Edge NetworksTo address these limitations, we propose a multi-agent reinforcement learning (MARL) approach for queue-level scheduling in TSN-enabled VEC.
- Oct 7, 2026 · Research paper · 1 sourceCOPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement LearningAsynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories.
- Oct 7, 2026 · Research paper · 1 sourceAsk the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber DefensePolicy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces.
- Oct 6, 2026 · Research paper · 1 sourceRLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithmsLLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn.
- Oct 6, 2026 · Research paper · 1 sourceTeaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task SelectionReinforcement Learning (RL) has enabled legged robots to perform a range of skills in single-task settings.
- Oct 6, 2026 · Research paper · 1 sourceFlashNeRD: Performance-First Contact-Rich Neural Robot DynamicsCompared with analytical physics, learned dynamics models promise robot simulation that is faster, inherently differentiable, and easily adaptable to real data.
- Oct 6, 2026 · Research paper · 1 sourceConvex-Concave Reinforcement LearningPolicy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today.
- Oct 6, 2026 · Research paper · 1 sourceLearning to Report Unsafe Tasks in a Multi-Agent GameWhen agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task.
- Oct 6, 2026 · Research paper · 1 sourceVETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM AgentsWe introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic.
- Oct 6, 2026 · Research paper · 1 sourceEnergy-Aware Path Following: Comparative Analysis of Reinforcement Learning and NMPC for Electric VehiclesPath-following control strategies typically follow the bi-objective optimization dilemma: minimizing deviations from a reference path while maintaining smooth speed profiles.
- Oct 6, 2026 · Research paper · 1 sourceLearning in Dreams, Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBAWe learn a structured, multi-agent world model of a complete ten-hero MOBA (206 units, every hero acting every tick, games of up to 6,000 ticks), train a policy only inside it with 1,400-tick free-running imagined episodes, and measure that policy in the real game against the opponent the game ships with.
- Oct 6, 2026 · Research paper · 1 sourceRevisiting Numerical Forecasting Models for Language-Based Trajectory PredictionLanguage-based trajectory predictors represent coordinates as discrete tokens and learn auxiliary tasks such as destination and group reasoning.
- Oct 6, 2026 · Research paper · 1 sourceLearned Adaptive Multiresolution Diffusion ImagingWe introduce Learned Adaptive Multiresolution Diffusion Imaging (Learned AMDI), which preserves the AMDI fixed-tree propagator and hierarchy constraints while replacing the post-propagation selector with a shared local policy trained by proximal policy optimization.
- Oct 6, 2026 · Research paper · 1 sourceIndependent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World ModelsFully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication.
- Oct 6, 2026 · Research paper · 1 sourceA GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement LearningWe introduce Hebero (Heterogeneous Benchmark for Robot Learning), a GPU-parallel Isaac Lab benchmark that enables efficient joint training and evaluation of a single policy across all 40 heterogeneous tasks.