RO AI LGFeb 24, 2025

TDMPBC: Self-Imitative Reinforcement Learning for Humanoid Robot Control

Zifeng Zhuang, Diyuan Shi, Runze Suo, Xiao He, Hongyin Zhang, Ting Wang, Shangke Lyu, Donglin Wang

arXiv:2502.17322v19.54 citationsh-index: 10Has Code

Originality Incremental advance

AI Analysis

This addresses the problem of efficient robot control for robotics researchers, though it appears incremental as it builds on existing RL methods.

The paper tackles the challenge of reinforcement learning in high-dimensional spaces for humanoid robot control by proposing a self-imitative RL framework that emphasizes task-relevant trajectories, achieving a 120% performance improvement on HumanoidBench with only 5% extra computation overhead.

Complex high-dimensional spaces with high Degree-of-Freedom and complicated action spaces, such as humanoid robots equipped with dexterous hands, pose significant challenges for reinforcement learning (RL) algorithms, which need to wisely balance exploration and exploitation under limited sample budgets. In general, feasible regions for accomplishing tasks within complex high-dimensional spaces are exceedingly narrow. For instance, in the context of humanoid robot motion control, the vast majority of space corresponds to falling, while only a minuscule fraction corresponds to standing upright, which is conducive to the completion of downstream tasks. Once the robot explores into a potentially task-relevant region, it should place greater emphasis on the data within that region. Building on this insight, we propose the $\textbf{S}$elf-$\textbf{I}$mitative $\textbf{R}$einforcement $\textbf{L}$earning ($\textbf{SIRL}$) framework, where the RL algorithm also imitates potentially task-relevant trajectories. Specifically, trajectory return is utilized to determine its relevance to the task and an additional behavior cloning is adopted whose weight is dynamically adjusted based on the trajectory return. As a result, our proposed algorithm achieves 120% performance improvement on the challenging HumanoidBench with 5% extra computation overhead. With further visualization, we find the significant performance gain does lead to meaningful behavior improvement that several tasks are solved successfully.

View on arXiv PDF Code

Similar