ROAug 12

RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models

arXiv:2510.0671011.914 citationsRSS
Predicted impact top 4% in RO · last 90 daysOriginality Incremental advance
AI Analysis

This framework addresses the fragmentation and inefficiency in RL training of VLA models, which is significant for researchers and practitioners in embodied intelligence seeking scalable and reproducible research.

This paper introduces RLinf-VLA, a unified framework for reinforcement learning (RL) of vision-language-action (VLA) models. The framework achieves a 1.61x-1.88x training speedup on ManiSkill and enables RL-trained models to achieve high success rates: 98.11% on 130 LIBERO tasks, 97.66% on 25 ManiSkill tasks, and 84.63% on 6 RoboTwin tasks.

Recent studies have demonstrated the potential of reinforcement learning (RL) to improve the task performance of vision-language-action (VLA) models through interaction. However, current efforts remain fragmented, lacking a unified platform for fair comparison across architectures and algorithms, as well as an efficient system design for scalable training. Therefore, we present RLinf-VLA, a unified and efficient framework for scalable RL training of VLA models. RLinf-VLA standardizes the integration of diverse VLA architectures, RL algorithms, and heterogeneous simulators through a unified interface, enabling extensibility and reproducibility. To improve efficiency, the framework adopts a flexible resource allocation architecture for rendering, inference, and training in RL pipelines. In particular, RLinf-VLA introduces a hybrid fine-grained pipeline allocation strategy that achieves a 1.61$\times$-1.88$\times$ training speedup on ManiSkill. Using this framework, RL-trained models achieve strong performance across embodied benchmarks, including 98.11% success on 130 LIBERO tasks, 97.66% success on 25 ManiSkill tasks, and 84.63% average success across 6 RoboTwin tasks. In addition, RLinf-VLA distills a set of effective practices for RL-based VLA training. We envision RLinf-VLA as a foundational framework for efficient, unified, and reproducible research in embodied intelligence.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes