ROAIJun 12

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

arXiv:2606.14409v112.21 citations
Predicted impact top 28% in RO · last 90 daysOriginality Synthesis-oriented
AI Analysis

This work provides a complete, reproducible pipeline for embodied AI, but the components are largely incremental integrations of existing methods.

Hy-Embodied-0.5-VLA (HyVLA-0.5) is an end-to-end robot learning system covering data collection, model design, pre-training, fine-tuning, RL post-training, and deployment. The system achieves a 78% success rate on real-world manipulation tasks, outperforming prior baselines by 15%.

In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pre-training and supervised fine-tuning, RL post-training, and real-world deployment. Each component serves a distinct role in this stack.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes