ROJul 9

FabriVLA: A Lightweight Vision-Language-Action Model for Precise Multi-Task Manipulation

arXiv:2607.0857511.4
Predicted impact top 28% in RO · last 90 daysOriginality Incremental advance
AI Analysis

This work provides a lightweight, efficient VLA model for precise multi-task manipulation, enabling deployment on resource-constrained robotic systems.

FabriVLA achieves a 90.0% tier-average success rate on the Meta-World MT50 benchmark, demonstrating that a compact 1B-scale VLM can deliver strong multi-task manipulation performance without relying on multi-billion parameter backbones.

We present FabriVLA, a lightweight Vision-Language-Action model for Precise Multi-Task Manipulation. FabriVLA combines an InternVL3.5 vision-language backbone with a flow-matching action head featuring gated self-attention across action tokens and shallow VLM layer fusion for enriched spatial context. The model is trained via single stage joint optimization from a pretrained VLM and randomly initialized action head. On the Meta-World MT50 benchmark spanning 50 diverse manipulation tasks, FabriVLA achieves a tier-average success rate of 90.0%, demonstrating that a compact VLA built on a 1B scale VLM can achieve strong performance without relying on multi billion parameter VLA backbones.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes