ROAIJun 18

Vesta: A Generalist Embodied Reasoning Model

arXiv:2606.2090535.9
Predicted impact top 1% in RO · last 90 daysOriginality Highly original
AI Analysis

This work demonstrates that a single generalist model can match or exceed specialist models in embodied AI, offering a scalable and efficient alternative for open-world robotics.

Vesta is a unified embodied generalist model that consolidates localization, spatial reasoning, navigation, and long-horizon planning into a single foundation model, outperforming individual SOTA baselines by >20% and an ensemble of per-category-best baselines by >10% across diverse benchmarks, and improving real-world robotic task success by >35%.

Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deploying a multi-model stack is computationally expensive and prone to cascading errors. We present Vesta, a unified embodied generalist that consolidates these capabilities into a single foundation model. Our approach combines a diverse and massive curated corpus designed to induce spatial grounding and a simple multimodal memory harness that enables reasoning over extended time horizons. Across diverse benchmarks, Vesta on average beats individual SOTA baselines by >$20\%$ and beats an ensemble of per-category-best baselines by $>10\%$ -- thus demonstrating that a generalist model can match or exceed specialists. On real-world robotic tasks requiring memory and reasoning, Vesta improves task success by >35\%. Our work thus demonstrates that a single generalist is a feasible, scalable, and arguably preferable alternative to combining specialists.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes