ROAIJul 3

HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control

arXiv:2607.0344918.8
Predicted impact top 10% in RO · last 90 daysOriginality Highly original
AI Analysis

This work addresses the frequency-competence paradox in VLA models for robotic manipulation, enabling real-time control with long-term reasoning.

HiMe introduces a hierarchical embodied memory framework that decouples robotic control into high-frequency execution, working memory, and long-term planning, achieving significant improvements in long-horizon task success rates and enabling self-correction based on human preferences.

Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance on immediate observations. Existing solutions face a ''frequency-competence paradox,'' where stronger reasoning models are too slow for real-time control, while faster models lack sufficient reasoning capabilities. To resolve this architectural misalignment, we propose HiMe, a Hierarchical Embodied Memory framework that decouples embodied intelligence into a high-frequency Executor for execution, a Sentry for working memory, and a Planner for long-term strategy. We also introduce a dynamic knowledge system based on cross-modal semantic schemas and active management mechanisms, allowing robots to maintain memory plasticity through ''Add, Update, and Delete'' operations. This hierarchical design effectively balances the conflict between real-time execution and slow thinking planning, significantly improving success rates in long-horizon tasks. Experiments demonstrate that this approach not only outperforms flat memory baselines but also exhibits the novel ability to self-correct its internal knowledge based on human preferences.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes