ROAICLJun 23

RAVEN: Long-Horizon Reasoning & Navigation with a Visuo-Spatio-Temporal Memory

arXiv:2606.2520612.4
Predicted impact top 27% in RO · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for scalable, fine-grained memory in long-term robot deployment, enabling efficient semantic, spatial, and temporal retrieval without lossy captioning.

RAVEN introduces an agentic memory system that stores visual embeddings with pose and time in a vector database for long-horizon robotic question answering and navigation, surpassing caption-based systems and matching frontier VLMs at 10x lower retrieval cost.

Long-term robot deployment requires a compact and scalable memory that preserves fine-grained visual semantics, grounds observations in space and time, and enables efficient storage and retrieval. In this paper, we propose RAVEN, an agentic memory system for long-horizon robotic question answering and navigation. RAVEN stores visual embeddings with pose and time in a vector database, and grounds retrieval in a spatial map to answer queries and navigate to goals. By operating directly on visual embeddings, RAVEN avoids lossy image-to-text captioning and enables accurate semantic, spatial, and temporal retrieval at scale. Across several simulated and real-world video question-answering benchmarks, RAVEN consistently surpasses caption-based memory systems and matches frontier VLMs on long-horizon tasks at 10$\times$ lower retrieval cost. Finally, we instantiate RAVEN on a Unitree Go1 robot for the task of long-horizon navigation for natural language goal-reaching, and show successful deployment over several large indoor environments.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes