ROJun 12

FloVerse: Floor Plan-Guided Multi-Modal Navigation

arXiv:2606.14267v18.6
Predicted impact top 50% in RO · last 90 daysOriginality Incremental advance
AI Analysis

This work provides a unified benchmark and method for floor plan-guided navigation, enabling agents to leverage spatial priors for more efficient navigation in unseen environments.

FloVerse introduces a unified floor plan-guided navigation task covering PointNav, ObjectNav, and ImageNav, with a large-scale dataset of 1.6K scenes and 240K expert trajectories. The proposed ThreeDiff policy achieves improved navigation performance across all goal modalities, demonstrating that floor plan priors enhance spatial reasoning.

Floor plans encapsulate compact spatial priors, enabling agents to navigate unseen scenes more efficiently. While prior work has explored floor plan-guided navigation, it has focused mainly on PointNav and a limited set of environments. To bridge this gap, we introduce FloVerse, a new task for floor plan-guided embodied navigation that unifies PointNav, ObjectNav, and ImageNav. To support FloVerse, we assemble FloVerse-1.6K, a large-scale dataset of 1.6K scenes from HM3D and Gibson 4+, paired with corresponding floor plans, comprising 240K expert trajectories and 12M RGBD frames. We further propose ThreeDiff, a two-stage imitation learning policy comprising a planner, a diffusion-based multimodal goal-reasoning module trained via masked-modality modeling, and a refiner, a depth-based trajectory-refinement module for safe execution. Extensive experiments demonstrate that (1) floor-plan priors improve navigation performance across all goal modalities, and (2) ThreeDiff implicitly captures spatial information from floor plans. These results underscore the effectiveness of spatial priors and validate our proposed unified approach for floor plan-guided embodied navigation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes