CVROJun 15

MVOFormer: Flow-Semantic Transformer for Robust Monocular Visual Odometry

arXiv:2606.164748.6
Predicted impact top 58% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the need for robust monocular visual odometry in autonomous navigation, offering improved zero-shot generalization across diverse environments.

MVOFormer introduces a transformer framework for monocular visual odometry that combines geometric motion cues with semantic priors to improve robustness and cross-domain generalization. Without target-domain fine-tuning, it outperforms prior learning-based methods on multiple benchmarks including TartanAir, KITTI, TUM-RGBD, and ETH3D-SLAM.

Monocular visual odometry (MVO) is foundational to autonomous navigation and robotic localization. However, existing learning-based MVO approaches often struggle with either a lack of interpretable, complementary features or overly complex multi-stage architectures. These limitations inherently restrict their robustness and cross-domain generalization. In this work, we propose MVOFormer, a novel transformer framework for robust monocular visual odometry. Our architecture features a Flow-Semantic Dual Branch Encoder that synergizes dense geometric motion cues with object-centric semantic priors, explicitly distinguishing static structures from dynamic distractors. These representations are then fused by an Iterative Multimodal Decoder, enabling coarse-to-fine pose refinement while dynamically suppressing attention on unreliable regions. Extensive evaluations demonstrate that, without any target-domain fine-tuning, MVOFormer achieves superior zero-shot generalization and robustness, significantly outperforming prior learning-based frame-to-frame methods across diverse benchmarks including TartanAir, KITTI, TUM-RGBD, and ETH3D-SLAM.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes