ROJul 17

Difference-Based Relational Learning for Zero-Shot Object-Goal Visual Navigation With Direct Sim-to-Real Transfer

arXiv:2607.156424.4
Predicted impact top 69% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For embodied AI researchers, this work addresses the sim-to-real gap in visual navigation with a method that generalizes across appearance and FoV variations.

T-DRN improves zero-shot object-goal visual navigation by using a Siamese difference-based feature extractor and dual-frame temporal buffer, achieving better success rates in simulation and robust real-world performance with direct sim-to-real transfer.

End-to-end deep reinforcement learning (DRL) for zero-shot object-goal visual navigation remains challenged by the sim-to-real gap, particularly variations in object appearance and restricted camera field-of-view (FoV). This letter proposes a Temporal Difference-Relational Network (T-DRN) for robust zero-shot sim-to-real transfer. T-DRN combines a Siamese difference-based feature extractor, which computes relational difference between the target and observed objects to produce domain-independent representations, with a dual-frame temporal buffer that preserves short-term object continuity under narrow FoV. Extensive experiments in AI2-THOR demonstrate that T-DRN improves zero-shot generalization in terms of success rates over strong baselines. Furthermore, T-DRN is systematically validated on a physical wheeled robot, demonstrating robust performance under real sensing and actuation constraints and supporting the feasibility of direct sim-to-real transfer.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes