CVJun 22

Learning Stable Canonical Worlds for Novel View Synthesis and Beyond

arXiv:2606.2302712.9
Predicted impact top 34% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the problem of unstable and noisy scene representations in feed-forward Gaussian splatting for novel view synthesis and downstream perception tasks.

CanonicalGS introduces a feed-forward pipeline that maps multi-view observations into a stable scene-centric representation using uncertainty-aware fusion, achieving up to 2.5 dB PSNR improvement in novel view synthesis and 11% gain in semantic segmentation accuracy.

Feed-forward Gaussian splatting (FFGS) facilitates real-time novel view synthesis, yet current methods often remain tied to view-dependent predictions. As more input views are added, they may accumulate noisy or redundant evidence instead of converging to a stable scene representation. In this paper, we introduce CanonicalGS, a feed-forward pipeline that maps cluttered multi-view observations into a stable, scene-centric representation. CanonicalGS first extracts view-centric evidence from depth, semantic features, and uncertainty estimates, and then aggregates this evidence in a canonical latent world using uncertainty-aware fusion. By emphasizing reliable observations while suppressing uncertain or redundant ones, CanonicalGS produces representations that scale more effectively for novel view synthesis and transfer to downstream visual perception tasks. Experiments show up to a $2.5$ dB improvement in peak signal-to-noise ratio for synthesizing novel views and an $11\%$ gain in semantic segmentation accuracy.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes