Understanding Cross-Sensor Feature Variations for Generalizable 3D Perception
For practitioners of multi-modal 3D perception, this work provides a training-only regularization technique to improve cross-dataset generalization without target-domain samples.
This work addresses the performance degradation of radar-camera BEV 3D detectors across datasets by modeling source-domain variations in the frequency domain and using them to regularize the fusion space. The method achieves consistent improvements on cross-dataset detection between View-of-Delft and TJ4DRadSet over multiple BEV fusion backbones.
Radar-camera BEV perception often suffers from degraded performance when evaluated across datasets, as changes in driving scenes, sensor configurations, and environmental conditions can alter both the input observations and the internal fused representations. This work studies this issue from the perspective of source-domain variation modeling, aiming to improve the robustness of BEV-based 3D detectors without relying on target-domain samples. We introduce a framework that characterizes visual scene variations in the frequency domain and uses them to synthesize diverse source-domain views. By comparing the resulting fused BEV representations, the framework further captures how image-level variations influence multi-modal BEV features. These variation patterns are then used to regularize the detector, encouraging the learned fusion space to remain stable under latent scene changes. The proposed method is applied only during training and leaves the inference pipeline unchanged. Experiments on cross-dataset radar-camera 3D detection between View-of-Delft and TJ4DRadSet demonstrate consistent improvements over multiple BEV fusion backbones, and the gains remain effective when a small amount of target-domain data is available.