AIIVJul 8

Wavelet Phase Diffusion for Structurally and Semantically Consistent Sim-to-Real Translation

arXiv:2607.2162812.8
Predicted impact top 45% in AI · last 90 daysOriginality Incremental advance
AI Analysis

This work improves sim-to-real translation for autonomous driving, enabling more realistic and consistent synthetic-to-real image translation without expensive control modules or paired data.

Wavelet Phase Diffusion achieves structurally and semantically consistent sim-to-real translation by operating in the Dual-Tree Complex Wavelet Packet Transform domain with Low-Frequency Randomization, outperforming prior methods in realism and semantic consistency on vKITTI→KITTI and reducing VLM planner ADE and FDE by 5.4% and 5.1% on CARLA video translation.

Simulation-to-reality translation must bridge the appearance gap between synthetic and real domains while preserving structural and semantic consistency. Conditioning-based methods achieve spatial alignment but introduce computationally expensive control modules. Paired-data methods achieve realism but rely on complex synthesis pipelines, often altering scene geometry and semantics. Training-free editing methods avoid both constraints but lack a learned appearance prior, limiting their perceptual quality. Recently proposed phase-preserving diffusion presents a promising alternative, but Fourier-domain formulations are constrained by global spectral coupling. This coupling induces spatial artifacts such as ringing and boundary leakage, thereby degrading structural and semantic consistency. We introduce Wavelet Phase Diffusion, which addresses this through two components. First, we operate in the Dual-Tree Complex Wavelet Packet Transform domain, whose localized wavelet packets enable spatially adaptive phase injection without global spectral interference. Second, Low-Frequency Randomization (LFR) replaces the low-frequency packet, decoupling the model from the synthetic illumination prior and enabling in-distribution real-world appearance. Both components train on unpaired open-domain data, and introduce negligible inference overhead. The spatial locality further enables instance-level translation, where individual objects or regions are translated to photorealistic appearance independently while the surrounding scene remains untranslated. On vKITTI $\to$ KITTI image translation, ours outperforms prior methods in realism and semantic consistency while maintaining competitive structural alignment. For CARLA video translation, ours approaches the realism of paired-data methods while reducing VLM planner ADE and FDE by $5.4\%$ and $5.1\%$, respectively.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes