LGJun 18

Multi-Modal Contrastive Learning for Implicit Earth Embeddings via Location Tying

arXiv:2606.201678.7
Predicted impact top 50% in LG · last 90 daysOriginality Incremental advance
AI Analysis

For researchers working on spatial prediction with limited labelled data, this work shows that multimodal contrastive learning can match existing two-modality methods but reveals that the location encoder itself is a bottleneck, indicating incremental progress.

The paper proposes two multimodal contrastive learning architectures (MELT and SALT) for location encoders that extend beyond two modalities using unpaired geospatial data. Both methods match the performance of the strongest two-modality baseline (SATCLIP) across four downstream tasks, but increasing modalities does not consistently improve performance, suggesting the location encoder is the main limitation.

Spatial prediction tasks are often limited by a lack of high-quality labelled ground-truth observations. To overcome this challenge, self-supervised pre-training is a possible solution, with contrastive learning dominant for location encoders. Those approaches usually align geographic coordinates with just one additional modality. We propose two multimodal contrastive learning architectures: Multimodal Embedding via Location Tying (MELT) and Sequential Alternating Location Training (SALT). These architectures expand this framework beyond two modalities by utilising unpaired geospatial data. Both methods are technically viable and match the performance of the strongest two-modality baseline (SATCLIP) across four downstream tasks. However, increasing the number of modalities does not consistently improve performance, suggesting that the chosen location encoder is the main limitation - the contrastive objective reaches its peak early, regardless of modality diversity or pre-training volume. MELT provides more stable training than SALT and presents a stronger foundation for future scaling.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes