CVJun 25

See & Sniff: Learning Visuo-Olfactory Representations

arXiv:2606.273077.9
Predicted impact top 64% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the lack of paired visuo-olfactory data, enabling multimodal learning with olfaction for the first time, though it is an incremental step using synthetic pairing.

The paper introduces a scalable visuo-olfactory dataset and a self-supervised framework for learning joint representations, achieving a 7% improvement in smell classification over smell-only baselines and enabling cross-modal retrieval and smell localization.

While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. We introduce SmellNet-V, a scalable visuo-olfactory dataset built on the insight that odor identity is largely invariant to visual transformations within a semantic category. This allows us to synthetically pair smell-only samples with semantically aligned in-the-wild web images, converting a unimodal olfactory dataset into a cross-modal benchmark without costly co-collection. Building on this dataset, we propose See & Sniff, a self-supervised framework that learns joint visuo-olfactory representations via dense local alignment and naturally produces smell saliency maps for spatial grounding of odor sources. We further introduce pixel-level smell localization task and a benchmark for evaluation. Our method surpasses smell-only baselines by 7% in smell classification from smell alone and generalizes to cross-modal retrieval and smell localization, establishing visuo-olfactory learning as a new direction in multimodal perception.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes