CVLGJul 20

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

arXiv:2607.182379.0
Predicted impact top 12% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers needing context-aware perceptual similarity metrics, TPIPS provides a method to condition on specific semantic aspects, outperforming existing VLMs and enabling new applications.

Existing perceptual similarity metrics collapse context-dependent visual similarity into a single scalar. The authors introduce a large-scale dataset of human similarity judgments over image triplets with free-form semantic aspects, fine-tune a VLM to produce TPIPS, which aligns more closely with human perception and generalizes beyond training, enabling text-guided retrieval and fine-grained generative model evaluation.

Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to human annotators' consensus. Leveraging our data, we fine-tune a VLM to produce our Text-Prompted Image Perceptual Similarity (TPIPS) metric, capturing multiple senses of visual similarity depending on the specified text prompt. We demonstrate that TPIPS aligns more closely with human perception and generalizes reliably beyond the training distribution. Finally, we show that TPIPS unlocks new capabilities in text-guided retrieval, compositional search, and the fine-grained evaluation of generative models. Our code, data, and trained models are at https://peterwang512.github.io/TPIPS

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes