CLAICVJun 27, 2021

Visual Conceptual Blending with Large-scale Language and Vision Models

arXiv:2106.14127v11.015 citations
Originality Synthesis-oriented
AI Analysis

This work addresses the challenge of visual concept blending for AI and creative applications, but it is incremental as it applies existing models to a new task.

The paper tackled the problem of blending visual concepts by using large-scale language models to generate blend descriptions and text-based image generation models to create visual depictions, demonstrating quantitative and qualitative superiority over classical methods and prior models.

We ask the question: to what extent can recent large-scale language and image generation models blend visual concepts? Given an arbitrary object, we identify a relevant object and generate a single-sentence description of the blend of the two using a language model. We then generate a visual depiction of the blend using a text-based image generation model. Quantitative and qualitative evaluations demonstrate the superiority of language models over classical methods for conceptual blending, and of recent large-scale image generation models over prior models for the visual depiction.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes