Where did the ambiguity go? Examining how multimodal models interpret polysemous words
For researchers and developers of multimodal foundation models, this reveals a modality gap in meaning representation, highlighting that models' understanding of polysemy does not transfer equally across text and image generation.
This paper investigates how text-to-image and text-generation models handle polysemous words, finding that images exhibit far less semantic diversity than text (normalized entropy 0.10 vs. 0.25), and both are less varied than human imagination (0.47). Models also overestimate the diversity of their own outputs when asked to predict distributions.
Human language is highly polysemous. Many common words (e.g., 'bank' or 'palm') carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand this multiplicity of meaning, but much less is known about how polysemy surfaces in other modalities such as images. We study this across 17 text-to-image and 15 text-generation models by giving each a polysemous word with no context to fix its meaning and measuring which senses are produced over many samples. We find a clear multimodal gap, where within every model family, generated images settle on far fewer senses than generated sentences (normalized entropy 0.10 vs. 0.25), and both are far less varied than what people imagine for the same words (normalized entropy 0.47). However, when we instead ask a model to list how often it would generate outputs corresponding to each possible meaning of a word, it predicts distributions that are more diverse than the actual space of outputs. These results reveal a multimodal gap in how foundation models express meaning, and how their understanding may not transfer faithfully nor equally across modalities.