CLAIJan 10, 2025

Iconicity in Large Language Models

arXiv:2501.05643v13 citationsh-index: 3Digital Scholarship in the Humanities
Originality Incremental advance
AI Analysis

This addresses the problem of understanding how LLMs process linguistic features like iconicity, which is important for researchers in computational linguistics and AI, though it is incremental in exploring model capabilities.

The study investigated whether large language models (LLMs) can encode lexical iconicity by having GPT-4 generate iconic pseudowords and testing meaning guesses from humans and LLMs. Results showed that humans guessed meanings more accurately than with distant natural languages, and LLMs outperformed humans in this task.

Lexical iconicity, a direct relation between a word's meaning and its form, is an important aspect of every natural language, most commonly manifesting through sound-meaning associations. Since Large language models' (LLMs') access to both meaning and sound of text is only mediated (meaning through textual context, sound through written representation, further complicated by tokenization), we might expect that the encoding of iconicity in LLMs would be either insufficient or significantly different from human processing. This study addresses this hypothesis by having GPT-4 generate highly iconic pseudowords in artificial languages. To verify that these words actually carry iconicity, we had their meanings guessed by Czech and German participants (n=672) and subsequently by LLM-based participants (generated by GPT-4 and Claude 3.5 Sonnet). The results revealed that humans can guess the meanings of pseudowords in the generated iconic language more accurately than words in distant natural languages and that LLM-based participants are even more successful than humans in this task. This core finding is accompanied by several additional analyses concerning the universality of the generated language and the cues that both human and LLM-based participants utilize.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes