CLJun 12

Characterizing Cultural Localization in AI-Generated Stories

arXiv:2606.14626v113.8
Predicted impact top 75% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and developers of generative AI, this work highlights a limitation in cultural localization and potential biases in story generation.

The paper proposes a method to measure templated localization in AI-generated stories, finding that only 9-17% of vocabulary accounts for cross-national variation and that remaining narratives contain repeated templates. Cultural markers from 19 Global South countries are on average offensive.

The global use of artificial intelligence has increased interest in assessing the ability to generate culturally localized content, including stories. Cultural localization in stories often occurs through either templated localization -- the use of cultural markers (e.g., names, locations) in a generic narrative -- or holistic localization -- the variation of plots, values, and themes, in addition to cultural markers. We propose a method to measure the degree to which content was generated through templated localization. Specifically, we identify the lexical tokens that distinguish stories across nationalities and measure the similarity of the narratives that remain after removing them. In stories generated by five models on 125 topics for 193 nationalities, our method is able to detect that only a small subset (9-17%) of the vocabulary accounts for the variation across nationalities and that the narratives that remain after removing them contain repeated multi-word sequences, suggesting the presence of a shared culturally-agnostic narrative template. Finally, we characterize the cultural markers for their stereotypicality and offensiveness, finding that markers from 19 countries, mostly located in the Global South, are on average offensive.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes