LLMs Infer Cultural Context but Fail to Apply It When Responding
Identifies a gap between cultural knowledge and culturally adaptive language generation in LLMs, which is important for developers aiming to reduce cultural bias in AI systems.
LLMs can infer cultural background and recall relevant conventions, but often fail to apply this knowledge to adapt their responses (e.g., using local measurement units) unless explicitly prompted to do so sequentially. The study introduces the CAPRI dataset and finds that models' priors are not culture-neutral, sometimes aligning with the model's country of origin.
Recent work has shown that LLMs overrepresent dominant cultures, particularly Western ones, while marginalizing others. We investigate whether this affects models' ability to generate culturally adapted responses by evaluating their use of local measurement units based on the user's perceived cultural background. We introduce Cultural and Pragmatic Response Inference (CAPRI), a dataset of conversations with varying levels of cultural cues. Experiments with state-of-the-art LLMs show that models can infer cultural background and recall relevant conventions, but often fail to utilize the information to adapt their answers to the relevant cultural conventions, unless explicitly prompted to perform the tasks sequentially. We further evaluate adaptation to the interpretation of time and quantity expressions, two subjective language grounding dimensions that are affected by culture. We find that models increasingly adapt their answers as cultural cues accumulate, but their priors are not culture-neutral, sometimes aligning with the model's country of origin. Overall, CAPRI provides a resource for future research aimed at narrowing the gap between cultural knowledge and culturally adaptive language generation.