SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context
For researchers building personalized AI assistants, this benchmark identifies a critical gap in inferring revealed preferences from multimodal social-media traces rather than relying on explicit user statements.
SocialPersona benchmarks whether multimodal LLMs can infer user preferences from longitudinal social-media timelines and use them in dialogue. Experiments show models identify broad interests but struggle with fine-grained and recent preferences, especially when personalizing dialogue, revealing a key challenge in cross-modal user modeling.
Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogue? More comprehensive personalization demands a harder capability -- inferring what users care about from the multimodal traces they naturally leave behind. We introduce SocialPersona, a benchmark for evaluating whether multimodal large language models (MLLMs) can recover revealed preferences from longitudinal social-media timelines and use them in dialogue. Built from longitudinal timelines of 171 everyday, non-promotional social-media users, SocialPersona contains text, images, timestamps, and 2,597 human-verified preference tags across seven interest domains, separating stable interests from recent interests. It supports two tasks: constructing structured user profiles from multimodal context and generating responses aligned with inferred profiles. Experiments with proprietary and open-weight MLLMs show that models can identify broad interest domains, yet their performance drops on fine-grained and recent interests and degrades further when inferred profiles must be used to personalize dialogue. Together with evidence that text and images provide complementary preference signals, these results indicate that robust cross-modal, long-horizon user modeling remains a key challenge, and that SocialPersona can help measure and advance progress toward assistants that infer and act on revealed preferences.