CLJun 22

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

arXiv:2606.2309216.0Has Code
Predicted impact top 61% in CL · last 90 daysOriginality Incremental advance
AI Analysis

It provides a new evaluation benchmark for a previously unexplored multimodal reasoning capability, highlighting a gap in current MLLMs.

The paper introduces PIVOTS, the first benchmark for evaluating fine-grained interpersonal relationship reasoning in multimodal large language models, and finds that current MLLMs significantly underperform humans, with the best model achieving only 55% accuracy on the main task.

Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently multimodal, it remains largely unexplored by existing multimodal large language models (MLLMs). To address this gap, we introduce PIVOTS, the first benchmark built from Social-IQ 2.0 and YouTube data to evaluate MLLMs' ability to predict bidirectional interpersonal relationship dimensions grounded in established psychology research. In addition, PIVOTS includes auxiliary tasks that assess models' ability to identify and leverage the critical visual cues underlying such predictions. We evaluate both proprietary and open-source MLLMs and conduct detailed ablation studies to analyze the effects of visual modalities and explicit social role information in conversational utterances. We further examine how joint and pairwise prediction settings benefit MLLMs in scoring bidirectional PIVOTS dimensions. Project page and resources: https://flynnzhangsx.github.io/PIVOTSBench/ .

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes