SIAIJun 3

Towards Multi-Agent-Simulation-Based Community Note Evaluation

arXiv:2606.1826810.7
Predicted impact top 11% in SI · last 90 daysOriginality Incremental advance
AI Analysis

For social media platforms relying on community-based fact-checking, this work provides a scalable automated evaluation method to reduce reliance on slow human ratings.

The paper tackles the challenge of delay and low-ratio of cross-consensus community fact-checks by proposing MultiCom, a persona-guided multi-agent rating framework that simulates diverse raters. MultiCom achieves 84.7% average accuracy (balanced accuracy 68.3%, macro-F1 60.1%) on community note evaluation.

Community-based fact-checking that relies on cross-consensus is expanding rapidly on social media platforms. However, the delay and low-ratio of cross-consensus community fact-checks rated by human contributors remains a significant challenge. To address this, we first created ComRate, a large-scale dataset comprising 2.5 million community notes and over 209 million ratings sourced from $\mathbb{X}$. We then propose MultiCom, a persona-guided multi-agent rating framework for community note evaluation. MultiCom simulates diverse rater population by clustering contributors in a matrix-factorized rater space and prompting persona agents to generate structured assessments based on the official community notes rating schema. These agents output structured and explainable judgments, such as confidence, agreement signals and reasons. An out-of-fold calibrated aggregation algorithm combines features such as raw votes and diagnostic reason signals for reliable prediction. Extensive evaluations demonstrate that MultiCom outperforms alternative methods, achieving an average accuracy of 84.7% (balanced accuracy 68.3%, macro-F1 60.1%) on the evaluation set.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes