CL AI MASep 5, 2025

Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate

Andrea Wynn, Harsh Satija, Gillian Hadfield

arXiv:2509.05396v226.644 citationsh-index: 8

Originality Incremental advance

AI Analysis

This identifies critical failure modes in multi-agent debate for AI reasoning, cautioning against naive applications that could degrade performance.

The paper investigates how diversity in model capabilities affects multi-agent debate, finding that debate can reduce accuracy over time as models shift from correct to incorrect answers to agree with peers, even when stronger models outnumber weaker ones.

While multi-agent debate has been proposed as a promising strategy for improving AI reasoning ability, we find that debate can sometimes be harmful rather than helpful. Prior work has primarily focused on debates within homogeneous groups of agents, whereas we explore how diversity in model capabilities influences the dynamics and outcomes of multi-agent interactions. Through a series of experiments, we demonstrate that debate can lead to a decrease in accuracy over time - even in settings where stronger (i.e., more capable) models outnumber their weaker counterparts. Our analysis reveals that models frequently shift from correct to incorrect answers in response to peer reasoning, favoring agreement over challenging flawed reasoning. We perform additional experiments investigating various potential contributing factors to these harmful shifts - including sycophancy, social conformity, and model and task type. These results highlight important failure modes in the exchange of reasons during multi-agent debate, suggesting that naive applications of debate may cause performance degradation when agents are neither incentivised nor adequately equipped to resist persuasive but incorrect reasoning.

View on arXiv PDF

Similar