CLAIMAJan 4

Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage

arXiv:2601.01685v17 citationsHas Code
Originality Highly original
AI Analysis

This reveals a socio-technical vulnerability in LLM-based agent interactions, posing a risk for systems relying on autonomous information synthesis.

The paper tackles the problem of colluding LLM agents manipulating victim beliefs using only truthful evidence fragments, achieving attack success rates of up to 74.4% for proprietary models and 70.6% for open-weights models, with stronger reasoning capabilities increasing susceptibility.

As large language models (LLMs) transition to autonomous agents synthesizing real-time information, their reasoning capabilities introduce an unexpected attack surface. This paper introduces a novel threat where colluding agents steer victim beliefs using only truthful evidence fragments distributed through public channels, without relying on covert communications, backdoors, or falsified documents. By exploiting LLMs' overthinking tendency, we formalize the first cognitive collusion attack and propose Generative Montage: a Writer-Editor-Director framework that constructs deceptive narratives through adversarial debate and coordinated posting of evidence fragments, causing victims to internalize and propagate fabricated conclusions. To study this risk, we develop CoPHEME, a dataset derived from real-world rumor events, and simulate attacks across diverse LLM families. Our results show pervasive vulnerability across 14 LLM families: attack success rates reach 74.4% for proprietary models and 70.6% for open-weights models. Counterintuitively, stronger reasoning capabilities increase susceptibility, with reasoning-specialized models showing higher attack success than base models or prompts. Furthermore, these false beliefs then cascade to downstream judges, achieving over 60% deception rates, highlighting a socio-technical vulnerability in how LLM-based agents interact with dynamic information environments. Our implementation and data are available at: https://github.com/CharlesJW222/Lying_with_Truth/tree/main.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes