Taming Toxic Talk: Using chatbots to intervene with users posting toxic comments

Jeremy Foote, Deepak Kumar, Bedadyuti Jha, Ryan Funkhouser, Loizos Bitsikokos, Hitesh Goel, Hsuen-Chi Chiu

arXiv:2601.20100v1

Originality Synthesis-oriented

AI Analysis

This addresses the problem of toxic behavior in online communities by testing a non-punitive intervention, but it is incremental as it builds on prior lab findings without achieving practical impact.

The study explored using generative AI chatbots to conduct rehabilitative conversations with users who posted toxic content online, finding that while many participants engaged sincerely and expressed remorse, there was no significant reduction in toxic behavior compared to a control group.

Generative AI chatbots have proven surprisingly effective at persuading people to change their beliefs and attitudes in lab settings. However, the practical implications of these findings are not yet clear. In this work, we explore the impact of rehabilitative conversations with generative AI chatbots on users who share toxic content online. Toxic behaviors -- like insults or threats of violence, are widespread in online communities. Strategies to deal with toxic behavior are typically punitive, such as removing content or banning users. Rehabilitative approaches are rarely attempted, in part due to the emotional and psychological cost of engaging with aggressive users. In collaboration with seven large Reddit communities, we conducted a large-scale field experiment (N=893) to invite people who had recently posted toxic content to participate in conversations with AI chatbots. A qualitative analysis of the conversations shows that many participants engaged in good faith and even expressed remorse or a desire to change. However, we did not observe a significant change in toxic behavior in the following month compared to a control group. We discuss possible explanations for our findings, as well as theoretical and practical implications based on our results.

View on arXiv PDF

Similar