Chatbots Outperform Human Fundraisers and Debaters in Persuasion Study
New research has demonstrated the surprising persuasive power of artificial intelligence, with chatbots proving more effective than their human counterparts in both fundraising and debating scenarios. The study, which is currently a preprint and awaiting peer review, pitted commercial AI models against professional human fundraisers working for the non-profit organization Save the Children.
Over more than a thousand conversations with potential donors, the AI model Claude Opus 4.6 exhibited remarkable success. It was nearly three times more effective at securing donations compared to the human fundraisers. Furthermore, the donations elicited by Claude were, on average, 13% higher.
The researchers also conducted a separate experiment to assess persuasiveness in a debate setting. Here, AI models, including Claude and Google's Gemini 2.5 Pro, outperformed elite competitive debaters by a margin of 4.6 percentage points.
According to Kobi Hackenburg, a PhD student at the University of Oxford and the lead author of the study, the chatbots' advantage stemmed from their capacity to rapidly process and articulate claims. The AI models were able to present significantly more information, spewing out messages nearly five times the length of those from human fundraisers. In the debate trials, they deployed approximately 37 facts per conversation, lasting 15 to 20 minutes, whereas human debaters managed only about three facts initially.
However, the study introduced a critical caveat: when the AI models were restricted to the same word count as the human debaters, their persuasive edge largely disappeared. This suggests that the chatbots' superior performance was heavily reliant on their ability to generate and deliver a high volume of information, rather than a more nuanced understanding or emotional appeal.
An interesting finding also emerged regarding the accuracy of the information provided by the chatbots. While greater truthfulness did not directly correlate with persuasiveness, the study found varying degrees of accuracy among different models. OpenAI's GPT 5.4 scored an average of 89 in truthfulness, while xAI's Grok scored 26, as graded by an LLM-powered system. The inaccuracies were often subtle, such as citing reports that sounded plausible but did not actually exist.
Read original source