→ Back to Home
Large Language Models

Researchers tested chatbots against elite fundraisers and debaters. Here's who won.

New research has demonstrated that large language model (LLM) powered chatbots possess a remarkable ability to persuade, even outperforming human experts in specific scenarios. The study, which is currently a preprint and awaiting peer review, involved two primary tests: one focusing on fundraising and another on competitive debating. In the fundraising experiment, professional fundraisers were pitted against commercial AI models to raise money for the nonprofit organization Save the Children. The chatbots proved to be more effective, not only convincing a higher percentage of participants to donate but also securing an average donation 13 percent larger than that achieved by human fundraisers. This success was partly attributed to the chatbots' capacity to generate messages nearly five times longer than those crafted by their human counterparts. The second test involved elite competitive debaters facing off against AI models like Claude and Google's Gemini 2.5 Pro. The AI models outperformed the debaters by a margin of 4.6 percentage points. Researchers noted that the chatbots deployed approximately 37 facts per conversation, lasting 15 to 20 minutes, whereas human debaters initially managed only about three. However, when the AI models were restricted to the same word count as the human debaters, their persuasive advantage largely disappeared, suggesting that sheer volume of information plays a significant role in their effectiveness. Despite their persuasive power, the study also highlighted varying levels of factual accuracy among the chatbots. An LLM-powered system designed to grade claims found that OpenAI's GPT 5.4 scored an average of 89 for truthfulness, while xAI's Grok scored a mere 26. Interestingly, greater truthfulness did not directly correlate with increased persuasiveness in the aggregate. The inaccuracies were often subtle, presenting plausible-sounding but factually incorrect information that would be difficult for a human to detect without immediate verification. This raises important questions about the ethical implications of using highly persuasive, yet potentially inaccurate, AI systems.
#llm performance#ai ethics#chatbot capabilities#persuasion#research
Read original source