Images Amplify Misinformation Spread in Vision-Language Models
New research presented at the International AAAI Conference on Web and Social Media has uncovered a critical vulnerability in Vision-Language Models (VLMs): the presence of images amplifies the sharing of misinformation. The study found that images increase the resharing rates of false news by 14.5% and true news by 5.3%, demonstrating a clear bias towards visual content that can exacerbate the spread of inaccurate information.
The researchers developed a novel prompting strategy to bypass VLMs' default refusals to engage with controversial news, enabling them to analyze resharing decisions across various topics and elicited traits. This allowed for a comprehensive evaluation of how images influence VLMs' behavior. The findings suggest that VLMs replicate human-like biases in their response to visual cues, underscoring the complex ethical and societal challenges associated with deploying multimodal AI systems in information-sharing environments.
Furthermore, the study explored the impact of "persona conditioning," revealing that certain traits, such as those associated with the Dark Triad, can further amplify the resharing of false news. Conversely, Republican-aligned profiles showed reduced sensitivity to veracity. Among the state-of-the-art VLMs evaluated, Claude-3-Haiku emerged as the most robust against visual misinformation. These results emphasize the urgent need for enhanced evaluation frameworks and mitigation strategies to address visual influence and persona-driven variability in multimodal AI, particularly as these systems increasingly shape public discourse.
Read original source