Model Releases
Evaluating VLMs on Multimodal Aristotelian Persuasion Tasks
arXiv:2608.01238v1 Announce Type: new Abstract: Vision Language Models (VLMs) have demonstrated exceptional performance across various tasks. However, they have not yet been thoroughly evaluated on mo
arXiv:2608.01238v1 Announce Type: new Abstract: Vision Language Models (VLMs) have demonstrated exceptional performance across various tasks. However, they have not yet been thoroughly evaluated on more complex tasks. The Persuasion Model, conceived by Aristotle, resembles a triangle shape, which highlights its inherent challenges related to personal biases. To assess the progress of VLMs on these complex tasks, we use the ImageArg datasets, focusing on the Logos, Ethos, and Pathos detection tasks. Our findings indicate that models from the Qwen family achieve improved F1 scores, with Qwen3 performing exceptionally well on the Logos and Pathos tasks, while Qwen2 exhibits competitive performance on the more complex Ethos detection task. We release the code to foster research in this direction.
Related
- Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
- OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning
- Efficient Task Adaptation in Large Language Models via Selective Parameter Optimization
Source: arXiv cs.CL | 2026-08-04