LLM Persuasion Is in the Eye of the Evaluation
Large language models (LLMs) have already been shown to match or exceed human experts in persuasion.
ProofPaper ↗
Key points
- That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments.
- In this study, we adapt nine published automated methods to a shared setup, run them on the same fifteen LLMs, and ask whether their rankings agree and why.
- We find that the methods agree only weakly (mean Spearman $ρ= 0.25$).
- More broadly, our results suggest that persuasion scores combine a model's ability to persuade with its willingness to do so.
Sources (1)
- [1]LLM Persuasion Is in the Eye of the EvaluationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:19 PM
Large language models (LLMs) have already been shown to match or exceed human experts in persuasion.
That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments.
Extractive summary: sentences quoted from the sources.