ResearchResearch paperLarge Language Models · Applications · Evaluation & Benchmarks1 source · Oct 7, 2026

LLM Persuasion Is in the Eye of the Evaluation

Large language models (LLMs) have already been shown to match or exceed human experts in persuasion.

Key points

  • That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments.
  • In this study, we adapt nine published automated methods to a shared setup, run them on the same fifteen LLMs, and ask whether their rankings agree and why.
  • We find that the methods agree only weakly (mean Spearman $ρ= 0.25$).
  • More broadly, our results suggest that persuasion scores combine a model's ability to persuade with its willingness to do so.

Sources (1)

  • [1]LLM Persuasion Is in the Eye of the Evaluation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 03:19 PM
    Large language models (LLMs) have already been shown to match or exceed human experts in persuasion.
    That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments.

Extractive summary: sentences quoted from the sources.

Related