ResearchResearch paperMultimodal Models · Speech & Audio · Robotics & Embodied AI1 source · Oct 8, 2026

Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation

Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify.

Key points

  • We present a perceptually grounded and semantics-aware benchmark that combines standardized model comparison, human-centered metric validation, and fine-grained semantic evaluation.
  • For the semantic-appropriateness category, we propose a new metric, Semantic Gesture Preservation (SGP), which measures how far semantic gestures in the ground truth are preserved in the generated gestures.
  • For this, we augment the BEAT2 dataset's annotations using a multi-modal LLM.
  • We then conduct a perceptual study where 101 participants score generated gestures among five dimensions, including human-likeness, motion diversity, absence of animation errors, speech timing and content match.

Sources (1)

  • [1]Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 10:43 AM
    Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify.
    We present a perceptually grounded and semantics-aware benchmark that combines standardized model comparison, human-centered metric validation, and fine-grained semantic evaluation.

Extractive summary: sentences quoted from the sources.

Related