Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture Generation
Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify.
ProofPaper ↗
Key points
- We present a perceptually grounded and semantics-aware benchmark that combines standardized model comparison, human-centered metric validation, and fine-grained semantic evaluation.
- For the semantic-appropriateness category, we propose a new metric, Semantic Gesture Preservation (SGP), which measures how far semantic gestures in the ground truth are preserved in the generated gestures.
- For this, we augment the BEAT2 dataset's annotations using a multi-modal LLM.
- We then conduct a perceptual study where 101 participants score generated gestures among five dimensions, including human-likeness, motion diversity, absence of animation errors, speech timing and content match.
Sources (1)
- [1]Perceptually Grounded and Semantics-Aware Evaluation for Holistic Co-Speech Gesture GenerationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 10:43 AM
Holistic and semantics-aware co-speech gesture generation has advanced rapidly, yet evaluation remains behind: objective metrics do not consistently reflect human perception, and semantic appropriateness remains difficult to quantify.
We present a perceptually grounded and semantics-aware benchmark that combines standardized model comparison, human-centered metric validation, and fine-grained semantic evaluation.
Extractive summary: sentences quoted from the sources.