Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.
Key points
- Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery.
- Therefore, we collect 10 human annotations for each of 3k real and synthetic responses derived from the CANDOR corpus.
- CVAM is supervised finetuned on synthesized aesthetic descriptions and labels, then optimized with Group Relative Policy Optimization on human judgments.
- Together, we demonstrate the importance of grounding voice aesthetics in human perception and propose a principled framework for human alignment.
Sources (1)
- [1]Conversational Voice Aesthetic Model with Reinforcement Learning from Human ListenersarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:13 PM
We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.
Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery.
Extractive summary: sentences quoted from the sources.