AION
Research paperSpeech & Audio · Large Language Models · Multimodal Models1 source · Oct 7, 2026

Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners

We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.

Key points

  • Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery.
  • Therefore, we collect 10 human annotations for each of 3k real and synthetic responses derived from the CANDOR corpus.
  • CVAM is supervised finetuned on synthesized aesthetic descriptions and labels, then optimized with Group Relative Policy Optimization on human judgments.
  • Together, we demonstrate the importance of grounding voice aesthetics in human perception and propose a principled framework for human alignment.

Sources (1)

  • [1]Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:13 PM
    We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.
    Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery.

Extractive summary: sentences quoted from the sources.