AION
Research paperLarge Language Models · Speech & Audio1 source · Oct 8, 2026

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).

Key points

  • We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance.
  • We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model.
  • On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines.
  • LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.

Sources (1)

  • [1]Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:12 AM
    Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).
    We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance.

Extractive summary: sentences quoted from the sources.