Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).
Key points
- We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance.
- We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model.
- On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines.
- LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.
Sources (1)
- [1]Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-SpeecharXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:12 AM
Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS).
We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance.
Extractive summary: sentences quoted from the sources.