Steerspeech: Activation Steering For Emotion Control In Generated Speech
We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations.
ProofPaper ↗
Key points
- Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training.
- For each target emotion we train a lightweight low-rank transform, using a multi-expert objective that encourages monotonic emotion control while preserving speaker identity and linguistic content, constraining steering drift, and keeping the TTS backbone frozen.
- To optimize through discrete speech tokens, we introduce a two-pass generation-and-replay pipeline using a straight-through estimator to backpropagate expert supervision through sampled tokens.
- Objective and subjective evaluations with Qwen3-TTS across seen, unseen, and accented speakers show stronger continuous emotion control with limited speaker and content degradation.
Sources (1)
- [1]Steerspeech: Activation Steering For Emotion Control In Generated SpeecharXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:56 PM
We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations.
Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
- Oct 5, 2026perplexity-ai/pplx-decider-v1.1-27b
- Oct 4, 2026nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
- Oct 2, 2026alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF
- Oct 1, 2026nvidia/PixelUMM