Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders
Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space.
Key points
- We combine a TopK sparse autoencoder with route-specific supervision and cross-factor adversaries.
- Across frozen SPEAR and WavLM encoders, independent probes show factor-specific retention and suppression: linguistic information remains stronger in the linguistic route, while paralinguistic factors, including speaker identity, emotion, and prosody, are retained in the paralinguistic route and substantially reduced in the linguistic route.
- The route organisation learned on LibriSpeech persists on MSP-Podcast without representation-side retraining.
- These results show consistent route-selective separation across encoders, corpora, independent probes, and representation-level interventions.
Sources (1)
- [1]Disentangling Linguistic and Paralinguistic Information with Routed Sparse AutoencodersarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 08:10 PM
Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space.
We combine a TopK sparse autoencoder with route-specific supervision and cross-factor adversaries.
Extractive summary: sentences quoted from the sources.