ResearchResearch paperSpeech & Audio · Interpretability1 source · Oct 6, 2026

Feature Encoding in VAE-based Audio Decoders: Effects of Input, Depth and Distribution

Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood.

Key points

  • We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types.
  • For RAVE, we find that synthetic stimuli are encoded well across models and audio features (pitch |\r{ho}|=0.45, 5.1x the null, BPM |\r{ho}| = 0.76, 8.6x the null).
  • We find the best cross-layer cluster improves the strength (r = 0.65, p = 0.006) and prevalence (r = 0.75, p = 0.001) of BPM encoding when compared against the best whole layers within the same section, with no effect for joint encoding.
  • These findings advance the interpretability of neural audio models and inform targeted control strategies for neural synthesis.

Sources (1)

  • [1]Feature Encoding in VAE-based Audio Decoders: Effects of Input, Depth and Distribution
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 08:32 AM
    Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood.
    We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types.

Extractive summary: sentences quoted from the sources.

Related