Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges.
Key points
- Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction.
- We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD.
- CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher.
- Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective.
Sources (1)
- [1]Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal GenerationHugging Face Daily Papers · Sep 29, 12:00 AM
Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges.
Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction.
Extractive summary: sentences quoted from the sources.