AION
Research paperMultimodal Models · Image, Video & 3D Generation · Reinforcement Learning1 source · Sep 29, 2026

Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation

Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges.

Key points

  • Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction.
  • We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD.
  • CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher.
  • Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective.

Sources (1)

  • [1]Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
    Hugging Face Daily Papers · Sep 29, 12:00 AM
    Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges.
    Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction.

Extractive summary: sentences quoted from the sources.