ResearchResearch paperMultimodal Models1 source · Oct 7, 2026

MUNITE: Unified Multimodal Latent Inference for Any-to-Any Multimodal Generation

We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence.

Key points

  • Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation.
  • A shared latent sample captures variation that must remain consistent across generated targets, while modality-specific generative decoders model the remaining uncertainty independently.
  • To learn these conditional distributions from incomplete training examples, we extend conditional flow matching through self-distillation: predictions conditioned on richer available observations supervise the same model conditioned on smaller subsets at the same intermediate latent state.
  • Across PolyMNIST-D-Q, FFHQ64, and image-text-audio, MUNITE achieves competitive or better generation quality and source-target alignment, with higher joint-generation coherence.

Sources (1)

  • [1]MUNITE: Unified Multimodal Latent Inference for Any-to-Any Multimodal Generation
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 11:27 AM
    We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence.
    Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation.

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 7, 2026MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
  2. Oct 7, 2026Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
  3. Oct 7, 2026On-Policy Distillation Teaches New Skills but Not New Knowledge
  4. Oct 7, 2026UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
  5. Oct 6, 2026Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
  6. Sep 29, 2026[AINews] AMD buys World Labs for $8.2B, as Atlas solves sparse reconstruction problem for robotics, design and more

Related