MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos
Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference.
ProofPaper ↗
Key points
- We redefine egocentric modality missingness as temporally localized sensor outages within untrimmed, multi-event observations, with whole-clip absence as the limiting case.
- We introduce MacJEPA, a missing-modality-robust Masked-context query JEPA that recognizes visual actions and acoustic events from supplied interval queries over audio-visual context.
- MacJEPA further repurposes masking in JEPA from a self-supervised pretext into a supervised robustness objective, aligning masked and clean latent representations of both multimodal content tokens and the task-conditioned queries.
- MacJEPA thus unifies strong full-input recognition with temporal missing-modality robustness in a single model operating on untrimmed multi-event videos.
Sources (1)
- [1]MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric VideosarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:42 AM
Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference.
We redefine egocentric modality missingness as temporally localized sensor outages within untrimmed, multi-event observations, with whole-clip absence as the limiting case.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 5, 2026Anthropic Subscriptions Offer 5x+ More Value Than OpenAI
- Oct 5, 2026Sharing AI progress in mathematics
- Oct 3, 2026[AINews] not much happened today
- Aug 26, 2026huggingface/transformers v5.16.0: Release: v5.16.0
- Jul 3, 2026huggingface/transformers v5.13.0: Release v5.13.0
- Jun 10, 2026huggingface/transformers v5.11.0: Release v5.11.0