SV-TAD: Native Sparse Convs for Efficient Temporal Action Detection
To adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules.
Key points
- While effective for parameter-efficient training, existing adapters do not reduce inference-time computation, leaving scalability with respect to video length largely unaddressed.
- Token selection can reduce attention cost by pruning redundant tokens, but it breaks the spatial grid structure required by convolutional adapters.
- We address this by introducing native sparse 2D convolutions, a primitive that allows these adapters, for the first time, to operate directly and efficiently on dynamically pruned token sets.
- We integrate this primitive into SV-TAD, an adapter framework for temporal action detection, reducing VideoMAEv2-L computation by up to 64% and achieving 2.2x faster inference, while maintaining state-of-the-art accuracy on THUMOS-14 and ActivityNet-1.3.
Sources (1)
- [1]SV-TAD: Native Sparse Convs for Efficient Temporal Action DetectionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:31 AM
To adapt billion-parameter Vision Transformers for long-video understanding, recent methods freeze the backbone and train lightweight convolutional modules.
While effective for parameter-efficient training, existing adapters do not reduce inference-time computation, leaving scalability with respect to video length largely unaddressed.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
- Oct 8, 2026One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
- Oct 8, 2026Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
- Oct 7, 2026Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Oct 6, 2026EmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 6, 2026huggingface/transformers v5.19.0: Release v5.19.0