TAM: Task-Aware Memory Distillation for Efficient Spatiotemporal Prediction
Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student.
Key points
- We propose TAM, a Task-Aware Memory Distillation framework that organizes a frozen teacher's knowledge into a bounded, retrievable history.
- Memory entries encode latent features, forecast changes, or flow residuals, while task-specific selection rules identify relevant historical references.
- We evaluate TAM on video prediction, weather forecasting, and traffic flow prediction across multiple teacher-student configurations.
- These results demonstrate the utility of historical teacher supervision across distinct forecasting tasks without additional student inference cost.
Sources (1)
- [1]TAM: Task-Aware Memory Distillation for Efficient Spatiotemporal PredictionarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 09:58 AM
Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student.
We propose TAM, a Task-Aware Memory Distillation framework that organizes a frozen teacher's knowledge into a bounded, retrievable history.
Extractive summary: sentences quoted from the sources.