No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping
Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.
Key points
- And it must run online: each frame emitted from audio observed up to the current time, at interactive rates.
- Audio-conditioned facial motion occupies a comparatively low-dimensional manifold, a regime where a single-pass GAN suffices.
- We show that a causal, time-invariant generator driven by i.i.d. noise cannot suppress its output spectrum over a band without collapsing its per-step innovation.
- FaceGAN emits expression and head pose in a single forward pass per frame and matches or outperforms state-of-art approaches in generation quality.
Sources (1)
- [1]No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise ShapingarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:30 AM
Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.
And it must run online: each frame emitted from audio observed up to the current time, at interactive rates.
Extractive summary: sentences quoted from the sources.