AION
Research paperImage, Video & 3D Generation · Speech & Audio · Multimodal Models1 source · Oct 8, 2026

No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping

Audio-driven facial animation underpins real-time avatars, telepresence, and embodied virtual agents.

Key points

  • And it must run online: each frame emitted from audio observed up to the current time, at interactive rates.
  • Audio-conditioned facial motion occupies a comparatively low-dimensional manifold, a regime where a single-pass GAN suffices.
  • We show that a causal, time-invariant generator driven by i.i.d. noise cannot suppress its output spectrum over a band without collapsing its per-step innovation.
  • FaceGAN emits expression and head pose in a single forward pass per frame and matches or outperforms state-of-art approaches in generation quality.

Sources (1)

Extractive summary: sentences quoted from the sources.