ResearchResearch paperSpeech & Audio · Image, Video & 3D Generation · Multimodal Models1 source · Oct 6, 2026

Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses

In this paper, we propose LipDA, a novel framework for joint LipSync Detection and Attribution, which takes advantage of the inconsistency between head and lip.

Key points

  • Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks.
  • However, existing defense strategies struggle against LipSync forgeries, as advanced LipSync generation methods not only achieve better lip synchronization but also eliminate visual artifacts.
  • An important reason is that they overlook an inherent biological coupling between lip movements and head poses in natural speech videos.
  • We conduct extensive experiments on two challenging LipSync datasets as well as our own proposed large-scale and multi-generator dataset.

Sources (1)

Extractive summary: sentences quoted from the sources.

Related