AION
Research paperLarge Language Models · Multimodal Models · Speech & Audio1 source · Oct 7, 2026

Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering

We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.

Key points

  • Linking people's appearance and actions to character identities is essential for understanding video narratives.
  • Starting from LSMDC v2 movie clips, our pipeline matches detected faces to actor reference images, tracks characters across frames, and builds inputs with identity-linked bounding boxes.
  • We study five grounding strategies combining textual coordinates with visual face or estimated person boxes across Video-MLLM families at roughly 2B, 4B, and 8B parameters and larger frontier models.
  • We introduce BAC by LoRA fine-tuning Qwen models at 2B, 4B, and 8B scales on about 32K identity-aware captioned clips.

Sources (1)

Extractive summary: sentences quoted from the sources.