Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering
We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.
Key points
- Linking people's appearance and actions to character identities is essential for understanding video narratives.
- Starting from LSMDC v2 movie clips, our pipeline matches detected faces to actor reference images, tracks characters across frames, and builds inputs with identity-linked bounding boxes.
- We study five grounding strategies combining textual coordinates with visual face or estimated person boxes across Video-MLLM families at roughly 2B, 4B, and 8B parameters and larger frontier models.
- We introduce BAC by LoRA fine-tuning Qwen models at 2B, 4B, and 8B scales on about 32K identity-aware captioned clips.
Sources (1)
- [1]Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question AnsweringarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:41 PM
We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation.
Linking people's appearance and actions to character identities is essential for understanding video narratives.
Extractive summary: sentences quoted from the sources.