ResearchResearch paperLarge Language Models · Speech & Audio · Multimodal Models1 source · Oct 7, 2026

TiTok: Audio-Visual LLM for Multi-Segment Temporal Grounding

Audio-visual multi-segment grounding (AV-MSG) in untrimmed videos, reasoning over audio-visual evidence and predicting multiple segments for a query, is a fundamental problem but remains challenging.

Key points

  • Visual-only models overlook complementary acoustic cues, while audio-visual models often fail to calibrate the number of events - a phenomenon we refer to as count miscalibration.
  • We present TiTok, an audio-visual large language model (AV-LLM) that localizes an arbitrary number of temporal event segments for each query.
  • For precise boundary prediction, we introduce the Time Token Interleaving (TTI) method, which explicitly injects special time tokens into the audio-visual stream to align input-side temporal perception with output-side temporal prediction.
  • We further propose decoupled, multi-segment-oriented rewards for reinforcement learning, consisting of global, local, count, precision, and format rewards, optimized with Group reward-Decoupled Normalization Policy Optimization (GDPO).

Sources (1)

  • [1]TiTok: Audio-Visual LLM for Multi-Segment Temporal Grounding
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 04:08 AM
    Audio-visual multi-segment grounding (AV-MSG) in untrimmed videos, reasoning over audio-visual evidence and predicting multiple segments for a query, is a fundamental problem but remains challenging.
    Visual-only models overlook complementary acoustic cues, while audio-visual models often fail to calibrate the number of events - a phenomenon we refer to as count miscalibration.

Extractive summary: sentences quoted from the sources.

Related