Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness
We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output.
ProofPaper ↗
Key points
- Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived.
- Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it.
- Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe.
- We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead.
Sources (1)
- [1]Have I Seen Enough? Frozen Video-Language Models Encode Evidence ReadinessarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 03:43 PM
We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output.
Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived.
Extractive summary: sentences quoted from the sources.