FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?
Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios.
Key points
- Under bounded context budgets, models must balance temporal history, spatial resolution, and temporal granularity; sparse sampling at 1--2 FPS misses fast events.
- We introduce FastBench to evaluate high-dynamic perception in real-world video streams.
- We also present ProactiveFrame, a training-free baseline that adjusts incoming frame rates through text tokens.
- FastBench provides a testbed for high-dynamic streaming video understanding.
Sources (1)
- [1]FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 05:55 PM
Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios.
Under bounded context budgets, models must balance temporal history, spatial resolution, and temporal granularity; sparse sampling at 1--2 FPS misses fast events.
Extractive summary: sentences quoted from the sources.