Why VLMs Miss Small Objects, and When Zooming In Is Safe
Vision-language models (VLMs) often miss small objects in large images.
ProofPaper ↗
Key points
- We ask three questions: what limits them, which of these limits better models can remove, and whether the classical way of handling large images, local decomposition, still has a future.
- We answer them with a theory built on two quantities of the image interface: S, the number of visual tokens across an object's side, and L, the content a call must cover.
- The limits: a W x H image sent whole within N tokens gives an object of side m at most msqrt(N/(WH)) tokens per side.
- Assuming only that more tokens per object and less content per call do not hurt on average, splitting an image cannot lower recall if no view zooms out relative to the whole image and views overlap by one object.
Sources (1)
- [1]Why VLMs Miss Small Objects, and When Zooming In Is SafearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 7, 02:10 AM
Vision-language models (VLMs) often miss small objects in large images.
We ask three questions: what limits them, which of these limits better models can remove, and whether the classical way of handling large images, local decomposition, still has a future.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 6, 2026[AINews] Reflection Beam - 501B-A23B American Open Model
- Oct 5, 2026Our approach to EU text provenance rules
- Oct 5, 2026Sharing AI progress in mathematics
- Oct 5, 2026Building advertising for the way people use AI
- Sep 30, 2026Gemini 4 Argon: our next era of frontier intelligence
- Sep 30, 2026[AINews] OpenAI DevDay 2026: Dots, 6.1 Sol, Ultrafast, Decisions API, Agents API, Spaces, Marketplace, and 1.2 Billion ChatGPT WAU