Visual Orchestration Tax in Agentic VLM Pipelines: Auditing and Certifying Visual Evidence Reuse
Agentic VLM pipelines increasingly pass the same static visual evidence through multiple specialist agents and tools.
ProofPaper ↗
Key points
- This design creates an orchestration-level redundancy mode: semantically unchanged images are repeatedly reconstructed as image-conditioned requests at the VLM API boundary.
- We call this phenomenon visual orchestration tax and develop a measurement-to-certification framework for visual evidence reuse in agentic VLM pipelines.
- The certification side introduces SharedVisCache, a contract-aware evidence reuse hook keyed by image content, preprocessing fingerprint, and encoder assumptions.
- The results position visual reuse as a measurable, behavior-preserving property of agent orchestration and define an agent-layer contract that makes backend prefix or token reuse semantically interpretable.
Sources (1)
- [1]Visual Orchestration Tax in Agentic VLM Pipelines: Auditing and Certifying Visual Evidence ReusearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 6, 11:25 AM
Agentic VLM pipelines increasingly pass the same static visual evidence through multiple specialist agents and tools.
This design creates an orchestration-level redundancy mode: semantically unchanged images are repeatedly reconstructed as image-conditioned requests at the VLM API boundary.
Extractive summary: sentences quoted from the sources.