ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression Segmentation
We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks.
ProofPaper ↗
Key points
- Generalized referring expression segmentation (GRES) requires dynamically balancing high-level semantics for identifying a variable number of language-specified referents with fine-grained visual evidence for precise boundary delineation.
- This requirement challenges existing cascaded vision-language architectures, which typically rely on static feature interfaces and single-pass mask prediction, limiting adaptive perception and geometric correction.
- Evolution-Aware Semantic Scheduling (EASS) couples contour-guided bidirectional boundary sampling with state-conditioned routing of multilevel multimodal features, adapting perception to each contour state.
- Following supervised initialization, Dustbin-Augmented Entropic Credit Transport GRPO (DECT-GRPO) jointly optimizes discrete grounding and continuous contour actions with instance-level credits.
Sources (1)
- [1]ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression SegmentationarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:06 PM
We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks.
Generalized referring expression segmentation (GRES) requires dynamically balancing high-level semantics for identifying a variable number of language-specified referents with fine-grained visual evidence for precise boundary delineation.
Extractive summary: sentences quoted from the sources.
Before this
- Oct 8, 2026Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
- Oct 8, 2026Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
- Oct 8, 2026SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
- Oct 8, 2026Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
- Oct 7, 2026On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
- Oct 7, 2026Q-Learning with Scalar Adjoint Matching