AION
Research paperMultimodal Models · Robotics & Embodied AI1 source · Oct 8, 2026

DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception

To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks.

Key points

  • Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving.
  • However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements.
  • Specifically, we first transform diverse perceptual representation (i.e., 2D bounding boxes, 2D masks, and 3D bounding boxes) into 1D vector sequences, which are then mapped to compact discrete tokens in the high-dimensional space.
  • Then, a lightweight de-tokenizer enables seamless integration with MLLMs by decoding output tokens back to original 2D and 3D perceptual representations.

Sources (1)

  • [1]DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:32 PM
    To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks.
    Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving.

Extractive summary: sentences quoted from the sources.