AION
Research paperRobotics & Embodied AI · Multimodal Models · Computer Vision1 source · Oct 8, 2026

The Operator Mismatch Problem: Deploying BEV Perception with Portable GPU Compute

We present BEVPIPE, a framework for deploying multimodal BEV perception pipelines using portable GPU compute APIs and integrating them with production inference runtimes.

Key points

  • Modern autonomous driving systems rely on bird's-eye-view (BEV) perception models that fuse camera and LiDAR inputs to detect objects in 3D space.
  • The reason is an operator mismatch between dense convolutions (which runtimes handle well), sparse 3D convolutions (which runtimes cannot represent), and geometric scatter operations (which runtimes have no vocabulary for).
  • Today, every sparse convolution library is CUDA-only and PyTorch-coupled, locking BEV deployment to a single vendor's hardware and a single execution framework.
  • BEVPIPE partitions the model into runtime-managed dense subgraphs and three external operator extensions (voxelizer, sparse encoder, BEV projector), connected through a shared GPU memory space.

Sources (1)

  • [1]The Operator Mismatch Problem: Deploying BEV Perception with Portable GPU Compute
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 08:42 AM
    We present BEVPIPE, a framework for deploying multimodal BEV perception pipelines using portable GPU compute APIs and integrating them with production inference runtimes.
    Modern autonomous driving systems rely on bird's-eye-view (BEV) perception models that fuse camera and LiDAR inputs to detect objects in 3D space.

Extractive summary: sentences quoted from the sources.