AION
Research paperRobotics & Embodied AI · Agents & Tool Use · Multimodal Models2 sources · Oct 8, 2026

SuperNav: An Agentic Navigation System for Any Task in Any Scene

General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.

Key points

  • Some existing methods fine-tune multimodal large language models (MLLMs) to predict navigation actions, making their behavior dependent on the coverage of navigation training data and potentially limiting generalization to new requests and environments.
  • To realize this idea, we introduce SuperNav, which equips a pretrained MLLM with a specialized agent harness without navigation-specific fine-tuning of the MLLM.
  • A unified visual-point interface connects decision-making to motion by allowing the model to specify destinations directly in images and revise its decisions from execution feedback.
  • Category-level evaluation on HM3D and deployment on a real quadruped robot further demonstrate its applicability across environments.

Sources (2)

  • [1]SuperNav: An Agentic Navigation System for Any Task in Any Scene
    Hugging Face Daily Papers · Oct 8, 12:00 AM
    General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.
    Some existing methods fine-tune multimodal large language models (MLLMs) to predict navigation actions, making their behavior dependent on the coverage of navigation training data and potentially limiting generalization to new requests and environments.
  • [2]SuperNav: An Agentic Navigation System for Any Task in Any Scene
    arXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:19 PM · same content

Extractive summary: sentences quoted from the sources.