SuperNav: An Agentic Navigation System for Any Task in Any Scene
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.
Key points
- Some existing methods fine-tune multimodal large language models (MLLMs) to predict navigation actions, making their behavior dependent on the coverage of navigation training data and potentially limiting generalization to new requests and environments.
- To realize this idea, we introduce SuperNav, which equips a pretrained MLLM with a specialized agent harness without navigation-specific fine-tuning of the MLLM.
- A unified visual-point interface connects decision-making to motion by allowing the model to specify destinations directly in images and revise its decisions from execution feedback.
- Category-level evaluation on HM3D and deployment on a real quadruped robot further demonstrate its applicability across environments.
Sources (2)
- [1]SuperNav: An Agentic Navigation System for Any Task in Any SceneHugging Face Daily Papers · Oct 8, 12:00 AM
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality.
Some existing methods fine-tune multimodal large language models (MLLMs) to predict navigation actions, making their behavior dependent on the coverage of navigation training data and potentially limiting generalization to new requests and environments.
- [2]SuperNav: An Agentic Navigation System for Any Task in Any ScenearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 03:19 PM · same content
Extractive summary: sentences quoted from the sources.