Tell Robot What Not to Do: A Negation Understanding Perspective
To this end, we propose NegaAlign, a parameter-efficient, plug-and-play framework that extends pretrained VLAs to follow negated instructions through image-language supervision alone.
Key points
- Instruction following enables robots to perform diverse tasks specified in natural language, making it a fundamental capability for human-robot interaction.
- We investigate how to enable vision-language-action models (VLAs) to follow negated instructions, where robots must accomplish task goals while respecting explicit exclusions.
- Specifically, we introduce Negation Transformation Layers into selected layers of the vision-language backbone to reshape intermediate instruction representations.
- We further introduce NegaBench, a simulation benchmark spanning 10 scenarios across five domains for systematically evaluating manipulation under negated constraints.
Sources (1)
- [1]Tell Robot What Not to Do: A Negation Understanding PerspectivearXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 01:39 PM
To this end, we propose NegaAlign, a parameter-efficient, plug-and-play framework that extends pretrained VLAs to follow negated instructions through image-language supervision alone.
Instruction following enables robots to perform diverse tasks specified in natural language, making it a fundamental capability for human-robot interaction.
Extractive summary: sentences quoted from the sources.