SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning
Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation.
Key points
- We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory.
- Building on this monitor, we propose SafeInferCom, a formal verifier-guided framework that preserves valid intermediate plans and directs error correction during generation.
- Experiments across multiple LRLMs and planning domains reveal reasoning-response inconsistency and limited self-correction under one-shot inference.
- SafeInferCom improves planning success and accelerates error correction relative to one-shot inference.
Sources (1)
- [1]SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task PlanningarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 04:19 AM
Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation.
We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory.
Extractive summary: sentences quoted from the sources.