Explicit task reasoning empowering robotic manipulation
Abstract
Currently, robots still struggle to perform human-like explicit task reasoning. Most existing vision-language-action (VLA) approaches heavily rely on large-scale demonstration datasets and carefully designed models, which pose significant challenges for model interpretability, generalization, and efficiency in real-world robotic applications. We present a novel VLA paradigm that performs explicit task reasoning directly on the robot. We decouple vision, language, reasoning, and action, interconnecting these four modules through proposed numerical signals. Leveraging the large language model (LLM)’s common sense, mathematical, and physical reasoning abilities, the system generates a complete action plan purely through explicit logic. Benefiting from this simple yet effective paradigm, our method requires neither task-specific demonstration data nor manipulation policy training, relying solely on explicit task reasoning to achieve a wide range of robotic manipulations, with all robot behaviors being interpretable. We validate the approach with six designed experimental suites assessing 3D scene understanding, qualitative and quantitative manipulation, language interaction and complex reasoning, and long-horizon multi-step autonomy. Across 240 real-world manipulation trials, the robot achieved a 91.67% success rate. Our proposed paradigm offers new insights for VLA research, circumventing the inefficiencies of large-scale data collection and model training, and instead leveraging the efficiency of explicit task reasoning.
// Source
Institutions: Hong Kong Polytechnic University, Southeast University