Engineering & Technologyarticle2026-08-11

From Reactive to Proactive: World‑Model Enhanced VLA for Mobile Manipulation

Open access0 citations

Abstract

Mobile robot manipulation demands tight coordination of locomotion and dexterous arm control under changing viewpoints and contacts, making the action space far larger than in fixed-base settings. Vision-Language-Action (VLA) models map perceptual and linguistic inputs directly to whole-body actions, but they suffer from two key limitations: open‑loop execution accumulates control errors, and their reactive nature lacks an internal model of physical dynamics. World models offer a predictive simulator that enables robots to anticipate how the environment evolves before acting. However, existing world action models (WAMs) still struggle with coarse temporal granularity, coupled navigation‑manipulation modeling, and train‑test mismatches, often missing fine‑grained contact information and drifting over long horizons. Recent work bridges VLA and world models: DreamTrajectory predicts intent‑level trajectories and scores them with a lightweight world model for online action selection; CheckVLA uses an action‑conditioned world model to verify execution without breaking real‑time constraints; ABot‑M0.5 aligns granularity, action space, and consistency for unified mobile‑manipulation modeling. Together, these advances show that integrating VLA with explicit world prediction is a promising path toward more robust, autonomous, and generalizable mobile manipulation systems.

// Source

View paper (DOI)Open access versionOpenAlexCalifornia Digital LibraryPublished 2026-08-11

Authors: Youpeng Wen

Institutions: Chinese University of Hong Kong