Learning to sail without wind information: reinforcement learning-based control for unmanned sailing vessels under partial observability
Abstract
Unmanned sailing vessels commonly rely on real-time wind-field measurements as key inputs for multi-level decision-making, including path planning and control. During long-endurance offshore missions, however, wind-sensing systems may be affected by harsh environmental conditions, external mounting constraints, and maintenance difficulties. These issues motivate complementary control strategies that can support autonomous navigation while reducing dependence on real-time wind-measurement channels. In this study, we formulate sail–rudder control under partial observability, where the policy uses onboard observable vessel states and target-relative information without receiving measured wind speed or wind direction. A recurrent SAC-LSTM policy is employed as a wind-input-free control implementation, in which a Long Short-Term Memory (LSTM) module is incorporated into the Soft Actor-Critic (SAC) framework to encode recent observation–action history. This design is used to examine whether temporal context from vessel-motion evolution and previous sail–rudder actions can support continuous sail–rudder decision-making when direct wind variables are excluded from the policy input. In randomized-wind simulations evaluated over five independent training seeds, SAC-LSTM achieved a task success rate of 99.36%, compared with 39.64% for the memoryless wind-input-free SAC baseline and 95.24% for the wind-informed SAC reference. Under the same wind-input-free current-observation setting, SAC-LSTM also reduced the average successful-episode navigation time and rudder-angle variation rate relative to the memoryless SAC baseline. In the present randomized-wind benchmark, SAC-LSTM obtained a higher task success rate than the wind-informed SAC reference without using measured wind speed or wind direction, while showing similar successful-episode navigation time and trajectory length. These results indicate that recurrent observation–action history can provide useful temporal context for wind-input-free sail–rudder control under the tested 4-DOF still-water simulation conditions.
// Source
Authors: Chao Yang, Buyu Guo, Yulu Zhang, Yanzhen Gu, Yanjun Li, Sheng Huang, Huang Hui, Jiang Dong, Peiliang Li
Institutions: Zhejiang Ocean University, Sanya University, Institute of Navigation, Hainan Meteorological Service