Safe and adaptive control of non-stationary stochastic systems via Lyapunov-constrained distributional reinforcement learning
Abstract
Abstract Non-stationary environments pose significant challenges for reinforcement learning (RL), particularly in safety-critical applications like robotics and energy systems, where adaptability, stability, and robustness to uncertainty are essential. This paper introduces Lyapunov-Guided Distributional Reinforcement Learning with Entropy Regularization (LG-DRL-ER), a novel framework designed to address these challenges in non-stationary Markov Decision Processes (MDPs). LG-DRL-ER integrates three key components: distributional RL to model the full return distribution, capturing reward uncertainty; Lyapunov stability constraints to ensure safe convergence to a target state; and entropy regularization to promote adaptive exploration. We present a rigorous stability analysis, supported by lemmas and theorems, demonstrating convergence in probability under dynamic conditions. We further establish finite-time convergence guarantees with explicit bounds on the convergence time. An adaptive parameter update rule ensures robustness to changing dynamics, validated through theoretical guarantees and empirical evaluations. Simulations in chaotic and hyperchaotic systems showcase LG-DRL-ER’s superior performance in achieving stable, adaptive control compared to existing methods. This framework advances safe RL by bridging data-driven learning with control-theoretic guarantees, offering significant implications for autonomous systems and dynamic environments.
// Source
Authors: Mohammad Ali Labbaf Khaniki, Morteza Mirzaee, Elahe Moradi
Institutions: Islamic Azad University South Tehran Branch, Islamic Azad University North Tehran Branch, Islamic Azad University, Tehran, K.N.Toosi University of Technology