Barrier-Function-Constrained Residual RL-Based Current Control for LCL-Filtered NPC Inverters under Weak-Grid Operation
Abstract
As renewable energy penetration continues to grow, grid-connected inverters are increasingly required to operate reliably under weak-grid conditions, where low short-circuit ratios, changing grid impedances, and rapid disturbances reveal the limitations of conventional fixed-gain PI current controllers. Reinforcement learning (RL) offers a more adaptive control approach; however, its application to fast inner current loops remains limited by stringent safety requirements and practical rediness challenges. This paper presents a control framework that combines a model-free Proximal Policy Optimization (PPO)-based residual reinforcement learning controller with a model-based Control Lyapunov Function–Control Barrier Function (CLF–CBF) quadratic-program (QP) safety filter for an LCL-filtered three-phase neutral-point-clamped (NPC) grid-connected inverter. Although residual RL and CLF–CBF safety methods have been explored independently, this work integrates a model-free PPO-based residual controller with a model-based CLF–CBF safety filter within a unified architecture for fast inner-loop current control under varying grid-strength and operating conditions. Instead of replacing the conventional PI controller, the residual policy generates bounded corrective voltage commands that complement a fixed PI baseline, while the CLF–CBF–QP safety filter provides analytical enforcement of stability and operational constraints independently of the learned policy. A two-stage training strategy, consisting of unconstrained warm-start followed by safety-constrained fine-tuning, enables efficient policy learning while maintaining safe operation. Evaluated across eight representative operating scenarios with five independent trials per condition, the proposed framework achieves a maximum total harmonic distortion (THD) reduction of 49.7% under extreme thermal stress and 44.2% under nominal operation, with an average THD reduction of 25.7% across all scenarios. The framework also reduces aggregate constraint violations by 35.7% across the evaluated conditions, indicating that the integration of model-free residual reinforcement learning with a model-based CLF–CBF safety layer can improve harmonic performance while providing model-based supervision of current and voltage operating constraints.
// Source
Authors: Jabala Nur Fahima, Kayes Hasan, Md. Rifat Hazari, Shameem Ahmad, Chowdhury Akram Hossain, Mohammad Abdul Mannan, Emanuele Ogliari
Institutions: Politecnico di Milano, BRAC University, RMIT University, American International University-Bangladesh