Optimal Policy Analysis for Monotonic (PO)MDPs and an Application to Fishery Management
Abstract
<title>Abstract</title> We study the monotonicity properties of optimal policies for a class of fully/partially observable Markov Decision Processes (MDPs) motivated by renewable natural resource management. We first introduce a new class of deterministic monotonic MDPs with continuous state and action spaces, and we prove that their optimal policies are monotonic in the state. These monotonic MDPs are particularly interesting, because they do not satisfy certain supermodularity assumptions on the reward function and the transition dynamics in previous monotonicity analyses for the optimal policies of MDPs, and our proof requires an original inductive analysis. Our theoretical results are numerically illustrated on a fully observable fishery management problem, using a modified value iteration algorithm that computes near-optimal monotonic policies with a max-min smoothing strategy. We then introduce a generalization of our monotonic MDPs to handle partial observability and stochas-ticity, and we conjecture the existence of a near-optimal policy that is monotonic in the expected state. Our conjecture is supported by numerical results on a partially observable fishery management problem, using algorithms for optimizing over a class of monotonic policies called multi-threshold policies. Moreover, when the environment is unknown and needs to be learned through interactions, multi-threshold policies appear to be more robust to model learning error than general policies in a model-based offline reinforcement learning approach.
// Source
Authors: 俊 鳥居, Jerzy A. Filar, Dirk P. Kroese, Nan Ye
Institutions: The University of Queensland