AI & Computingreview2026-08-21

Adaptive Decision-Making in Audio Classification: A Systematic Review of Reinforcement Learning and Multi-Armed Bandit Frameworks

Open access0 citations

Abstract

Adaptive decision-making has emerged as an important direction in audio classification. Reinforcement learning (RL) and multi-armed bandit (MAB) methods offer principled frameworks for sequential and localized decision-making. However, despite growing interest in these approaches and the availability of review studies on audio classification and audio-based learning, their role within the audio classification pipeline remains fragmented and has not been systematically synthesized. This study presents a systematic review of RL- and MAB-based approaches in audio classification, to characterize how adaptive decision-making is formulated, where it is applied within the pipeline, and what impact it has on system performance, robustness, and efficiency. The review followed PRISMA guidelines. A literature search across IEEE Xplore, Scopus, Nature, and Google Scholar identified 1896 records. Thirty-one studies met the predefined inclusion criteria and data were extracted for analysis. Reporting quality was assessed using the TRIPOD framework, and methodological reliability was evaluated using a domain-specific risk-of-bias tool. The findings show that adaptive decision-making in audio classification is evolving toward a control-centric paradigm. RL-based optimization and control approaches dominated the literature, whereas bandit-based approaches remained comparatively underused. A central finding of the review is a structural mismatch between problem type and method choice. Many adaptive tasks are inherently local and repeated decision problems, yet they are predominantly addressed using full RL frameworks rather than lighter bandit formulations. This review introduces a taxonomy of adaptive decision-making mechanisms and proposes a unified framework that reframes audio classification as a layered decision-making process under uncertainty. The evidence suggests that the future of audio classification lies not only in improved prediction, but in adaptive decision-making architectures capable of managing uncertainty, variability, and real-world deployment constraints.

// Source

View paper (DOI)Open access versionOpenAlexMachine Learning and Knowledge ExtractionPublished 2026-08-21

Authors: Geofrey Owino, Timothy Kamanu, John Ndiritu

Institutions: University of Nairobi