Non-invasive detection of obstructive sleep apnea from speech in a limited clinical dataset
Abstract
Abstract This study investigates non-invasive detection of obstructive sleep apnea (OSA) from daytime speech in a limited polysomnography-labeled clinical cohort. The objective was to evaluate whether speaker-verification-inspired representations can support robust speech-based OSA screening under constrained clinical data conditions. We compared conventional mel-frequency cepstral coefficient (MFCC)-based classifiers with i-vector systems using cosine-similarity, support vector machine (SVM), and neural-network (NN) back ends, and included an exploratory ECAPA-TDNN embedding experiment using a pretrained frozen speaker model with prototype-based cosine scoring. Speech recordings from 93 participants were collected under controlled conditions and evaluated using subject-independent validation. Supervised i-vector back ends achieved the strongest performance. The i-vector SVM system obtained the highest numerical accuracy of 96.3% and the strongest threshold-independent performance, with AUC-ROC of 0.987 and AUC-PR of 0.992, while the i-vector NN system showed closely comparable performance with a non-significant AUC-ROC difference from i-vector SVM. ECAPA-TDNN showed moderate class separability but did not outperform the i-vector systems. Demographic-only and speech-demographic fusion analyses indicated that sex, age, and body mass index carried predictive information but did not fully explain the best speech-based performance. These findings support speech-based OSA screening as a low-burden complementary approach, while emphasizing the need for larger, balanced, and externally validated cohorts.
// Source
Authors: Martina Škapová, Pavol Partila, Jaromir Tovarek, Tereza Lubojacka, Jana Slonkova, Ondrej Hruby, Samuel Genzor, Jan Mizera, Petr Mooz
Institutions: VSB - Technical University of Ostrava, University Hospital Ostrava, University Hospital Olomouc