Identification of features associated with dysmenorrhea in women aged 20–50 years using logistic regression and synthetic minority oversampling technique
Abstract
Dysmenorrhea (Dys) is a prevalent gynecological condition that significantly impairs women’s quality of life. This study investigates the key predictive features associated with Dys in Taiwanese women aged 20–50, using logistic regression (LR) with and without the Synthetic Minority Oversampling Technique (SMOTE) to address class imbalance. Data were obtained from the MJ Health Database, involving 1,038 women. Dys was defined via self-report. A total of 41 clinical, biochemical, and lifestyle features were analyzed. Performance metrics—including accuracy, sensitivity, specificity, precision, F1 score, and area under the ROC and PR curves—were computed via 5-fold cross-validation. Logistic regression coefficients and odds ratios were calculated to evaluate feature associations. Without SMOTE, the LR model showed high specificity (97.9%) and accuracy (93.2%) but low sensitivity (51.9%). With SMOTE, sensitivity improved significantly (83.2%), although precision and specificity declined, demonstrating a trade-off between detecting positive cases and generating false positives. Both models maintained high AUC scores (ROC AUC ~ 0.92; PR AUC ~ 0.68). Variable importance analysis, based on standardized coefficients, highlighted six key predictive features: childbirth history, age, glutamic pyruvic transaminase (GPT), glutamic oxaloacetic transaminase (GOT), menstrual flow volume, and diastolic blood pressure (DBP). SMOTE-enhanced logistic regression offers a highly sensitive model for detecting Dys in imbalanced datasets. The identified predictive features could inform targeted screening interventions in primary care settings, though external validation is required.
// Source
Authors: Jian-Lun Su, Yu-Huan Weng, Chi-Kang Lin, Ta-Wei Chu, Yu-Chi Wang, Jyh-Gang Leu
Institutions: Fu Jen Catholic University, National Defense Medical Center, Memorial Hospital