Engineering context-based keywords from sentiment and domain features for enhanced review-level sentiment classification in MOOCs
Abstract
Massive Open Online Courses (MOOCs) generate large volumes of learner feedback that provide valuable insights for educational improvement. However, the domain-specific language used in these reviews often limits the effectiveness of conventional sentiment analysis methods. To address this issue, this study proposes a context-aware feature engineering framework for MOOC sentiment classification. The proposed approach integrates sentiment-based features derived from SentiWordNet with domain-specific features extracted through noun phrase chunking. A co-occurrence-based refinement strategy is subsequently applied to reduce noise and retain meaningful educational concepts. Six feature representations are evaluated using classical machine learning classifiers, including Naïve Bayes, Logistic Regression, Random Forest, and Decision Tree. Experimental results demonstrate that the proposed feature sets consistently improve performance over a TF-IDF baseline. In particular, the Sentiment + DF3 configuration achieves an accuracy of 75.74% using Logistic Regression, representing competitive performance compared to deep learning models such as LSTM (76.23%) and DistilBERT (78.80%). Despite slightly lower accuracy, the proposed framework offers improved interpretability and reduced feature dimensionality. Overall, the results indicate that integrating sentiment lexicons with refined domain-specific features provides an effective and interpretable alternative for MOOC sentiment analysis.
// Source
Authors: Raed Kamil Naser, Keng Hoon Gan, Alaa Thamer Mahmood
Institutions: Universiti Sains Malaysia, Ministry of Defense, Ministry of Defence, Middle Technical University