A contrastive learning framework with HuBERT for violin music emotion recognition
Abstract
Music emotion recognition (MER) is essential for affective computing and intelligent recommendation, yet solo instrumental music—particularly violin performance—remains underexplored due to the absence of lyrics, the subtlety of emotional expression, and the lack of dedicated datasets. To address these gaps, this study adopts the VioMusic dataset, a publicly available violin performance dataset with phrase-level Valence-Arousal (VA) annotations, comprising approximately 7.5 h of recordings and 1926 segmented phrases. Furthermore, this study proposes a CL-HuBERT (Contrastive Learning with HuBERT) framework that performs domain-adaptive continued pre-training of a HuBERT encoder via contrastive learning on unlabeled violin audio, followed by supervised fine-tuning for valence-arousal prediction. Experimental results demonstrate that the proposed framework achieves superior performance, with Pearson correlation coefficients of 0.718 for arousal and 0.701 for valence, and MAE values of 0.089 and 0.095, respectively, under the optimal feature fusion setting. Systematic comparisons of audio features, feature fusion strategies, data preprocessing methods, and deep feature extractors further validate the effectiveness of the proposed approach. This study provides an effective self-supervised contrastive learning framework and establishes a strong benchmark for solo instrumental MER.
// Source
Authors: Xinyue Liu
Institutions: Anyang Institute of Technology, Anyang Normal University