Exploring Speech Emotion Recognition: Techniques and Emerging Trends
Abstract
Emotion is a fundamental aspect of human communication, and the automatic recognition of emotions from speech has gained significant importance in the development of safety-critical and service-oriented applications. This survey paper presents a comprehensive review of recent developments in Speech Emotion Recognition (SER), focusing primarily on research contributions from 2021 to 2025. The paper discusses the general SER framework, major emotional speech datasets, commonly used acoustic and spectral features, traditional machine learning techniques, and modern deep learning approaches. Furthermore, it analyzes recent advancements in transformer-based and cross-corpus SER methods, robustness enhancement techniques, and multi-modal emotion recognition systems. The survey also highlights current research trends and open challenges. By synthesizing existing literature and identifying research gaps, this article aims to provide researchers with a structured understanding of the current state of the field. Keywords: Speech Emotion Recognition, Deep Learning, Transformer Models, Cross-Corpus Generalization,Self-Supervised Learning, Acoustic Features, Affective Computing
// Source
Authors: Suman Giri, ANITHA N
Institutions: PES University