Challenges and opportunities of generative artificial intelligence models in audio/acoustic domain: a comprehensive survey
Abstract
Abstract Generative artificial intelligence (AI) has transformed image and text processing, but its adoption in the audio/acoustic domain remains underexplored due to inherent challenges in modeling long-term temporal dependencies in one-dimensional signals and achieving human-perceptible coherence. This comprehensive survey addresses these gaps by systematically reviewing state-of-the-art generative AI models, including generative adversarial networks (GANs), diffusion/flow-matching models, variational autoencoders (VAEs), recurrent neural networks (RNNs), transformers, and Neural Codec Language Models (Codec LMs), organized around three primary application domains: (1) speech synthesis , encompassing text-to-speech conversion, neural vocoding, voice conversion, and zero-shot voice cloning; (2) music generation , covering both symbolic and acoustic composition, multi-track generation, and style transfer; and (3) general audio synthesis, sound effects, and source separation , including text-to-audio generation, audio restoration and enhancement, data augmentation, and conditional source separation. We provide a detailed taxonomy of architectures, functionalities, comparative strengths, and limitations, supported by common evaluation metrics. Structured comparisons with existing surveys demonstrate that this is the first work to jointly cover all three audio domains and all generative model families, while also providing dedicated evaluation-metric analysis and cross-domain comparative assessments. We further highlight emerging opportunities in AI applications such as healthcare monitoring (e.g., symptom analysis and mental health assessment) and biometric authentication, demonstrating the potential of synthetic audio to address real-life challenges. Through research gap identification, model efficacy comparison, and future direction outlining, this survey serves as a foundational reference for advancing generative AI techniques across the audio domain.
// Source
Authors: Sayanton Dibbo, Sudip Vhaduri, Chia-Hua Lin
Institutions: University of Tennessee at Knoxville, Knoxville College, Purdue University West Lafayette, University of Alabama, Institute of Electrical and Electronics Engineers