Journal Article
A Generative Adversarial Network-Based Approach for Music Genre Style Transfer With Spectral Entropy Segment Selection
Journal of the Audio Engineering Society · Volume 74 · Issue 9 · pp. 659–680 · September 2026
Abstract
Music style transfer (MST) finds applications in music creation, music therapy, and music recommendation. Human auditory perceptual ability enables the recognition of musical genres by perceiving typical instrument types and rhythmic patterns. To simulate this perceptual process, a generative adversarial network (GAN)-based method for genre MST with spectral entropy segment selection (SESSGAN-MST) is proposed. First, spectral entropy is employed to model the importance of music segments that carry genre features. A dual-threshold spectral entropy segment selection mechanism is developed that integrates global and local spectral entropy information to identify two types of critical segments: moderate-contribution feature segments and critical-contribution feature segments of the music. Subsequently, a dynamic adaptive weight optimization approach is adopted, and in conjunction with GANs, a more accurate music style transformation function is constructed for distinct music segments. Experimental results on the GTZAN dataset of pure music demonstrate that the proposed method achieves the highest preference rate in the majority of style preference test scenarios and outperforms three comparison methods GAN -voice conversion using mel_spectrograms, variational auto cycle-consistent GAN, and feature-specific loss self-attentive GAN–voice conversion) on Fréchet Audio Distance metrics in the majority of conversion scenarios, reducing the mean Fréchet Audio Distance from 20.26 ± 4.84 to 14.68 ± 2.80 (a 27.54% reduction).
