Express Paper
Music generation model based on global emotional featureperception
Express Paper · Paper 405 · May 2026
Abstract
The rapid development of artificial intelligence composition technology has brought innovation to music creation. However, current deep learning music generation models often neglect the global correlation of emotional features, resulting in fragmented emotional expression in generated works and insufficient alignment with human emotional perception, making it difficult to meet the core demand for emotional conveyance in diverse music creation. This study aims to propose a music generation method that integrates a global perception mechanism for emotional features. Taking the EMOPIA and VGMIDI preprocessed datasets as the research objects, an improved model based on EMelodyGen is constructed: a GLU network layer is introduced in the feature extraction stage to enhance the models ability to filter and represent emotion-related features; an improved PPO-Clip algorithm is integrated in the training process, and a multi-dimensional emotional reward function is designed to achieve global dynamic perception and optimization of emotional features. Experimental results show that the music21 parsing rate of the EMelodyGen-PPO model on the target dataset is 3% and 4% higher than that of the baseline model, respectively. An automated quality assessment system based on fluency, rhythm stability, harmony richness, melodic smoothness, and structural integrity verifies that the comprehensive score of the models generated works is significantly better than that of the comparative model. This study provides an efficient technical path for emotion-oriented music generation, which can empower grassroots cultural workers and independent musicians at low cost, facilitate diverse music creation practices and emotional audio content dissemination, and align with the diversity and innovative development concept of the AES audio community.
