Opens in a new tab

AES E-Library

← Back to search

Convention Paper Open Access

A Scalable Two-Stage Automatic Mixing System Integrating Machine Learning and Domain Knowledge

Authors: Shi, Jinjie; Xie, Kunzhu; Ma, Yinghao; Reiss, Joshua

Convention Paper · Paper 10232 · October 2025

Abstract

Music mixing involves transforming clean, individual tracks into a cohesive final mix using audio effects and expert knowledge. While rule-based and machine learning methods have shown promise, scaling them to real-world situations remains challenging. We propose a two-stage mixing architecture that combines domain knowledge with deep learning, enabling the system to handle over 100 input tracks with high perceptual quality.
The first stage uses a rule-based level balancing system to mix grouped tracks into stems. The second stage employs a differentiable mixing style transfer model guided by a reference mix. To enhance intra-group (within subgroup) robustness, we refine loudness estimation by incorporating spectral centroid and fundamental frequency features, addressing limitations of Loudness Units relative to Full Scale (LUFS) on narrowband signals.
Subjective listening tests demonstrate that our enhanced intra-group mixing approach consistently outperforms LUFS-based baselines across multiple musical genres. Furthermore, our proposed two-step system enables deep learning to successfully handle projects with over 100 tracks for the first time, achieving mixing results that significantly surpass those of traditional rule-based systems. Code and audio examples are available at https://doi.org/10.5281/zenodo.17171082.

Details

AES Convention
159
Paper number
10232
Publication date
October 14, 2025
Affiliation
Queen Mary University of London; Queen Mary University of London; Queen Mary University of London; Wuhan University of Communication (See document for exact affiliation information.)
Type
Convention Paper