Opens in a new tab

AES E-Library

← Back to search

Convention Paper

Deep Learning-Based Lower-Layer Upmixing

Authors: Páez, Thaddeus; Madrid, Luis; Souza-Blanes, Ema; Bharitkar, Sunil

Convention Paper · Paper 10283 · May 2026

Abstract

This paper introduces a novel approach for generating a lower layer in multichannel audio upmixing, addressing a gap in existing methods that primarily focus on mid and top layers. Leveraging Harmonic-Percussive Separation (HPS), the proposed framework dynamically adjusts key parameters (separation factor, harmonic attenuation, and phase shift) to enhance percussive components while diffusing harmonic elements. We compared three neural network architectures for this task: LSTM, TCN, and Transformer. Experimental results show comparable perceptual quality and objective metrics across all models, with the TCN being the most balanced and suitable for deployment on edge devices.

Details

AES Convention
160
Paper number
10283
Publication date
May 28, 2026
Session subject
AI and Machine Learning in Audio, Audio Processing, Immersive Audio
Affiliation
Samsung Research Tijuana; Samsung Research Tijuana; DMS Audio Lab, Samsung Research America; DMS Audio Lab, Samsung Research America (See document for exact affiliation information.)
Type
Convention Paper