Opens in a new tab

AES E-Library

← Back to search

Express Paper Open Access

Deep Learning Based Voice Extraction and Primary-Ambience Decomposition for Stereo to Surround Upmixing

Authors: Paez Amaro, Ricardo Thaddeus; Tejeda Ocampo, Carlos; Souza Blanes, Ema; Bharitkar, Sunil; Madrid Herrera, Luis

AES Convention 154 · Paper 62 · May 2023

Abstract

Surround systems have gained popularity in home entertainment despite the fact that most of the cinematic content is delivered in two-channel stereo format. Although there are several upmixing options, it has proven challenging to deliver an upmixed signal that approximates the original directionality and timbre intended by the mixing artist. The aim of this work is to design a two-to-five channels upmixer using a novel upmixing strategy combining voice extraction and primary-ambience decomposition. Results from a modified-MUSHRA test show that our proposed upmixer outperforms established alternatives for cinematic upmixing in perceived spatial and timbral quality.

Details

Published in
AES Convention 154
AES Convention
154
Paper number
62
Publication date
May 6, 2023
Session subject
Neural Networks
Affiliation
Samsung Research Tijuana, Mexico; Samsung Research Tijuana, Mexico; Samsung Research America, Mountain View, CA, USA; Samsung Research America, Mountain View, CA, USA; Samsung Research Tijuana, Mexico (See document for exact affiliation information.)
Type
Express Paper