Opens in a new tab

AES E-Library

← Back to search

Convention Paper

Extension of Monaural to Stereophonic Sound Based on Deep Neural Networks

Authors: Chun, Chan Jun; Jeong, Seok Hee; Park, Su Yeon; Kim, Hong Kook

AES Convention 139 · Paper 9400 · October 2015

Abstract

In this paper we propose a method of extending monaural into stereophonic sound based on deep neural networks (DNNs). First, it is assumed that monaural signals are the mid signals for the extended stereo signals. In addition, the residual signals are obtained by performing the linear prediction (LP) analysis. The LP coefficients of monaural signals are converted into the line spectral frequency (LSF) coefficients. After that, the LSF coefficients are taken as the DNN features, and the features of the side signals are estimated from those of the mid signals. The performance of the proposed method is evaluated using a log spectral distortion (LSD) measure and a multiple stimuli with a hidden reference and anchor (MUSHRA) test. It is shown from the performance comparison that the proposed method provides lower LSD and higher MUSHRA score than a conventional method using hidden Markov model (HMM).

Details

Published in
AES Convention 139
AES Convention
139
Paper number
9400
Publication date
October 6, 2015
Session subject
Signal Processing
Affiliation
Gwangju Institute of Science and Technology (GIST), Gwangju, Korea; City University of New York, New York, NY, USA (See document for exact affiliation information.)
Type
Convention Paper