C. Uhle, M. Kratschmer, A. Travaglini, and B. Neugebauer, “Clean dialogue loudness measurements based on Deep Neural Networks,” in Proc. AES Conference: 2020 AES International Conference on Audio for Virtual and Augmented Reality (August 2020), Aug. 2020, Paper 10479. [Online]. Available: https://aes.org/publications/elibrary-page/?id=21156
Uhle C, Kratschmer M, Travaglini A, Neugebauer B. Clean dialogue loudness measurements based on Deep Neural Networks. In: AES Conference: 2020 AES International Conference on Audio for Virtual and Augmented Reality (August 2020). Audio Engineering Society; 2020. Paper 10479. Available from: https://aes.org/publications/elibrary-page/?id=21156
@inproceedings{Uhle2020_21156,
author = {Uhle, Christian and Kratschmer, Michael and Travaglini, Alessandro and Neugebauer, Bernhard},
title = {{Clean dialogue loudness measurements based on Deep Neural Networks}},
booktitle = {AES Conference: 2020 AES International Conference on Audio for Virtual and Augmented Reality (August 2020)},
note = {Paper 10479},
year = {2020},
month = aug,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=21156}
}
TY - CPAPER
TI - Clean dialogue loudness measurements based on Deep Neural Networks
AU - Uhle, Christian
AU - Kratschmer, Michael
AU - Travaglini, Alessandro
AU - Neugebauer, Bernhard
T2 - AES Conference: 2020 AES International Conference on Audio for Virtual and Augmented Reality (August 2020)
M1 - Paper 10479
PY - 2020
DA - 2020/08/06
UR - https://aes.org/publications/elibrary-page/?id=21156
PB - Audio Engineering Society
LA - en
AB - Loudness normalization based on clean dialogue loudness improves consistency of the dialogue level compared to the loudness of the full program measured at speech or signal activity. Existing loudness metering methods can not estimate clean dialogue loudness from mixture signals comprising speech and background sounds, e.g. music, sound effects or environmental sounds. This paper proposes to train deep neural networks with input signals and target values obtained from isolated speech and backgrounds to estimate the clean dialogue loudness. Furthermore, the proposed method outputs estimates for loudness levels of background and mixture signal, and Voice Activity Detection. The presented evaluation reports a mean absolute error of 1.5 LU for momentary loudness, 0.5 LU for short-term and 0.27 LU for long-term loudness of the clean dialogue given the mixture signal.
ER -