I. Thoidis, N. Vryzas, L. Vrysis, R. Kotsakis, G. Kalliris, and C. Dimoulas, “Disentangled estimation of reverberation parameters using temporal convolutional networks,” in Proc. AES Convention 152, May 2022, Paper 10593. [Online]. Available: https://aes.org/publications/elibrary-page/?id=21706
Thoidis I, Vryzas N, Vrysis L, Kotsakis R, Kalliris G, Dimoulas C. Disentangled estimation of reverberation parameters using temporal convolutional networks. In: AES Convention 152. Audio Engineering Society; 2022. Paper 10593. Available from: https://aes.org/publications/elibrary-page/?id=21706
@inproceedings{Thoidis2022_21706,
author = {Thoidis, Iordanis and Vryzas, Nikolaos and Vrysis, Lazaros and Kotsakis, Rigas and Kalliris, George and Dimoulas, Charalampos},
title = {{Disentangled estimation of reverberation parameters using temporal convolutional networks}},
booktitle = {AES Convention 152},
note = {Paper 10593},
year = {2022},
month = may,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=21706}
}
TY - CPAPER
TI - Disentangled estimation of reverberation parameters using temporal convolutional networks
AU - Thoidis, Iordanis
AU - Vryzas, Nikolaos
AU - Vrysis, Lazaros
AU - Kotsakis, Rigas
AU - Kalliris, George
AU - Dimoulas, Charalampos
T2 - AES Convention 152
M1 - Paper 10593
PY - 2022
DA - 2022/05/06
UR - https://aes.org/publications/elibrary-page/?id=21706
PB - Audio Engineering Society
LA - en
AB - Reverberation is ubiquitous in everyday listening environments, from meeting rooms to concert halls and record-ing studios. While reverberation is usually described by the reverberation time, getting further insight concerning the characteristics of a room requires to conduct acoustic measurements and calculate each reverberation param-eter manually. In this study, we propose ReverbNet, an end-to-end deep learning-based system to non-intrusively estimate multiple reverberation parameters from a single speech utterance. The proposed approach is evaluated using simulated room reverberation by two popular effect processors. We show that the proposed approach can jointly estimate multiple reverberation parameters from speech signals and can generalise to unseen speakers and diverse simulated environments. The results also indicate that the use of multiple branches disentangles the embedding space from misalignments between input features and subtasks.
ER -