Y. Wang, X. Wu, and T. Qu, “UP-WGAN: Upscaling Ambisonic Sound Scenes Using Wasserstein Generative Adversarial Networks,” in Proc. AES Convention 152, May 2022, Paper 10577. [Online]. Available: https://aes.org/publications/elibrary-page/?id=21690
Wang Y, Wu X, Qu T. UP-WGAN: Upscaling Ambisonic Sound Scenes Using Wasserstein Generative Adversarial Networks. In: AES Convention 152. Audio Engineering Society; 2022. Paper 10577. Available from: https://aes.org/publications/elibrary-page/?id=21690
@inproceedings{Wang2022_21690,
author = {Wang, Yiwen and Wu, Xihong and Qu, Tianshu},
title = {{UP-WGAN: Upscaling Ambisonic Sound Scenes Using Wasserstein Generative Adversarial Networks}},
booktitle = {AES Convention 152},
note = {Paper 10577},
year = {2022},
month = may,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=21690}
}
TY - CPAPER
TI - UP-WGAN: Upscaling Ambisonic Sound Scenes Using Wasserstein Generative Adversarial Networks
AU - Wang, Yiwen
AU - Wu, Xihong
AU - Qu, Tianshu
T2 - AES Convention 152
M1 - Paper 10577
PY - 2022
DA - 2022/05/06
UR - https://aes.org/publications/elibrary-page/?id=21690
PB - Audio Engineering Society
LA - en
AB - Sound field reconstruction using spherical harmonics (SH) has been widely used. However, order-limited summation leads to an inaccurate reconstruction of sound pressure when the reconstructed region is large. The reconstruction performance also degrades when it comes to high frequency. Upscaling ambisonic sound scenes is used to overcome the limitations. In this work, a deep-learning-based method for upscaling is proposed. Specifically, the generative adversarial network (GAN) is introduced. Instead of estimating the SH coefficients, a U-Net-based fully convolutional generator is introduced, which directly outputs the two-dimensional sound pressure. Results show that the proposed method significantly improves the upscaling results compared with the previous deep-learning-based method.
ER -