I. Nawfal et al., “Ambisonics Super-Resolution Using A Waveform-Domain Neural Network,” in Proc. AES 2024 International Conference on Audio for Virtual and Augmented Reality, Aug. 2024, Paper 4. [Online]. Available: https://aes.org/publications/elibrary-page/?id=22653
Nawfal I, Delikaris Manias S, Souden M, Merimaa J, Atkins J, McMullin E, Pirhosseinloo S, Phillips D. Ambisonics Super-Resolution Using A Waveform-Domain Neural Network. In: AES 2024 International Conference on Audio for Virtual and Augmented Reality. Audio Engineering Society; 2024. Paper 4. Available from: https://aes.org/publications/elibrary-page/?id=22653
@inproceedings{Nawfal2024_22653,
author = {Nawfal, Ismael and Delikaris Manias, Symeon and Souden, Mehrez and Merimaa, Juha and Atkins, Joshua and McMullin, Elisabeth and Pirhosseinloo, Shadi and Phillips, Daniel},
title = {{Ambisonics Super-Resolution Using A Waveform-Domain Neural Network}},
booktitle = {AES 2024 International Conference on Audio for Virtual and Augmented Reality},
note = {Paper 4},
year = {2024},
month = aug,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=22653}
}
TY - CPAPER
TI - Ambisonics Super-Resolution Using A Waveform-Domain Neural Network
AU - Nawfal, Ismael
AU - Delikaris Manias, Symeon
AU - Souden, Mehrez
AU - Merimaa, Juha
AU - Atkins, Joshua
AU - McMullin, Elisabeth
AU - Pirhosseinloo, Shadi
AU - Phillips, Daniel
T2 - AES 2024 International Conference on Audio for Virtual and Augmented Reality
M1 - Paper 4
PY - 2024
DA - 2024/08/05
UR - https://aes.org/publications/elibrary-page/?id=22653
PB - Audio Engineering Society
LA - en
AB - Ambisonics is a spatial audio format describing a sound field. First-order Ambisonics (FOA) is a popular format comprising only four channels. This limited channel count comes at the expense of spatial accuracy. Ideally one would be able to take the efficiency of a FOA format without its limitations. We have devised a data-driven spatial audio solution that retains the efficiency of the FOA format but achieves quality that surpasses conventional renderers. Utilizing a fully convolutional time-domain audio neural network (Conv-TasNet), we created a solution that takes a FOA input and provides a higher order Ambisonics (HOA) output. This data driven approach is novel when compared to typical physics and psychoacoustic based renderers. Quantitative evaluations showed a 0.6dB average positional mean squared error difference between predicted and actual 3rd order HOA. The median qualitative rating showed an 80% improvement in perceived quality over the traditional rendering approach.
ER -