Opens in a new tab

AES E-Library

← Back to search

Conference Paper

Ambisonics Super-Resolution Using A Waveform-Domain Neural Network

Authors: Nawfal, Ismael; Delikaris Manias, Symeon; Souden, Mehrez; Merimaa, Juha; Atkins, Joshua; McMullin, Elisabeth; Pirhosseinloo, Shadi; Phillips, Daniel

AES 2024 International Conference on Audio for Virtual and Augmented Reality · Paper 4 · August 2024

Abstract

Ambisonics is a spatial audio format describing a sound field. First-order Ambisonics (FOA) is a popular format comprising only four channels. This limited channel count comes at the expense of spatial accuracy. Ideally one would be able to take the efficiency of a FOA format without its limitations. We have devised a data-driven spatial audio solution that retains the efficiency of the FOA format but achieves quality that surpasses conventional renderers. Utilizing a fully convolutional time-domain audio neural network (Conv-TasNet), we created a solution that takes a FOA input and provides a higher order Ambisonics (HOA) output. This data driven approach is novel when compared to typical physics and psychoacoustic based renderers. Quantitative evaluations showed a 0.6dB average positional mean squared error difference between predicted and actual 3rd order HOA. The median qualitative rating showed an 80% improvement in perceived quality over the traditional rendering approach.

Details

Published in
AES 2024 International Conference on Audio for Virtual and Augmented Reality
Paper number
4
Publication date
August 5, 2024
Session subject
Audio for Virtual and Augmented Reality
Affiliation
Apple; Apple; Apple; Apple; Apple; Apple; Apple; Apple (See document for exact affiliation information.)
Type
Conference Paper