Conference Paper
S^3MASH: Spatial Sound Scene Matching using Single-Channel Audio
AES 2024 International Conference on Audio for Virtual and Augmented Reality · Paper 15 · August 2024
Abstract
This paper describes a novel approach for recording and binaurally reproducing spatial sound scenes using the audio from a single microphone. This is realised by recording the sound scene using both a microphone array, which potentially comprises more affordable and lower quality capsules, and a monophonic microphone, possibly featuring a higher quality capsule. By adopting a perceptually motivated sound-field model and estimating the models spatial parameters, it is possible to define target time-frequency-dependent binaural spatial covariance matrices (SCMs). The actual binaural signals can then be synthesised using an adaptive SCM matching renderer, which takes only the higher-quality monophonic audio signal as input. A perceptual study was conducted to compare this novel processing approach, using a tetrahedral array and an omnidirectional microphone, against binaural renderings achieved through traditional Ambisonic means, when using four- and 32-channel arrays. The results show that, despite utilising only a monophonic signal for the spatialisation, the proposed approach yielded binaural renderings that are perceptually in-between the two conventional Ambisonic array renderings, with regards to their perceived spatial accuracy.
