Conference Paper
Open Access
Complex Ratio Mask Ambisonics-to-Binaural Rendering with Intensity Vector Features and Perceptual Multi-Objective Optimization
AVARIG 2026: Audio for Virtual and Augmented Reality and Immersive Games · Paper 504 · June 2026
Abstract
This paper presents a neural network for binaural rendering of first-order Ambisonics (FOA) signals, enabling immersive spatial audio over headphones without individualized HRTF measurements at inference time. The model operates in the STFT domain using Complex Ratio Masks (CRM): a shared mask pair (left/right ear) is applied via complex-valued multiplication to all four FOA channels with directional weighting, preserving both magnitude and phase. The input representation extends standard spectral features with acoustic intensity vector channels
encoding sound arrival direction at each time-frequency bin. Training uses a multi-objective loss combining SI-SDR, multi-resolution spectral reconstruction, and interaural level/phase difference terms. The four-level UNet backbone with channel-spatial attention totals approximately four million parameters. Evaluation against eight baseline configurationsspanning magnitude-masking variants (28M parameters), a waveform-domain Wave-U-Net (143 M), and the A2B time-domain renderershows that the CRM variant achieves the strongest spectral and spatial fidelity in the cross-domain setting with a seven-fold parameter reduction over the magnitude-masking
baseline while maintaining real-time CPU operation.
