Journal Article
Efficient Audio Enhancement With a Differentiable Psychoacoustic Loss
Journal of the Audio Engineering Society · Volume 74 · Issue 6 · pp. 430–444 · June 2026
Abstract
Audio enhancement aims to improve the perceived quality of audio signals. Initially, to address bandwidth extension, this work proposes AEROMambaP, an efficient variant of the AERO super-resolution architecture in which attention and long short-term memory layers are replaced by the Mamba state-space model, together with a newly developed differentiable perceptual loss derived from the perceptual audio quality measure (PAQM). During training, the architecture requires two to fours times less GPU memory than the baseline; during inference, it achieves a 14× speedup while using one-fifth of the GPU memory. When upsampling both a piano dataset and MUSDB18 from 11.025 kHz to 44.1 kHz, subjective listening tests show that AEROMambaP outperforms AERO by 15% in perceived quality scores. Next, to enhance of audio signals highly compressed by lossy coding, it is further proposed AEROMambaPS̄, which applies the same framework but replaces short-time Fourier transform reconstruction losses with the PAQM loss, specifically to enhance MP3-encoded audio at 32 kbps. In listening evaluations, AEROMambaPS̄ achieves 52% higher quality rating than AEROMambaP when restoring compressed audio. These results demonstrate that PAQM-driven training coupled with lightweight state-space modeling yields high perceptual quality and computational efficiency in both band-limited and compressed audio scenarios.
