Opens in a new tab

AES E-Library

← Back to search

Express Paper Open Access

The Ambisonic Denoising Paradox: U-Net Processing Degrades ASR Transcription Quality for Medical Speech

Authors: Zaporowski, Szymon; Mróz, Bartłomiej

Express Paper · Paper 451 · May 2026

Abstract

Spatial audio recording using higher-order Ambisonics offers rich directional information for medical speech capture, yet challenging hospital acoustic environments motivate preprocessing with neural denoising algorithms. This study investigates whether U-Net-based denoising of third-order ambisonic recordings improves automatic speech recognition (ASR) quality for medical applications. We developed the Medical Immersive Audio Corpus (MIAC), comprising 1,759 utterances (6.43 hours) of Polish medical speech recorded with a Zylia ZM-1 microphone in uncontrolled hospital environments, capturing 16-channel third-order Ambisonics across multiple specializations including thyroid ultrasonography, surgical procedures, and general diagnostics. We applied a U-Net architecture with dual attention mechanisms trained using the Noise2Noise paradigm to denoise the corpus, then evaluated transcription quality using ten Whisper ASR models ranging from 39 million to 1.55 billion parameters, including domain-adapted medical variants. Surprisingly, we discovered a noise reduction paradox where denoising degraded transcription quality for seven of ten models, with statistically significant increases in Word Error Rate (WER) and Character Error Rate (CER) for general-purpose base, small, and medium models. Only the domain adapted whisper-medium-68000-abbr model showed statistically significant improvement (p = 0:0008), while large-scale models (large-v2, large-v3) exhibited robustness with negligible changes. Effect sizes remained small (Cohens d < 0:2) across all models. These counterintuitive findings suggest modern ASR systems implicitly utilize background noise characteristics as informative features, and that preprocessing pipelines should be reconsidered for domain-specific applications. Our results provide practical guidance for medical speech processing system design.

Details

AES Convention
160
Paper number
451
Publication date
May 28, 2026
Session subject
AI and Machine Learning in Audio, Audio Processing, Recording, Production, and Reproduction
Affiliation
Department of Multimedia Systems, Gda´nsk University of Technology; Department of Multimedia Systems, Gda´nsk University of Technology (See document for exact affiliation information.)
Type
Express Paper