Express Paper
Construction of a Japanese air traffic control communication corpus assisted with automatic speech recognition
Express Paper · Paper 381 · October 2025
Abstract
We present a phonetically transcribed corpus of Air Traffic Control (ATC) communications of Japan airports comprising about two hours of conversations between pilots and air traffic controllers. Audio recordings were manually annotated by naive Japanese listeners with limited English proficiency (B2 or above), assisted with three offline Automatic Speech Recognition (ASR) systems: regular Whisper, Whisper fine-tuned with ATC transcriptions from mostly European airports, and Parakeet. For each annotator, a recording was presented along with a single ancillary transcription produced by one of these systems, so that all annotators were equally exposed to the three systems. Automatic and human annotations were compared with those made by an air traffic expert to establish Word Error Rates (WERs). Significant differences in WER were observed between the outputs of ASR systems, with fine-tuned Whisper achieving the best accuracy. Human transcriptions more than halved the WER of automatic ones, achieving similar performance among them regardless of ASR system employed and English proficiency of the transcriber. Together, our findings suggest that ASR-assisted transcriptions could be effective in specialized jargon domains or when field experts are not available. In addition, the new publicly available dataset supports future research in aviation speech recognition.
