Convention Paper
Open Access
Confidently Wrong: Evaluating AudioSet-Trained Models Under Real-World Deployment
Convention Paper · Paper 10274 · May 2026
Abstract
Audio event-classification models trained on AudioSet are widely adopted and form a central component of the state of the art in machine listening, yet their behavior when deployed in complex, open acoustic environments remains largely unexplored. In this study, we evaluate several AudioSet-pretrained architecturesparticularly models from the PANNs family, including MobileNetV2, Wavegram LogMel CNN14, and the transformer-based PaSST modelwhen applied to a real operational scenario at the commercial Port of Valencia, Spain. We observed
a recurring but model-dependent behavior: the models frequently assigned relatively high probability to the class Music for non-musical industrial and transportation sounds. These events included train wheels squealing, motorcycle acceleration, emergency sirens, and reversing beepssound categories that are common in port logistics environments but acoustically different from music. By analyzing the probability distributions output by the models, we demonstrate that this erroneous Music activation is not limited to isolated cases but is observed across multiple
architectures, with varying intensity depending on the model and sound category. Our findings highlight limitations in the robustness and domain generalization of AudioSet-derived models and emphasize the need for targeted adaptation techniques when deploying them in real industrial settings.
