L. Pham et al., “A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators,” in Proc. Convention Paper, May 2026, Paper 10264. [Online]. Available: https://aes.org/publications/elibrary-page/?id=23179
Pham L, Vu K, Tran D, Freitter S, Hasenbalg M, Antonutti D, Fischinger D, Schindler A, Boyer M, McLoughlin I. A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators. In: Convention Paper. Audio Engineering Society; 2026. Paper 10264. Available from: https://aes.org/publications/elibrary-page/?id=23179
@inproceedings{Pham2026_23179,
author = {Pham, Lam and Vu, Khoi and Tran, Dat and Freitter, Simon and Hasenbalg, Marcel and Antonutti, Davide and Fischinger, David and Schindler, Alexander and Boyer, Martin and McLoughlin, Ian},
title = {{A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators}},
note = {Paper 10264},
year = {2026},
month = may,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=23179}
}
TY - CPAPER
TI - A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators
AU - Pham, Lam
AU - Vu, Khoi
AU - Tran, Dat
AU - Freitter, Simon
AU - Hasenbalg, Marcel
AU - Antonutti, Davide
AU - Fischinger, David
AU - Schindler, Alexander
AU - Boyer, Martin
AU - McLoughlin, Ian
M1 - Paper 10264
PY - 2026
DA - 2026/05/28
UR - https://aes.org/publications/elibrary-page/?id=23179
PB - Audio Engineering Society
LA - en
AB - In this paper, we analyze how the provision of bonafide resources and AI generated utterances in training and testing datasets affect the performance and the generality of Deepfake Speech Detection (DSD) models. To this end, we first propose a deep-learning based model based on state-of-the-art architectures, referred to as the baseline. We then conducted experiments using the baseline to determine how the combination of bonafide resources and AI generated utterances affect the threshold score used to detect fake or bonafide input audio during the inference process. Given the experimental results, a dataset, which re-uses public Deepfake Speech Detection (DSD) datasets and balances between Bonafide Resources (BR) and AI-Generated (AG) utterances, is proposed. We then train various deep-learning based models on the proposed dataset and conduct cross-dataset evaluation on different benchmark datasets. The cross-dataset evaluation results prove that the balance of BR and AG utterances is a key factor in training and achieving generalisability in a Deepfake Speech Detection (DSD) model.
ER -