Opens in a new tab

AES E-Library

← Back to search

Convention Paper

A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators

Authors: Pham, Lam; Vu, Khoi; Tran, Dat; Freitter, Simon; Hasenbalg, Marcel; Antonutti, Davide; Fischinger, David; Schindler, Alexander; Boyer, Martin; McLoughlin, Ian

Convention Paper · Paper 10264 · May 2026

Abstract

In this paper, we analyze how the provision of bonafide resources and AI generated utterances in training and testing datasets affect the performance and the generality of Deepfake Speech Detection (DSD) models. To this end, we first propose a deep-learning based model based on state-of-the-art architectures, referred to as the baseline. We then conducted experiments using the baseline to determine how the combination of bonafide resources and AI generated utterances affect the threshold score used to detect fake or bonafide input audio during the inference process. Given the experimental results, a dataset, which re-uses public Deepfake Speech Detection (DSD) datasets and balances between Bonafide Resources (BR) and AI-Generated (AG) utterances, is proposed. We then train various deep-learning based models on the proposed dataset and conduct cross-dataset evaluation on different benchmark datasets. The cross-dataset evaluation results prove that the balance of BR and AG utterances is a key
factor in training and achieving generalisability in a Deepfake Speech Detection (DSD) model.

Details

AES Convention
160
Paper number
10264
Publication date
May 28, 2026
Session subject
AI and Machine Learning in Audio, Audio Applications and Technologies
Affiliation
Austrian Institute of Technology; Austrian Institute of Technology; Austrian Institute of Technology; Austrian Institute of Technology; Austrian Institute of Technology; Austrian Institute of Technology; Austrian Institute of Technology; FPT University; FPT University; Singapore Institute of Technology (See document for exact affiliation information.)
Type
Convention Paper