Opens in a new tab

AES E-Library

← Back to search

Conference Paper Open Access

Calibrating neural networks for synthetic speech detection: A likelihood-ratio-based approach

Authors: Cuccovillo, Luca; Aichroth, Patrick; Köllmer, Thomas

AES 2024 International Conference on Audio Forensics · Paper 3 · June 2024

Abstract

In this paper, we introduce a calibration procedure designed to convert the uncalibrated output scores of neural networks for synthetic speech detection into calibrated and interpretable likelihood ratios. This procedure is based on the assumption that the networks subject to calibration are deterministic and have undergone training until they reached convergence. Provided these conditions are satisfied, it is then possible to transform their output values into likelihood ratios using a minimal set of validation and calibration data, eliminating the need for retraining the models. We successfully tested the entire workflow on a state-of-the-art network example, demonstrating not only its effectiveness in calibration but also its ability to enhance fault tolerance against inadequate inputs.

Details

Published in
AES 2024 International Conference on Audio Forensics
Paper number
3
Publication date
June 17, 2024
Session subject
Synthetic Speech Detection; Likelihood-ratio Analysis; Explainable Artificial Intelligence; Neural Networks
Affiliation
Fraunhofer Institute for Digital Media Technology IDMT; Fraunhofer Institute for Digital Media Technology IDMT; Fraunhofer Institute for Digital Media Technology IDMT (See document for exact affiliation information.)
Type
Conference Paper