Opens in a new tab

AES E-Library

← Back to search

Convention Paper

Robustness of Speaker Recognition from Noisy Speech Samples and Mismatched Languages

Authors: Al-Noori, Ahmed; Li, Francis F.; Duncan, Philip J.

AES Convention 140 · Paper 9577 · May 2016

Abstract

Speaker recognition systems can typically attain high performance in ideal conditions. However, significant degradations in accuracy are found in channel-mismatched scenarios. Non-stationary environmental noises and their variations are listed at the top of speaker recognition challenges. Gammtone frequency cepstral coefficient method (GFCC) has been developed to improve the robustness of speaker recognition. This paper presents systematic comparisons between performance of GFCC and conventional MFCC-based speaker verification systems with a purposely collected noisy speech data set. Furthermore, the current work extends the experiments to include investigations into language independency features in recognition phases. The results show that GFCC has better verification performance in noisy environments than MFCC. However, the GFCC shows a higher sensitivity to language mismatch between enrollment and recognition phase.

Details

Published in
AES Convention 140
AES Convention
140
Paper number
9577
Publication date
May 6, 2016
Session subject
Human Factors and Interfaces
Affiliation
University of Salford, Salford, Greater Manchester, UK (See document for exact affiliation information.)
Type
Convention Paper