Opens in a new tab

AES E-Library

← Back to search

Conference Paper

Quantifying the Speaking Voice: Further Investigation Into Speaker Identification by a Simple Code-Matching Technique

Authors: Sanders, Richard W.; Smith, Jeff M.

AES Conference: 33rd International Conference: Audio Forensics-Theory and Practice · Paper 5 · June 2008

Abstract

This paper reports on the techniques refined for a method of speaker identification through the automated comparison of spectral, timbral, and temporal features unique to an individual’s speech production. This method was first described in Convention Paper 7274 presented by the co-author of this paper, Richard Sanders, at the 123rd Convention of the Audio Engineering Society. Since its first publication, the system (now referred to as SIDNI or Speaker Identification by Numerical Imprint) has improved from 79% correct identifications in 78 comparisons from the speech of 26 males to 100% correct identifications in 150 comparisons from the speech of 50 males. This paper will provide more information on these results and the results of several other tests while also elaborating on the specific speech characteristics exploited by the system and their potential for identification. Some characteristics include: average fundamental speaking frequency, ratio of spectral densities below 1 kHz to those above 1 kHz (Alpha ratio), average rate of vowels, jitter, and shimmer.

Details

Published in
AES Conference: 33rd International Conference: Audio Forensics-Theory and Practice
Paper number
5
Publication date
June 6, 2008
Session subject
Audio Forensics: Voice Identification
Affiliation
University of Denver, Colorado (See document for exact affiliation information.)
Type
Conference Paper