P. S. Popolo, R. W. Sanders, and I. R. Titze, “Quantifying the Speaking Voice: Generating a Speaker Code as a Means of Speaker Identification Using a Simple Code-Matching Technique,” in Proc. AES Convention 123, Oct. 2007, Paper 7274. [Online]. Available: https://aes.org/publications/elibrary-page/?id=14332
Popolo PS, Sanders RW, Titze IR. Quantifying the Speaking Voice: Generating a Speaker Code as a Means of Speaker Identification Using a Simple Code-Matching Technique. In: AES Convention 123. Audio Engineering Society; 2007. Paper 7274. Available from: https://aes.org/publications/elibrary-page/?id=14332
@inproceedings{Popolo2007_14332,
author = {Popolo, Peter S. and Sanders, Richard W. and Titze, Ingo R.},
title = {{Quantifying the Speaking Voice: Generating a Speaker Code as a Means of Speaker Identification Using a Simple Code-Matching Technique}},
booktitle = {AES Convention 123},
note = {Paper 7274},
year = {2007},
month = oct,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=14332}
}
TY - CPAPER
TI - Quantifying the Speaking Voice: Generating a Speaker Code as a Means of Speaker Identification Using a Simple Code-Matching Technique
AU - Popolo, Peter S.
AU - Sanders, Richard W.
AU - Titze, Ingo R.
T2 - AES Convention 123
M1 - Paper 7274
PY - 2007
DA - 2007/10/06
UR - https://aes.org/publications/elibrary-page/?id=14332
PB - Audio Engineering Society
LA - en
AB - This paper looks at a methodology of quantifying the speaking voice, by which temporal and spectral features of the voice are extracted and processed to create a numeric code that identifies speakers, so those speakers can be searched in a database much like fingerprints. The parameters studied include: (1) average fundamental frequency (F0) of the speech signal over time, (2) standard deviation of the F0, (3) the slope and (4) sign of the FO contour, (5) the average energy, (6) the standard deviation of the energy, (7) the spectral energy contained from 50 Hz to 1,000 Hz, (8) the spectral energy from 1,000 Hz to 5,000 Hz, (9) the Alpha Ratio, (10) the average speaking rate, and (11) the total duration of the spoken sentence.
ER -