M. Piotrowska, G. Korvel, A. Kurowski, B. Kostek, and A. Czyzewski, “Machine Learning Applied to Aspirated and Non-Aspirated Allophone Classification–An Approach Based on Audio “Fingerprinting”,” in Proc. AES Convention 145, Oct. 2018, Paper 10070. [Online]. Available: https://aes.org/publications/elibrary-page/?id=19796
Piotrowska M, Korvel G, Kurowski A, Kostek B, Czyzewski A. Machine Learning Applied to Aspirated and Non-Aspirated Allophone Classification–An Approach Based on Audio “Fingerprinting”. In: AES Convention 145. Audio Engineering Society; 2018. Paper 10070. Available from: https://aes.org/publications/elibrary-page/?id=19796
@inproceedings{Piotrowska2018_19796,
author = {Piotrowska, Magdalena and Korvel, Grazina and Kurowski, Adam and Kostek, Bozena and Czyzewski, Andrzej},
title = {{Machine Learning Applied to Aspirated and Non-Aspirated Allophone Classification–An Approach Based on Audio “Fingerprinting”}},
booktitle = {AES Convention 145},
note = {Paper 10070},
year = {2018},
month = oct,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=19796}
}
TY - CPAPER
TI - Machine Learning Applied to Aspirated and Non-Aspirated Allophone Classification–An Approach Based on Audio “Fingerprinting”
AU - Piotrowska, Magdalena
AU - Korvel, Grazina
AU - Kurowski, Adam
AU - Kostek, Bozena
AU - Czyzewski, Andrzej
T2 - AES Convention 145
M1 - Paper 10070
PY - 2018
DA - 2018/10/06
UR - https://aes.org/publications/elibrary-page/?id=19796
PB - Audio Engineering Society
LA - en
AB - The purpose of this study is to involve both Convolutional Neural Networks and a typical learning algorithm in the allophone classification process. A list of words including aspirated and non-aspirated allophones pronounced by native and non-native English speakers is recorded and then edited and analyzed. Allophones extracted from English speakers’ recordings are presented in the form of two-dimensional spectrogram images and used as input to train the Convolutional Neural Networks. Various settings of the spectral representation are analyzed to determine adequate option for the allophone classification. Then, testing is performed on the basis of non-native speakers’ utterances. The same approach is repeated employing learning algorithm but based on feature vectors. The archived classification results are promising as high accuracy is observed.
ER -