G. Vines and E. Nemer, “Comparison of Audio Spectral Features in a Convolutional Neural Network,” in Proc. AES Convention 153, Oct. 2022, Paper 10634. [Online]. Available: https://aes.org/publications/elibrary-page/?id=21963
Vines G, Nemer E. Comparison of Audio Spectral Features in a Convolutional Neural Network. In: AES Convention 153. Audio Engineering Society; 2022. Paper 10634. Available from: https://aes.org/publications/elibrary-page/?id=21963
@inproceedings{Vines2022_21963,
author = {Vines, Greg and Nemer, Elias},
title = {{Comparison of Audio Spectral Features in a Convolutional Neural Network}},
booktitle = {AES Convention 153},
note = {Paper 10634},
year = {2022},
month = oct,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=21963}
}
TY - CPAPER
TI - Comparison of Audio Spectral Features in a Convolutional Neural Network
AU - Vines, Greg
AU - Nemer, Elias
T2 - AES Convention 153
M1 - Paper 10634
PY - 2022
DA - 2022/10/06
UR - https://aes.org/publications/elibrary-page/?id=21963
PB - Audio Engineering Society
LA - en
AB - Time-Frequency transformation and spectral representations of audio signals are commonly used in various machine learning applications. Typically the Mel-Spectrogram is used to create the input features to the network justified by the Mel scale’s human auditory system basis. In this paper, we compare several spectral features in a gender detection speech model comparing their performance and showing that the Mel-Spectrogram is not always the best choice for input features.
ER -