Opens in a new tab

AES E-Library

← Back to search

Convention Paper Open Access

Comparison of Audio Spectral Features in a Convolutional Neural Network

Authors: Vines, Greg; Nemer, Elias

AES Convention 153 · Paper 10634 · October 2022

Abstract

Time-Frequency transformation and spectral representations of audio signals are commonly used in various machine learning applications. Typically the Mel-Spectrogram is used to create the input features to the network justified by the Mel scale’s human auditory system basis. In this paper, we compare several spectral features in a gender detection speech model comparing their performance and showing that the Mel-Spectrogram is not always the best choice for input features.

Details

Published in
AES Convention 153
AES Convention
153
Paper number
10634
Publication date
October 6, 2022
Session subject
Applications in Audio
Affiliation
San Diego, CA, USA; San Diego, CA, USA; (See document for exact affiliation information.)
Type
Convention Paper