Opens in a new tab

AES E-Library

← Back to search

Convention Paper

Modal Representations for Audio Deep Learning

Authors: Skare, Travis; Abel, Jonathan S.; Smith, III, Julius O.

AES Convention 147 · Paper 10248 · October 2019

Abstract

Deep learning models for both discriminative and generative tasks have a choice of domain representation. For audio, candidates are often raw waveform data, spectral data, transformed spectral data, or perceptual features. For deep learning tasks related to modal synthesizers or processors, we propose new, modal representations for data. We experiment with representations such as an N-hot binary vector of frequencies, or learning a set of modal filterbank coefficients directly. We use these representations discriminatively–classifying cymbal model based on samples–as well as generatively. An intentionally naive application of a basic modal representation to a CVAE designed for MNIST digit images quickly yielded results, which we found surprising given less prior success when using traditional representations like a spectrogram image. We discuss applications for Generative Adversarial Networks, towards creating a modal reverberator generator.

Details

Published in
AES Convention 147
AES Convention
147
Paper number
10248
Publication date
October 6, 2019
Session subject
Posters: Audio Signal Processing
Affiliation
CCRMA, Stanford University, Stanford, CA, USA (See document for exact affiliation information.)
Type
Convention Paper