Opens in a new tab

AES E-Library

← Back to search

Engineering Brief

Generative Modeling of Metadata for Machine Learning Based Audio Content Classification

Authors: Bharitkar, Sunil G.

AES Convention 147 · Paper 564 · October 2019

Abstract

Automatic content classification technique is an essential tool in multimedia applications. Present research for audio-based classifiers look at short- and long-term analysis of signals, using both temporal and spectral features. In this paper we present a neural network to classify between the movie (cinematic, TV shows), music, and voice using metadata contained in either the audio/video stream. Towards this end, statistical models of the various metadata are created since a large metadata dataset is not available. Subsequently, synthetic metadata are generated from these statistical models, and the synthetic metadata is input to the ML classifier as feature vectors. The resulting classifier is then able to classify real-world content (e.g., YouTube) with an accuracy ˜ 90% with very low latency (viz., ˜ on an average 7 ms) based on real-world metadata.

Details

Published in
AES Convention 147
AES Convention
147
Paper number
564
Publication date
October 6, 2019
Session subject
Applications in Audio
Affiliation
HP Labs., Inc., San Francisco, CA, USA (See document for exact affiliation information.)
Type
Engineering Brief