Opens in a new tab

AES E-Library

← Back to search

Convention Paper

Audio Content Annotation, Description, and Management Using Joint Audio Detection, Segmentation, and Classification Techniques

Authors: Vegiris, Christos; Dimoulas, Charalampos; Papanikolaou, George

AES Convention 126 · Paper 7661 · May 2009

Abstract

The current paper focuses on audio content management by means of joint audio segmentation and classification. We concentrate on the separation of typical audio classes, such as silence / background noise, speech, music and their combinations. A compact feature-vector subset is selected by a Correlation feature selection subset evaluation algorithm after the use of EM clustering algorithm on an initial audio data set. Time and spectral parameters are extracted using filter-banks and wavelets in combination with sliding windows and exponential moving averaging techniques. Features are extracted on a point-to-point basis, using the finest possible time resolution, so that each sample can be individually classified to one of the available groups. Clustering algorithms like EM or Simple K-means are tested to evaluate the final point-to-point classification result, therefore the joint audio detection-classification indexes. The extracted audio detection, segmentation and classification results can be incorporated into appropriate description schemes that would annotate audio events / segments for content description and management purposes.

Details

Published in
AES Convention 126
AES Convention
126
Paper number
7661
Publication date
May 6, 2009
Session subject
Recording, Reproduction, and Delivery
Affiliation
Aristotle University of Thessaloniki, Thessaloniki, Greece (See document for exact affiliation information.)
Type
Convention Paper