Opens in a new tab

AES E-Library

← Back to search

Journal Article

Crowdsourcing Audio Semantics by Means of Hybrid Bimodal Segmentation with Hierarchical Classification

Authors: Vrysis, Lazaros; Tsipas, Nikolaos; Dimoulas, Charalampos; Papanikolaou, George

Journal of the Audio Engineering Society · Volume 64 · Issue 12 · pp. 1042–1054 · December 2016

Abstract

The task of general audio detection and segmentation is quite common in contemporary audio applications where computationally intensive processes are frequently involved. Machine learning is usually employed along with user-enabled data labeling that is intended to detect, segment, and semantically annotate the relevant audio events. This work focuses on a generic audio detection and classification method that combines hierarchical bimodal segmentation with hybrid pattern classification at different temporal resolutions. This paper presents the algorithmic perspective of a mobile back-end system to facilitate the construction, validation, and continuous update of generic audio ground-truth data. The goal is the implementation of a system that is capable of performing well in different conditions without relying on complicated pattern recognition systems and taxonomies. For this reason, minimal prior knowledge is necessary so that there is consistent behavior for different input signals and computational environments. Novel “classification confidence” metrics are implemented.

Details

Publication
Journal of the Audio Engineering Society
Volume
64
Issue
12
Pages
1042–1054
Publication date
December 6, 2016
Affiliation
Aristotle University of Thessaloniki, Thessaloniki, Greece (See document for exact affiliation information.)
Type
Journal Article