L. Pham, I. McLoughlin, H. Phan, R. Palaniappan, and Y. Lang, “Bag-of-Features Models Based on C-DNN Network for Acoustic Scene Classification,” in Proc. AES Conference: 2019 AES International Conference on Audio Forensics, Jun. 2019, Paper 12. [Online]. Available: https://aes.org/publications/elibrary-page/?id=20465
Pham L, McLoughlin I, Phan H, Palaniappan R, Lang Y. Bag-of-Features Models Based on C-DNN Network for Acoustic Scene Classification. In: AES Conference: 2019 AES International Conference on Audio Forensics. Audio Engineering Society; 2019. Paper 12. Available from: https://aes.org/publications/elibrary-page/?id=20465
@inproceedings{Pham2019_20465,
author = {Pham, Lam and McLoughlin, Ian and Phan, Huy and Palaniappan, Ramaswamy and Lang, Yue},
title = {{Bag-of-Features Models Based on C-DNN Network for Acoustic Scene Classification}},
booktitle = {AES Conference: 2019 AES International Conference on Audio Forensics},
note = {Paper 12},
year = {2019},
month = jun,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=20465}
}
TY - CPAPER
TI - Bag-of-Features Models Based on C-DNN Network for Acoustic Scene Classification
AU - Pham, Lam
AU - McLoughlin, Ian
AU - Phan, Huy
AU - Palaniappan, Ramaswamy
AU - Lang, Yue
T2 - AES Conference: 2019 AES International Conference on Audio Forensics
M1 - Paper 12
PY - 2019
DA - 2019/06/06
UR - https://aes.org/publications/elibrary-page/?id=20465
PB - Audio Engineering Society
LA - en
AB - This work proposes bag-of-features deep learning models for acoustic scene classi?cation (ASC) – identifying recording locations by analyzing background sound. We explore the effect on classi?cation accuracy of various front-end feature extraction techniques, ensembles of audio channels, and patch sizes from three kinds of spectrogram. The back-end process presents a two-stage learning model with a pre-trained CNN (preCNN) and a post-trained DNN (postDNN). Additionally, data augmentation using the mixup technique is investigated for both the pre-trained and post-trained processes, to improve classi?cation accuracy through increasing class boundary training conditions. Our experiments on the 2018 Challenge on Detection and Classi?cation of Acoustic Scenes and Events - Acoustic Scene Classi?cation (DCASE2018-ASC) subtask 1A and 1B signi?cantly outperform the DCASE2018 reference implementation and approach state-of-the-art performance for each task. Results reveal that the ensemble of multi-spectrogram features and data augmentation is bene?cial to performance.
ER -