M. Zanoni, S. Lusardi, P. Bestagini, A. Canclini, A. Sarti, and S. Tubaro, “Efficient Music Identification Approach Based on Local Spectrogram Image Descriptors,” in Proc. AES Convention 142, May 2017, Paper 9763. [Online]. Available: https://aes.org/publications/elibrary-page/?id=18639
Zanoni M, Lusardi S, Bestagini P, Canclini A, Sarti A, Tubaro S. Efficient Music Identification Approach Based on Local Spectrogram Image Descriptors. In: AES Convention 142. Audio Engineering Society; 2017. Paper 9763. Available from: https://aes.org/publications/elibrary-page/?id=18639
@inproceedings{Zanoni2017_18639,
author = {Zanoni, Massimiliano and Lusardi, Stefano and Bestagini, Paolo and Canclini, Antonio and Sarti, Augusto and Tubaro, Stefano},
title = {{Efficient Music Identification Approach Based on Local Spectrogram Image Descriptors}},
booktitle = {AES Convention 142},
note = {Paper 9763},
year = {2017},
month = may,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=18639}
}
TY - CPAPER
TI - Efficient Music Identification Approach Based on Local Spectrogram Image Descriptors
AU - Zanoni, Massimiliano
AU - Lusardi, Stefano
AU - Bestagini, Paolo
AU - Canclini, Antonio
AU - Sarti, Augusto
AU - Tubaro, Stefano
T2 - AES Convention 142
M1 - Paper 9763
PY - 2017
DA - 2017/05/06
UR - https://aes.org/publications/elibrary-page/?id=18639
PB - Audio Engineering Society
LA - en
AB - The diffusion of large music collections has determined the need for algorithms enabling fast song retrieval from query audio excerpts. This is the case of online media sharing platforms that may want to detect copyrighted material. In this paper we start from a proposed state-of-the-art algorithm for robust music matching based on spectrogram comparison leveraging computer vision concepts. We show that it is possible to further optimize this algorithm exploiting more recent image processing techniques and carrying out the analysis on limited temporal windows, still achieving accurate matching performance. The proposed solution is validated on a dataset of 800 songs, reporting an 80% decrease in computational complexity for an accuracy loss of about only 1%.
ER -