X. Zhou, F. Ke, C. Shu, G. Ren, and M. F. Bocko, “High-Precision Score-Based Audio Indexing Using Hierarchical Dynamic Time Warping,” in Proc. AES Convention 135, Oct. 2013, Paper 8931. [Online]. Available: https://aes.org/publications/elibrary-page/?id=16981
Zhou X, Ke F, Shu C, Ren G, Bocko MF. High-Precision Score-Based Audio Indexing Using Hierarchical Dynamic Time Warping. In: AES Convention 135. Audio Engineering Society; 2013. Paper 8931. Available from: https://aes.org/publications/elibrary-page/?id=16981
@inproceedings{Zhou2013_16981,
author = {Zhou, Xiang and Ke, Fangyu and Shu, Cheng and Ren, Gang and Bocko, Mark F.},
title = {{High-Precision Score-Based Audio Indexing Using Hierarchical Dynamic Time Warping}},
booktitle = {AES Convention 135},
note = {Paper 8931},
year = {2013},
month = oct,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=16981}
}
TY - CPAPER
TI - High-Precision Score-Based Audio Indexing Using Hierarchical Dynamic Time Warping
AU - Zhou, Xiang
AU - Ke, Fangyu
AU - Shu, Cheng
AU - Ren, Gang
AU - Bocko, Mark F.
T2 - AES Convention 135
M1 - Paper 8931
PY - 2013
DA - 2013/10/06
UR - https://aes.org/publications/elibrary-page/?id=16981
PB - Audio Engineering Society
LA - en
AB - We propose a novel audio signal processing algorithm of high-precision score-based audio indexing that accurately maps a music score with its corresponding audio. Specifically we improve the time precision of existing score-audio alignment algorithms to find the accurate positions of audio onsets and offsets. We achieve higher time precision by (1) improving the resolution of alignment sequences, and (2) admitting a hierarchy of spectrographic analysis results as audio alignment features. The performance of our proposed algorithm is testified by comparing the segmentation results with manually composed reference datasets. Our proposed algorithm achieves robust alignment results and enhanced segmentation accuracy and thus is suitable for audio engineering applications such as automatic music production and human-media interactions.
ER -