R. Schramm and E. Benetos, “Automatic Transcription of a Cappella recordings from Multiple Singers,” in Proc. AES Conference: 2017 AES International Conference on Semantic Audio, Jun. 2017, Paper 3-2. [Online]. Available: https://aes.org/publications/elibrary-page/?id=18757
Schramm R, Benetos E. Automatic Transcription of a Cappella recordings from Multiple Singers. In: AES Conference: 2017 AES International Conference on Semantic Audio. Audio Engineering Society; 2017. Paper 3-2. Available from: https://aes.org/publications/elibrary-page/?id=18757
@inproceedings{Schramm2017_18757,
author = {Schramm, Rodrigo and Benetos, Emmanouil},
title = {{Automatic Transcription of a Cappella recordings from Multiple Singers}},
booktitle = {AES Conference: 2017 AES International Conference on Semantic Audio},
note = {Paper 3-2},
year = {2017},
month = jun,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=18757}
}
TY - CPAPER
TI - Automatic Transcription of a Cappella recordings from Multiple Singers
AU - Schramm, Rodrigo
AU - Benetos, Emmanouil
T2 - AES Conference: 2017 AES International Conference on Semantic Audio
M1 - Paper 3-2
PY - 2017
DA - 2017/06/06
UR - https://aes.org/publications/elibrary-page/?id=18757
PB - Audio Engineering Society
LA - en
AB - This work presents a spectrogram factorisation method applied to automatic music transcription of a cappella performances with multiple singers. A variable-Q transform representation of the audio spectrogram is factorised with the help of a 6-dimensional sparse dictionary which contains spectral templates of vowel vocalizations. A post-processing step is proposed to remove false positive pitch detections through a binary classifier, where overtone-based features are used as input. Preliminary experiments have shown promising multi-pitch detection results when applied to audio recordings of Bach Chorales and Barbershop music. Comparisons made with alternative methods have shown that our approach increases the number of true positive pitch detections while the post-processing step keeps the number of false positives lower than those measured in comparative approaches.
ER -