D. F. Yela, S. Ewert, D. Fitzgerald, and M. Sandler, “On the Importance of Temporal Context in Proximity Kernels: A Vocal Separation Case Study,” in Proc. AES Conference: 2017 AES International Conference on Semantic Audio, Jun. 2017, Paper 1-2. [Online]. Available: https://aes.org/publications/elibrary-page/?id=18752
Yela DF, Ewert S, Fitzgerald D, Sandler M. On the Importance of Temporal Context in Proximity Kernels: A Vocal Separation Case Study. In: AES Conference: 2017 AES International Conference on Semantic Audio. Audio Engineering Society; 2017. Paper 1-2. Available from: https://aes.org/publications/elibrary-page/?id=18752
@inproceedings{Yela2017_18752,
author = {Yela, Delia Fano and Ewert, Sebastian and Fitzgerald, Derry and Sandler, Mark},
title = {{On the Importance of Temporal Context in Proximity Kernels: A Vocal Separation Case Study}},
booktitle = {AES Conference: 2017 AES International Conference on Semantic Audio},
note = {Paper 1-2},
year = {2017},
month = jun,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=18752}
}
TY - CPAPER
TI - On the Importance of Temporal Context in Proximity Kernels: A Vocal Separation Case Study
AU - Yela, Delia Fano
AU - Ewert, Sebastian
AU - Fitzgerald, Derry
AU - Sandler, Mark
T2 - AES Conference: 2017 AES International Conference on Semantic Audio
M1 - Paper 1-2
PY - 2017
DA - 2017/06/06
UR - https://aes.org/publications/elibrary-page/?id=18752
PB - Audio Engineering Society
LA - en
AB - Musical source separation methods exploit source-specific spectral characteristics to facilitate the decomposition process. Kernel Additive Modelling (KAM) models a source applying robust statistics to time-frequency bins as specified by a source-specific kernel, a function defining similarity between bins. Kernels in existing approaches are typically defined using metrics between single time frames. In the presence of noise and other sound sources information from a single-frame, however, turns out to be unreliable and often incorrect frames are selected as similar. In this paper, we incorporate a temporal context into the kernel to provide additional information stabilizing the similarity search. Evaluated in the context of vocal separation, our simple extension led to a considerable improvement in separation quality compared to previous kernels.
ER -