W. H. Nam, K.-R. Kim, J. Kim, and D. Eom, “Audio-Visual-Information-Based Speaker Matching Framework for Selective Hearing in a Recorded Video,” in Proc. AES Convention 153, Oct. 2022, Paper 17. [Online]. Available: https://aes.org/publications/elibrary-page/?id=21897
Nam WH, Kim KR, Kim J, Eom D. Audio-Visual-Information-Based Speaker Matching Framework for Selective Hearing in a Recorded Video. In: AES Convention 153. Audio Engineering Society; 2022. Paper 17. Available from: https://aes.org/publications/elibrary-page/?id=21897
@inproceedings{Nam2022_21897,
author = {Nam, Woo Hyun and Kim, Kyung-Rae and Kim, Jungkyu and Eom, Deokjun},
title = {{Audio-Visual-Information-Based Speaker Matching Framework for Selective Hearing in a Recorded Video}},
booktitle = {AES Convention 153},
note = {Paper 17},
year = {2022},
month = oct,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=21897}
}
TY - CPAPER
TI - Audio-Visual-Information-Based Speaker Matching Framework for Selective Hearing in a Recorded Video
AU - Nam, Woo Hyun
AU - Kim, Kyung-Rae
AU - Kim, Jungkyu
AU - Eom, Deokjun
T2 - AES Convention 153
M1 - Paper 17
PY - 2022
DA - 2022/10/06
UR - https://aes.org/publications/elibrary-page/?id=21897
PB - Audio Engineering Society
LA - en
AB - Selective hearing technology enables users to select a visual object in a video and focus on the desired sound of the object. Since the conventional audio-based voice separation technology generally uses only audio information, it was difficult to know which visual object in the video matched the separated voice. In this paper, to resolve this problem, we propose the audio-visual-information-based speaker matching framework. In this framework, to precisely quantify the matchness between the visual object and the separated voice, we designed the audio-visual feature matching algorithm based on convolutional neural network. The experimental results on the three categories (interview, script-reading, and sing-a-song) of datasets show that the proposed framework can provide the reliable and highly accurate matching relationship.
ER -