Opens in a new tab

AES E-Library

← Back to search

Express Paper

Audio-Visual-Information-Based Speaker Matching Framework for Selective Hearing in a Recorded Video

Authors: Nam, Woo Hyun; Kim, Kyung-Rae; Kim, Jungkyu; Eom, Deokjun

AES Convention 153 · Paper 17 · October 2022

Abstract

Selective hearing technology enables users to select a visual object in a video and focus on the desired sound of the object. Since the conventional audio-based voice separation technology generally uses only audio information, it was difficult to know which visual object in the video matched the separated voice. In this paper, to resolve this problem, we propose the audio-visual-information-based speaker matching framework. In this framework, to precisely quantify the matchness between the visual object and the separated voice, we designed the audio-visual feature matching algorithm based on convolutional neural network. The experimental results on the three categories (interview, script-reading, and sing-a-song) of datasets show that the proposed framework can provide the reliable and highly accurate matching relationship.

Details

Published in
AES Convention 153
AES Convention
153
Paper number
17
Publication date
October 6, 2022
Session subject
Applications in Audio
Affiliation
Samsung Research, Samsung Electronics; Samsung Research, Samsung Electronics; Samsung Research, Samsung Electronics; Samsung Research, Samsung Electronics (See document for exact affiliation information.)
Type
Express Paper