Opens in a new tab

AES E-Library

← Back to search

Convention Paper Open Access

Conversational Speech Separation: an Evaluation Study for Streaming Applications

Authors: Morrone, Giovanni; Cornell, Samuele; Zovato, Enrico; Brutti, Alessio; Squartini, Stefano

AES Convention 152 · Paper 10562 · May 2022

Abstract

Continuous speech separation (CSS) is a recently proposed framework which aims at separating each speaker
from an input mixture signal in a streaming fashion. Hereafter we perform an evaluation study on practical design considerations for a CSS system, addressing important aspects which have been neglected in recent works. In particular, we focus on the trade-off between separation performance, computational requirements and output latency showing how an offline separation algorithm can be used to perform CSS with a desired latency. We carry out an extensive analysis on the choice of CSS processing window size and hop size on sparsely overlapped data. We find out that the best trade-off between computational burden and performance is obtained for a window of 5 s.

Details

Published in
AES Convention 152
AES Convention
152
Paper number
10562
Publication date
May 6, 2022
Session subject
Television Audio
Affiliation
Università Politecnica delle Marche, Ancona, Italy; PerVoice S.p.A., Trento, Italy; Fondazione Bruno Kessler, Trento, Italy (See document for exact affiliation information.)
Type
Convention Paper