D. Damm, H. Grohganz, F. Kurth, S. Ewert, and M. Clausen, “SyncTS: Automatic Synchronization of Speech and Text Documents,” in Proc. AES Conference: 42nd International Conference: Semantic Audio, Jul. 2011, Paper 2-2. [Online]. Available: https://aes.org/publications/elibrary-page/?id=15962
Damm D, Grohganz H, Kurth F, Ewert S, Clausen M. SyncTS: Automatic Synchronization of Speech and Text Documents. In: AES Conference: 42nd International Conference: Semantic Audio. Audio Engineering Society; 2011. Paper 2-2. Available from: https://aes.org/publications/elibrary-page/?id=15962
@inproceedings{Damm2011_15962,
author = {Damm, David and Grohganz, Harald and Kurth, Frank and Ewert, Sebastian and Clausen, Michael},
title = {{SyncTS: Automatic Synchronization of Speech and Text Documents}},
booktitle = {AES Conference: 42nd International Conference: Semantic Audio},
note = {Paper 2-2},
year = {2011},
month = jul,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=15962}
}
TY - CPAPER
TI - SyncTS: Automatic Synchronization of Speech and Text Documents
AU - Damm, David
AU - Grohganz, Harald
AU - Kurth, Frank
AU - Ewert, Sebastian
AU - Clausen, Michael
T2 - AES Conference: 42nd International Conference: Semantic Audio
M1 - Paper 2-2
PY - 2011
DA - 2011/07/06
UR - https://aes.org/publications/elibrary-page/?id=15962
PB - Audio Engineering Society
LA - en
AB - In this paper, we present an automatic approach for aligning speech signals to corresponding text documents. For this sake, we propose to first use text-to-speech synthesis (TTS) to obtain a speech signal from the textual representation. Subsequently, both speech signals are transformed to sequences of audio features which are then time-aligned using a variant of greedy dynamic time-warping (DTW). The proposed approach is both efficient (with linear running time), computationally simple, and does not rely on a prior training phase as it is necessary when using HMM-based approaches. It benefits from the combination of a) a novel type of speech feature, being correlated to the phonetic progression of speech, b) a greedy left-to-right variant of DTW, and c) the TTS-based approach for creating a feature representation from the input text documents. The feasibility of the proposed method is demonstrated in several experiments.
ER -