H.-G. Kang, W. B. Kleijn, J. Skoglund, and M. Chinen, “Convolutional Transformer for Neural Speech Coding,” in Proc. AES Convention 155, Oct. 2023, Paper 10668. [Online]. Available: https://aes.org/publications/elibrary-page/?id=22249
Kang HG, Kleijn WB, Skoglund J, Chinen M. Convolutional Transformer for Neural Speech Coding. In: AES Convention 155. Audio Engineering Society; 2023. Paper 10668. Available from: https://aes.org/publications/elibrary-page/?id=22249
@inproceedings{Kang2023_22249,
author = {Kang, Hong-Goo and Kleijn, W. Bastiaan and Skoglund, Jan and Chinen, Michael},
title = {{Convolutional Transformer for Neural Speech Coding}},
booktitle = {AES Convention 155},
note = {Paper 10668},
year = {2023},
month = oct,
publisher = {Audio Engineering Society},
url = {https://aes.org/publications/elibrary-page/?id=22249}
}
TY - CPAPER
TI - Convolutional Transformer for Neural Speech Coding
AU - Kang, Hong-Goo
AU - Kleijn, W. Bastiaan
AU - Skoglund, Jan
AU - Chinen, Michael
T2 - AES Convention 155
M1 - Paper 10668
PY - 2023
DA - 2023/10/06
UR - https://aes.org/publications/elibrary-page/?id=22249
PB - Audio Engineering Society
LA - en
AB - In this paper, we propose a Convolutional-Transformer speech codec which utilizes stacks of convolutions and self-attention layers to remove redundant information at the downsampling and upsampling blocks of a U-Net-style encoder-decoder neural codec architecture. We design the Transformers to use channel and temporal attention with any number of attention stages and heads while maintaining causality. This allows us to take into consideration the characteristics of the input vectors and flexibly utilize temporal and channel-wise relationships at different scales when encoding the salient information that is present in speech. This enables our model to reduce the dimensionality of its latent embeddings and improve its quantization efficiency while maintaining quality. Experimental results demonstrate that our approach achieves significantly better performance than convolution-only baselines.
ER -