Opens in a new tab

AES E-Library

← Back to search

Journal Article

Dual-Residual Transformer Network for Speech Recognition

Authors: Duan, Zhikui; Gao, Guozhi; Chen, Jiawei; Li, Shiren; Ruan, Jinbiao; Yang, Guangguang; Yu, Xinmei

Journal of the Audio Engineering Society · Volume 70 · Issue 10 · pp. 871–881 · October 2022

Abstract

The Transformer, an attention-based encoder-decoder network, has recently become the prevailing model for automatic speech recognition because of its high recognition accuracy. However, the convergence speed of the Transformer is not that optimal. In order to address this problem, a structure called Dual-Residual Transformer Network (DRTNet), which has fast convergence speed, is proposed. In DRTNet, a direct path is added in the encoder and decoder layers to propagate features with the inspiration of the structure proposed in ResNet. Moreover, this architecture can also fuse features, which tends to improve the model performance. Specifically, the input of the current layer is the integration of the input and output of the previous layer. Empirical evaluation of the proposed DRTNet has been conducted on two public datasets, which are AISHELL-1 and HKUST, respectively. Experimental results on these two datasets show that DRTNet has faster convergence speed and better performance.

Details

Publication
Journal of the Audio Engineering Society
Volume
70
Issue
10
Pages
871–881
Publication date
October 6, 2022
Affiliation
Foshan University, Foshan, China (See document for exact affiliation information.)
Type
Journal Article