BIT's System for Multilingual Track

Zhipeng Wang; Yuhang Guo; Shuoying Chen

BIT's System for Multilingual Track

Zhipeng Wang, Yuhang Guo, Shuoying Chen

Add to Favorites

The 20th International Conference on Spoken Language Translation Long Paper

TLDR: This paper describes the system we submitted to the IWSLT 2023 multilingual speech translation track, with input being English speech and output being text in 10 target languages. Our system consists of CNN and Transformer, convolutional neural networks downsample speech features and extract local i

RocketChat
Abstract

You can open the #paper-IWSLT_52 channel in a separate window.

Abstract: This paper describes the system we submitted to the IWSLT 2023 multilingual speech translation track, with input being English speech and output being text in 10 target languages. Our system consists of CNN and Transformer, convolutional neural networks downsample speech features and extract local information, while transformer extract global features and output the final results. In our system, we use speech recognition tasks to pre-train encoder parameters, and then use speech translation corpus to train the multilingual speech translation model. We have also adopted other methods to optimize the model, such as data augmentation, model ensemble, etc. Our system can obtain satisfactory results on test sets of 10 languages in the MUST-C corpus.