Bridging the Modality Gap for Speech-to-Text Translation
Yuchen Liu, Junnan Zhu, Jiajun Zhang, Chengqing Zong
Introduction
Speech-to-Text translation (ST) aims at translating speech in one language into text in another language, which can be widely applied to conference speech, cross-border service, international business talk, academic forum, etc. Most existing approaches to speech-to-text translation are based on the pipeline paradigm, which first transforms the speech into text via an automatic speech recognition (ASR) system and then translates the transcribed text into the target language by a text-based machine translation model Ney (1999); Kanthak et al. (2005); Mathias and Byrne (2006).
Recently, the end-to-end speech-to-text translation model has attracted more attention due to its advantages over the pipeline paradigm, such as low latency, alleviation of error propagation, and fewer parameters Weiss et al. (2017); Bérard et al. (2018); Bansal et al. (2018); Jia et al. (2019); Sperber et al. (2019a). Previous studies employ an encoder-decoder structure to directly learn the mapping relationship between the speech input and text sequences in the target language. However, modality differences between speech and text result in the difficulty of training such a model, making its performance usually much inferior to the corresponding text-based neural machine translation (NMT) model.
We observe that there are three main modality differences between speech and text which affect the speech representational capacity. First, the length of frame-level speech features is much longer than that of the corresponding text data, which hinders the model from learning alignments between the input and the output sequence. Second, text data are generally digitalized into trainable embeddings, while speech features are calculated by hand-crafted filters and fixed during training, resulting in a lack of semantic information. Third, text data contain less noise with low uncertainty, while the speech is variational and easily affected by factors like the speech rate of the different speaker, silence, and noise, which leads to the poor robustness of the model. To make matter worse, the encoder in the ST model is overloaded since it requires to both learn acoustic information and extract semantic knowledge from the speech input simultaneously.
In order to release the burden of speech translation encoder and help it learn better representation, the text representation learned by the NMT model can be utilized as a guidance for speech representation learning. To achieve that, we propose a Speech-to-Text Adaptation for Speech Translation (STAST) model which aims to improve the end-to-end model performance by bridging the modality gap between speech and text. Specifically, the representations of two modalities need to be consistent in terms of length and share in the same latent space. Therefore, we decouple speech translation encoder into three parts, including an acoustic encoder, a shrink mechanism, and a semantic encoder. The shrink mechanism can filter redundant hidden states based on the spike-like label posterior probabilities generated by the CTC module, which has the ability to solve length inconsistency problems. To ensure the representation of two modalities being in the same semantic latent space, we apply the multi-task learning method and completely integrate the MT model into the ST model by sharing the semantic encoder and decoder. To further close the distance of two representations, we apply a cross-modal adaptation method, which can afford speech representation more semantic information. Experimental results on two speech translation corpora have illustrated the effectiveness of our proposed model.
The contributions of this paper are as follows:
We propose the STAST model which can improve speech translation performance by bridging the modality gap between speech and text.
Experimental results have demonstrated that our proposed method outperforms existing methods and achieves the new state-of-the-art performance.
Our method can easily leverage extra data and has shown especially superiority in low-resource scenarios.
Related Work
Speech Translation. Traditional studies on speech translation are based on the pipeline paradigm which consists of an ASR model and an MT model Ney (1999); Kanthak et al. (2005); Mathias and Byrne (2006). Focusing on how to improve the translation robustness on spoken language domain and combine the separate ASR and MT models, previous studies propose lattice-to-sequence models, synthetic data augment, and domain adaptation techniques Lavie et al. (1996); Waibel and Fugen (2008); Sperber et al. (2017, 2019b).
Recently, studies based on the end-to-end paradigm have emerged rapidly due to its advantages, such as lower latency, alleviation of error propagation, and fewer parameters. Zong et al. (1999) presume that it is possible to implement such end-to-end speech translation with the development of memory, computation speed, and representation methods. Then Bérard et al. (2016) give the first proof of the potential for an end-to-end ST model. Since then, pre-training, multitask learning, attention-passing, and knowledge distillation have been applied to improve the end-to-end model performance Anastasopoulos et al. (2016); Duong et al. (2016); Weiss et al. (2017); Bérard et al. (2018); Sperber et al. (2019a); Liu et al. (2019); Jia et al. (2019); Liu et al. (2020).
Considering the difficulty of modeling this cross-modal cross-lingual task in a single model, recent studies propose novel model structures or auxiliary tasks to enhance the model capability. Wang et al. (2020a) propose TCEN model which connects speech encoder and NMT encoder in tandem. This model aims to bridge the gap between pre-training and fine-tuning by reusing every sub-net. Considering the length inconsistency between speech encoder outputs and word embeddings, they lengthen the source text sentence by adding word repetitions and blank tokens to mimic the CTC output sequence. However, this process needs to train an extra sequence-to-sequence model and introduces much noise to the NMT model. They later propose a curriculum pre-training method which integrates two elementary courses to enable the encoder to understand the meaning of a sentence and map words in different languages Wang et al. (2020b). However, it conducts the force-alignment between the speech and the source word as well as the source-to-target word alignment which needs to train an extra ASR model and may introduce alignment errors. In total, these studies do not well solve the modality gap between speech and text. Consequently, much valuable semantic information learned by text-based MT model cannot be applied to the ST model, which limits the latter performance.
Cross-modal Adaptation. The paradigm of minimizing the difference between two models can be adopted in the transfer learning, where the knowledge embedded in one model can be transferred to another model, including output probabilities Hinton et al. (2015); Freitag et al. (2017), hidden representations Yim et al. (2017); Romero et al. (2015), and generated sequences Kim and Rush (2016). Such scheme has been applied in a variety of tasks, such as image classification Hinton et al. (2015); Li et al. (2017); Yang et al. (2018); Anil et al. (2018), speech recognition Hinton et al. (2015) and natural language processing Freitag et al. (2017); Kim and Rush (2016); Tan et al. (2019). Some related studies have attempted to transfer knowledge from the text model to the speech model by cross-modal adaptation. Cho et al. (2020) and Denisov and Vu (2020) adopts this method on spoken language understanding task, while Liu et al. (2019) apply it to speech translation task. However, they only transfer the knowledge from the output of the last model layer, ignoring the representation gap between speech and text modality.
Method
In this section, we first introduce the architecture of our proposed STAST model. Focusing on bridging the representation gap between speech and text modality, it decouples the encoder into three parts to transcribe the speech and extract semantic representation separately. To further close the semantic gap between two modalities, we apply a cross-modal adaptation method. Finally, we will give the training strategy for this task.
The speech translation corpus usually includes triplets of speech, transcription, and translation, denotes as , where is a sequence of speech features which are converted from the speech signals, is the corresponding transcription in source language, and denotes the translation in target language. , , are the length of speech features, transcription and translation, respectively, where . An extra ASR dataset can be leveraged for pre-training the ST model.
2 Model Architecture
STAST model adopts the encoder-decoder framework, as shown in Figure 1. Both the encoder and decoder adopt Transformer Vaswani et al. (2017) as the basic model structure since its superior property. Compared with the encoder in the ASR model which only needs to learn acoustic knowledge, the encoder in the end-to-end speech translation models is overloaded which requires to learn both acoustic and semantic knowledge of the source speech. To release its burden, we decouple the ST encoder into three parts, i.e. an acoustic encoder concatenated by a shrink mechanism, and a semantic encoder. Specifically, the acoustic encoder adopts a CTC module to learn speech representation and to predict the source transcription, which plays the role of an ASR model. To ensure the length of speech representation and text representations being consistent, a shrink mechanism is applied on the output of the acoustic encoder to filter redundant states based on the spike-like label posterior probabilities which are generated by the CTC module. This process can significantly reduce the length of speech input, as well as reserve most of the meaningful information in the source speech. Then the semantic encoder encodes the hidden states corresponding to the non-redundant positions to obtain better semantic representation, based on which decoder with an attention module generates the final translation. We introduce each part in detail as follows.
Acoustic Encoder. Since we decouple the ST encoder to executive different functions, the acoustic encoder here is mainly used to learn acoustic knowledge. It takes as input the sequence of speech features . For speech inputs, we first employ a speech pre-net to extract speech features, which is a linear layer here. The feature dimension is converted into the model hidden size . Then the acoustic representation is extracted by multiple stacked self-attention layers. The above process can be formalized as :
Shrink Mechanism. Note that the length of CTC output is the same as the input speech feature, which is still much larger than that of the corresponding source text. To bridge the length gap between speech and text representation, we apply a shrink mechanism.
As mentioned above, CTC paths are variation of the source transcription by allowing occurrences of blank tokens and repetitions. However, blank and repeated tokens do not contain any useful information and hinder the linguistics modeling. We assume that the triggered encode state sequence contains prior information of original speech input. Therefore, we only extract the encoded states in the acoustic encoder which corresponds to the CTC spike, as shown in the middle part of Figure 1. Specifically, we treat the label which has the largest value after the softmax layer in each state as the predicted label. Then, we only extract the hidden state whose label does not correspond to the blank label or consecutively repeated label and mask other states, which is similar with Yi et al. Yi et al. (2019) and Tian et al. Tian et al. (2020). Note that this process does not affect the back propagation of gradient. Then the shrunk hidden state can be formalized as follows, where denotes the corresponding index of token in the vocabulary.
Semantic Encoder and Decoder. The shrunk hidden state contains acoustic information but still lacks semantic knowledge. To obtain better semantic representation, we apply semantic encoder to encode the shrunk hidden states by another multiple stacked self-attention layers.
The decoder also follow the basic network structure of Transformer, it first adopts self-attention layers on target embeddings and then attends to the output of semantic encoder by cross-attention layers. Followed by feed-forward layer, the target token is predicted through a softmax layer based on the output of the decoder . The above process can be formalized as follows:
Finally, the distribution probability over a sequence of target tokens is calculated.
Integrated NMT Model. Since the NMT model has the same structure with the combination of semantic encoder and decoder, we integrate the whole NMT model into STAST. In NMT model, the source transcription is first embedded into word representation by looking up the embedding weight . Then semantic encoder extracts high-level semantic text representation , based on which the decoder performs translation task. To map speech representation and text representation into the same semantic space, we share the parameters of semantic encoder and decoder. Meanwhile, we also share the weight of softmax layer in acoustic encoder with the source word embedding weight and the weight of the decoder softmax layer to constrain the space gap, which means . Then the objective function of the MT task can be calculated by the cross-entropy loss as follow.
3 Cross-Modal Adaption
To further make the representation of speech and text modality closer, we propose a cross-modal adaptation method. Specifically, the cross-modal adaptation method is applied to transfer the semantic knowledge from text representation to speech representation. For each utterance, the semantic encoder encodes the shrunk output of acoustic encoder into semantic speech representation and encodes word embeddings into text representation , respectively. We take advantages of text representation as a regulation to constrain the space of speech representation during training. The supervision is implemented by minimizing the distance between two representations. We propose two adaptation methods, including sentence-level adaptation and word-level adaptation. The loss function can be calculated as:
where MSE is mean-squared error loss function used to evaluate the difference between the representation of speech and text, and are the average of two contextual representations.
4 Training Process
The final training objective function is the sum of four parts, including the CTC loss , the cross-entropy loss for ST task , the cross-entropy loss for MT task , and the cross-modal adaptation loss :
where , , , and are hyper-parameters, which denote the weight of each loss.
For training strategy, we first train the acoustic encoder by speech-transcription pairs in the ST corpus . Then we use speech-transcription-translation triplets to train the ST model and the NMT model by a multi-task learning framework, where transcriptions are only used during training. The module in our model is very flexible, where the CTC module and the integrated NMT model can perform auxiliary task to obtain better optimized parameters. Therefore, it can be easily trained by extra data, such as the part of acoustic encoder can be trained by extra ASR corpus to obtain better acoustic representation.
Experiments
We conduct experiments on two public speech translation datasets, including Augmented LibriSpeech English-French Corpus Kocabiyikoglu et al. (2018) and the MuST_C English-German TED Corpus Gangi et al. (2019b).
Augmented LibriSpeech En-Fr. This corpus is a subset of the LibriSpeech ASR corpus Panayotov et al. (2015), which is automatically aligned with e-books in French. This corpus is from clean audiobooks, which contains quadruplets, including English audios, manual transcriptions, French translations, and the translations obtained by Google Translate. The total audio contains 236 hours of speech. Following previous works Wang et al. (2020a), we only use the 100-hour clean training set and concatenate the aligned references with the provided Google translations, resulting in 90K utterances. We validate on the development set (1,071 utterances) and report the model performance on the test set (2,048 utterances). For pre-training, we use the total LibriSpeech ASR corpus as extra data, which includes 960 hours of speech.
MuST_C En-De. The MuST_C corpus is collected from TED talkshttps://www.ted.com, which includes the English speech, the corresponding transcription, and the target translations in different languages. We conduct experiments on English-German language direction. The speech in this corpus is recorded from the live presentation which contains more noise. The corpus contains a total of 408-hour speech with 234K translation pairs, which is divided into training set (400 hours with 229,703 utterances), development set, and test set. We report case-sensitive BLEU on the dev set (1,423 utterances) and tst-COMMON set (2,641 utterances).
2 Experimental Settings
The speech features are 80-dimensional log-Mel filterbanks extracted with a step size of 10ms and window size of 25ms, which are extended with mean subtraction and variance normalization. We adopt dimensionality reduction to downsample one frame every three frames.
For text data, we apply lowercase, punctuation normalization, and tokenization by Moses scriptshttps://www.statmt.org/moses/. Punctuations in English transcriptions are removed. We apply the BPE method Sennrich et al. (2016) on the combination of source and target text to obtain shared subword vocabulary. The number of merge operations in BPE is set to 8K for both tasks. In order to be comparable with other works, we employ case-insensitive BLEU computed using multi-bleu.pl scripthttps://github.com/moses-smt/mosesdecoder/scripts/generic/multi-bleu.perl as the evaluation of the translation task.
We use the base configuration of original Transformer Vaswani et al. (2017), where the number of self-attention layers in acoustic encoder, semantic encoder, and decoder is 6 with 512-dimensional hidden sizesWe compare to the ST model with a 12-layer single encoder, whose BLEU is 17.21 on the test set of Augmented LibriSpeech corpus., the filter size in feed-forward layer , the residual dropout and attention dropout are . We set , , , and in Equation 14 to 1.0, 1.0, 1.0, and 1.0 respectively. Samples are batched by approximate sequence length of 10,000-frame features. The STAST model is trained by Adam optimizer Kingma and Ba (2015) on one GPU. We save checkpoints every 1,000 steps and conduct the model average on the last 5 checkpoints as the final model. For inference, we perform beam search with a beam size of 4. Our code will be released after reviews.
Experimental Results
Results on Augmented LibriSpeech. Following previous work Wang et al. (2020b), we have two settings in this experiment. In the Base setting, we only use the data in the Augmented LibriSpeech corpus. Regarding the Expended setting, we use the 960-hour LibriSpeech ASR corpus as extra data to pre-train the acoustic encoder. Table 1 presents the results of MT models, pipeline systems, and end-to-end ST models in both settings.
Comparison with pipeline ST systems. We compare our end-to-end STAST model with text-based MT models and pipeline ST systems. The text-based MT model takes manual transcriptions as input, so its result can be regarded as the upper bound of the speech translation task. For the pipeline system, the ASR model and the NMT model are trained only by data in the ST corpus. The final translation is generated based on the output of the ASR model. As shown in Table 1, in the base setting our STAST model can achieve comparable or even better results than pipeline systems under the same data scale. However, STAST only uses a single model to combine ASR and MT tasks, which has much fewer parameters and computational cost than pipeline systems. When more ASR data is available, our method outperforms the best pipeline system by 0.9 BLEU and is slightly worse than text-based MT models. It indicates that with the use of extra ASR corpus, our method can obtain more acoustic information, which has benefits of extracting semantic knowledge and achieving better performances.
Comparison with end-to-end models. As shown in Table 1, our model outperforms all the previous end-to-end models and achieves the new state-of-the-art performance. Specifically, in the base setting our proposed model is much better than the LSMT-based ST model Bérard et al. (2018) and outperforms it by BLEU scores. Our model is also better than a widely used toolkit ESPnet by 1.1 BLEU, whose encoder and decoder are both pre-trained. Compared with Liu et al. Liu et al. (2019) who utilize an NMT model to teach an ST model on the probability of decoder output, our method achieves 0.8 better BLEU scores. We believe this is because their method only adopts knowledge distillation on the last output layer of model, which is hard to propagate the gradient back to the preceding networks. However, our method adopts cross-modal adaptation on speech and text representation, giving the model more guidance to bridge the gap between two modalities. STAST is also better than TCEN-LSTM proposed by Wang et al. Wang et al. (2020a). Compared with their method which requires to transform normal source sentences to noisy sentences in CTC path format, our method solves length inconsistency problem more effectively. Besides, the speech representation can obtain better semantic information with the cross-modal adaptation. Our method is also superior to the best previous model proposed by Wang et al. Wang et al. (2020b). In the expanded setting, with extra ASR data, our method improves the translation performance by 0.9 BLEU compared with that in the base setting. It also outperforms other previous end-to-end models by a large margin, which indicates the effectiveness of our method.
Results on MuST_C English-German. The results on MuST_C English-German corpus are listed in Table 2. Compared with the Augmented LibriSpeech corpus, the performance gap between text-based MT models and pipeline systems is much larger. This can be attributed to that the speech in this corpus is recorded from live presentations, which contains more noise, such as laughter, applause, and cheers, besides that the segment of this corpus is also noisy. As seen in Table 2, STAST model outperforms all of the previous end-to-end models in the base setting. We also find that the STAST model is slightly better than the pipeline system we implement but worse than the best system, which indicates the end-to-end ST model still lacks robustness to the noisy input. It is challenging to handle the speech under noisy circumstances and needs to design more powerful ST models.
2 Ablation Studies
To better evaluate the contribution of different proposed methods, we perform ablation studies on both corpora. The results in Table 3 show that all of the proposed methods have positive effects, and the benefits of these methods are accumulative. The performance drops if we remove the cross-modal adaption method. Without multi-task learning, the performance further drops 0.2 and 0.6 BLEU scores. It indicates that both cross-modal adaption and multi-task learning have the ability to transform the latent space of speech representation to text representation closely. With the guidance from the text-based NMT model, the ST model can learn more semantic knowledge and obtain better performance. We find if the shrunk output from the acoustic encoder is directly fed into decoder without the semantic encoder, the translation performance drops significantly. This means decoupling the ST encoder is necessary and the semantic encoder supplements the model with more capability to learn semantic information. If we further remove the shrink mechanism, the performance decreases by 0.4 and 0.8 BLEU scores on two corpora, which proves that blank and repeated tokens hinder the alignment learning between output sequence and input sequence. The shrink mechanism has the ability to bridge the length gap between speech and text representation and is indispensable for cross-modal adaptation. The discussion in detail will be depicted in Section 5.3. Finally, the CTC loss as an auxiliary loss is also beneficial, which is consistent with previous work in the ASR fields that CTC loss has the benefit of accelerating convergence and achieving better performance Kim et al. (2017).
3 Analyses
Effect of Shrink Mechanism on the CTC output. As shown in Figure 2, the histogram counts the difference between the length of the CTC output after shrink mechanism and that of the corresponding transcription on the training set of the Augmented LibriSpeech corpus. When the value is below zero, it means the length of shrunk hidden states is longer than that of transcription and vice versa. We find that the length of shrunk hidden states equals to that of the corresponding text transcription for 84.0% of data, and the difference is less than two for over 93.7% of data. Therefore, we conclude the shrink mechanism has the ability to solve length inconsistency problem. With more extra ASR data, we believe the accuracy of the shrunk length can be further improved.
Sequence-Level Adaptation v.s. Word-Level Adaptation. We compare two different adaptation methods as mentioned in Section 3.3, i.e. sequence-level adaptation and word-level adaptation. We conduct experiments on the Augmented LibriSpeech corpus and the results are reported in Table 4. We find that both adaptation methods have improvements compared with baselines. However, the performance of sequence-level adaptation is slightly better. This is consistent with Aldarmaki and Diab Aldarmaki and Diab (2019), where they find learning the aggregate mapping can yield a more optimal solution compared to word-level mapping.
The module in our proposed STAST model is flexible, which can easily leverage extra data. In Section 5.1, we have shown that extra ASR data can boost the model performance, here we analyze the effect of extra text data. To simulate lower-resource scenarios, we randomly select 10-hour, 50-hour, and 100-hour of speech training data (speech-transcription-translation triplets) from the MuST_C En-De dataset, which originally contains 400-hour speech. Then, we train the end-to-end ST baseline model (a simple encoder-decoder model), multi-task baseline model (the decoder of ST and MT is shared with separated encoders), and our STAST model under different scales of training data. The multi-task baseline model and our STAST model can have access to extra text data (transcription-translation pairs in the original corpus). Figure 3 shows the BLEU scores of three models under different scales of training data. We find that with the increase of the data size, all of the three models can obtain improvements. However, the end-to-end model is significantly inferior to the multi-task model and the STAST model. With only 10 hours of training data, our proposed model can achieve comparable performance with the end-to-end model trained on 100-hour data. The multi-task model is better than the end-to-end model but worse than our proposed model. The reason is that the multi-task model only shares part of parameters for different tasks but leaves the valuable semantic information learned by the NMT encoder unexploited by the ST model. While our method integrates the NMT model into the ST model and transfers speech representation to text representation more closely by the cross-modal adaptation method, which can obtain better performance.
Conclusions
In this paper, we propose the STAST model to improve end-to-end ST model. Considering the modality differences between speech and text, we propose ST encoder decoupling, length shrink mechanism, the NMT model integration, and cross-modal adaptation methods. Empirical studies have demonstrated that each proposed method has positive effects and the combination of them can achieve the new state-of-the-art result. In the future, we will explore how to effectively transfer more knowledge from the NMT model and the ST model. We also expect that the idea of bridging the representation gap between different modalities can be adopted on other tasks.