Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation
Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, Ann Lee
Introduction
Direct speech-to-speech translation (S2ST) aims at translating speech from one language into speech in another language without relying on text generation as an intermediate step . Compared to conventional cascaded approaches , which take advantage of automatic speech recognition (ASR), machine translation (MT) or end-to-end speech-to-text translation (S2T) followed by text-to-speech synthesis (TTS), direct S2ST has the advantage of faster inference and can support translation between languages without text writing systems .
Most recently, proposes to apply a self-supervised speech encoder pre-trained on unlabeled speech to convert target speech into discrete units and build a speech-to-unit translation (S2UT) model for direct S2ST. Self-supervised discrete targets can disentangle linguistic content from speaker identity and prosodic information in speech . Moreover, they enable opportunities for applying techniques from speech-to-text model training, such as ASR and S2T, to direct S2ST.
As we move along the spectrum from multi-stage to direct approaches, the amount of parallel data available becomes much more scarce. Pre-training, including initializing encoder or decoder trained from ASR or MT tasks, or from self-supervised pre-training with unlabeled data , multitask learning and data augmentation have been extensively studied for tackling the data scarcity issue for S2T. For direct S2ST, show that multitask learning is crucial for model convergence, and focus on incorporating pre-trained modules from ASR, S2T, MT or TTS tasks.
In this work, we take advantage of the S2UT framework proposed in and show that large-scale self-supervised pre-training with monolingual speech and text data and data augmentation techniques that benefit S2T model training can also be applied on S2UT model training. For pre-training, we transfer the technique from , which performs efficient finetuning with a wav2vec 2.0 speech encoder and an mBART text decoder , to S2UT with a wav2vec 2.0 speech encoder and an mBART decoder trained with discrete units extracted from unlabeled speech data. For data augmentation, we utilize ASR, MT and TTS models to create weakly supervised data .
The contributions of this work are as follows: we empirically demonstrate that self-supervised encoder and decoder pre-training with unlabeled speech and partial finetuning improve S2ST training under various setups, including training with synthetic single-speaker or real multi-speaker target speech, and low-resource setup (30-hr). We further improve the model via data augmentation with weakly supervised data and conduct experiments with the combination of multiple datasets to achieve strong performance over ASR+MT+TTS baseline systems.
Related Work
Self-supervised speech encoder pre-training has led to huge performance improvement on a wide range of applications such as ASR , S2T , speaker identification , etc. SpeechT5 proposes end-to-end pre-training, where the model learns to reconstruct the log Mel-filterbank of the masked regions of the input speech, to improve performance of speech-to-speech tasks such as voice conversion and speech enhancement, while it was not evaluated in the context of direct S2ST.
The improved quality of speech units discretized based on self-supervised speech representations also allows researchers to apply natural language processing (NLP) techniques on speech, such as spoken generative language modeling , and emotion conversion cast as a unit-to-unit translation task . further applies text-based autoencoder denoising on units for model pre-training. In this work, we model the target speech as discrete units, extend the monolingual unit pre-training in to a multilingual setup, and perform speech encoder and discrete unit decoder pre-training separately.
Data augmentation is a common method for increasing the size of training data. Synthetic speech data from TTS , self-training or back-translation from unlabeled data have been shown to be useful to ASR or S2T. In this work, we apply MT and TTS to prepare weakly supervised S2ST data from speech in the source language.
System
We follow and encode the target speech as discrete units with a HuBERT model trained on unlabeled speech followed by a k-means model learned from its hidden representations . Waveform input is encoded into a sequence of k-means cluster indices at every 20-ms frame, and we remove consecutive duplicate units to create a reduced unit sequence representing the target speech. In the end, the direct S2ST system consists of a sequence-to-sequence S2UT model with a speech encoder and a unit decoder, followed by a unit HiFi-GAN vocoder trained separately for unit-to-waveform conversion. In this work, we explore both encoder and decoder pre-training (Fig. 1).
2 Model pre-training
Wav2vec 2.0 is a self-supervised framework to learn speech representations from unlabeled audio data. It uses a multi-layer convolution neural network to encode the audio followed by a Transformer-based context encoder to build the contextualized representations. The model is trained via contrastive loss with masked spans on the input to the context encoder. In this work, we use Conformer instead of the Transformer for better model performance (Sec. 4.5.2).
2.2 Decoder pre-training: unit mBART
mBART was originally proposed for denoising autoencoder over text. During training, the sequence-to-sequence model predicts the original text given its noisy version, , created by randomly masking spans of . The starting position of each span is uniformly sampled from all positions, and the span lengths are sampled from a Poisson distribution (). The masking process is repeated until the total accumulated span lengths take up of the input. The model is trained on text data from multiple languages with language tags. In our case, we treat the reduced discrete units extracted from unlabeled speech data as text and apply mBART training with a Transformer-based encoder-decoder architecture .
3 Model finetuning
We combine the wav2vec 2.0 encoder and the unit mBART decoder and study the finetuning strategies in . A randomly initialized adaptor layer consisting of a single 1-D convolutional layer with stride 2 is added between the pre-trained modules to increase the model’s capacity to alleviate the mismatch between the learned representations, as well as length difference between the source audio and the reduced target units.
We examine both full and partial finetuning of the model. For the latter, we focus on the LayerNorm and Attention modules (dubbed as “LNA”) proposed in . The hypothesis is that LayerNorm parameters reflect the statistics of the pre-training data, and the encoder attention in the unit mBART decoder is optimized for unit sequence input. The adaptor layer is fully finetuned. We explore four finetuning strategies in total: 1. LNA-E: The LayerNorm and self attention parameters in the encoder and all the parameters in the decoder are finetuned. 2. LNA-D: The whole encoder and the LayerNorm and both encoder and self attention in the decoder are finetuned. We optionally freeze the encoder for the first updates. 3. LNA-E,D: Only LNA parameters are finetuned both on the encoder and the decoder side. 4. Full: We finetune the whole model end-to-end with an option of freezing the encoder for the first updates.
4 Data augmentation
We take advantage of speech from ASR data in the source language to increase the size of the parallel S2ST training data. We use a Transformer MT model to translate the text transcription in the source language to text in the target language. To convert text in the target language into target speech, we apply a text-to-unit (T2U) model https://github.com/pytorch/fairseq/blob/main/examples/speech_to_speech/docs/textless_s2st_real_data.md, which is a Transformer-based sequence-to-sequence model trained on text and the corresponding discrete unit sequence extracted from the paired audio. The T2U model is a way to bypass the TTS generation and HuBERT unit extraction pipeline for efficient generation of large-scale weakly supervised data. We choose to distill knowledge from MT models instead of pursuing self-training, since a three-stage cascaded system (ASR+MT+TTS) can take advantage of a large amount of data from each component during training and still outperforms existing direct S2ST systems .
Experiments
We conduct experiments on Spanish-English (Es-En) and English-Spanish (En-Es) translation with fairseq .
Table 1 summarizes the statistics of all the datasets used in the experiments. We experiment with two types of parallel S2ST data. First, we follow the convention of applying single-speaker TTS En: https://huggingface.co/facebook/tts_transformer-en-ljspeech, Es: https://huggingface.co/facebook/tts_transformer-es-css10 on the target text of S2T data (dubbed as “S2ST-syn”). We combine S2T datasets from multiple domains to improve the robustness of model training , resulting in 196-hr training data for Es-En and 571-hr for En-Es. We also combine the dev sets from all domains for tracking the training process and checkpoint selection, and conduct evaluation on all test sets. The second S2ST dataset is from VoxPopuli , which contains speech from European parliament plenary sessions and the oral interpretation (dubbed as “S2ST-real”). We use the same training data in and apply the speech normalizer from the 1-hr setup on the target speech.
The wav2vec 2.0 speech encoder and unit mBART are pre-trained on unlabeled speech data. For VoxPopuli , we remove utterances from the year 2012 and before to avoid overlap with the Europarl-ST dev and test data. For data augmentation, we use all the ASR data and the text transcriptions of the source speech in the S2T datasets to train the ASR models, and all the parallel text data to train the MT models.
2 Model setup
We use the multilingual HuBERT (mHuBERT) model, k-means model and unit-based HiFi-GAN vocoder from footnote 2, to encode target speech into a vocabulary of 1000 units. The mHuBERT and the k-means models are learned from the combination of En, Es and French unlabeled speech data from VoxPopuli , while we use them to encode En and Es target speech only.
We train the Conformer wav2vec 2.0 speech encoder with the LARGE configuration using Libri-light for En and VoxPopuli for Es, respectively, for 200k updates with a batch size of 19.4-hr for Es and 14.7-hr for En. We train the unit mBART with the LARGE configuration using the combination of all En and Es unlabeled speech for 500k updates with and , and we do not use the sentence permutation noise. During finetuning, we tune the hyper-parameters including learning rate ([5e-5, 1e-4]), dropout ([0.1, 0.3]), label smoothing ([0, 0.3]) and also encoder specific ones namely mask channel length (), mask probability ([0.1, 0.5]), channel mask probability ([0.1, 0.5]), layer drop ([0, 0.3]) and number of updates to freeze the wav2vec 2.0 encoder ([0, 5k]) on the dev sets.
3 Baselines
We build two cascaded baselines, ASR+MT+TTS and S2T+TTSfootnote 3, and two supervised S2UT baselines: 1. ASR: The En ASR model is finetuned with CTC from the Conformer wav2vec 2.0 model. We apply the same training for Es but find that a supervised ASR model with the s2t_transformer_l architecture in fairseq is better. 2. MT: As the ASR models are trained with normalized text (e.g. lowercase, digits in spoken form, etc.), we apply text normalization on both source and target texts as well to train Es-En and En-Es MT models. We use the transformer_wmt_en_de_big architecture in fairseq. The MT models are also used in data augmentation. 3. S2T: The S2T model consists of the pre-trained Conformer wav2vec 2.0 encoder and a randomly initialized text decoder with 6 Transformer layers, 8 attention heads, 256 embedding size and 2048 FFN embedding size, and is trained on the S2T datasets without multitask learning. 4. Supervised S2UT: We follow the same model configuration in to train Transformer-based S2UT models and explore both without and with multitask learning. For the latter we include two auxiliary tasks that use character sequences from source and target text transcripts as targets.
4 Evaluation
To evaluate the translation quality, we use open-sourced ASR modelsEn: https://huggingface.co/facebook/wav2vec2-large-960h-lv60-self, Es: https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53-spanish to transcribe the audios and compute BLEU scores using SacreBLEU . The reference text is normalized to lowercase, punctuation is removed, digits are converted to spoken forms, and all words in parentheses like “(Applause)” or “(Music)” are removed. We do not consider samples with empty translation after text normalization. To evaluate the naturalness of the speech output, we collect mean opinion scores (MOS) on a scale of 1 (the worst) to 5 (the best) from human listening tests on a set of 200 utterances randomly sampled across all the test sets for each system, and each sample is rated by 7 raters.
5 Results
Table 2 shows results from models trained with “S2ST-syn” data. We also provide BLEU from the synthetic targets (11) to demonstrate the impact of ASR errors. First, we see that without multitask, the supervised Es-En S2UT model (3) cannot converge properly with the combined 196-hr training set, while multitask learning helps model training (4).
Next, we see that with a pre-trained wav2vec 2.0 encoder and a randomly initialized decoder, we can achieve an average of 5.6 BLEU gain on En-Es test sets and 4.0 BLEU for Es-En (4 vs. 5). As we incorporate a unit mBART decoder, we find that LNA-D is the most effective finetuning strategy, yielding an average of 6.6 BLEU gain on En-Es test sets and 8.1 BLEU on Es-En compared to multitask learning (4 vs. 7). Our best En-Es S2ST model performs on par with the S2T+TTS baseline, and the Es-En S2ST model outperforms the cascaded system by 2.8 BLEU (1 vs. 7). Though the S2T system can be further improved with text-based pre-training, this is beyond our scope.
Further incorporating weakly supervised training data from ASR speech can bring +0.7 BLEU on En-Es and +3.1 BLEU on Es-En (7 vs. 10). We can also compare this system with the three-stage cascaded system, as they incorporate information from the same amount of supervised ASR and MT data. We see that the En-Es direct system can outperform ASR+MT+TTS, while there is a 1.4 BLEU gap for Es-En systems (2 vs. 10).
For MOS, Table 2 shows that the En-Es S2UT system produces more natural speech than Es TTS does (1 vs. 7), while the quality of the Es-En S2UT output is much worse. Note that the naturalness of the output speech is mainly controlled by the unit vocoder, and we use the models from without finetuning.
One advantage of self-supervised pre-training is that it only uses speech data and can work for unwritten languages. We finetune on “S2ST-real” data and compares with that incorporates an auto-encoding auxiliary task to predict the discrete units extracted from the source speech. We use dev and test sets from Europarl-ST, since it’s in the same domain as the “S2ST-real” data. Table 3 shows a 4.2 and 5 BLEU gain for En-Es and Es-En with LNA-D.
Finally, we study the effect of pre-training, with a decreasing amount of parallel data for finetuning by randomly sampling a subset from the “S2ST-syn” training data. Table 4 compares LNA-D finetuning with supervised S2UT models trained with multitask learning. When training on less than 100 hours of data, we reduce the supervised S2UT model to 8 encoder layers and 4 decoder attention heads. We see an average 3.4-13.0 BLEU gain from pre-training and LNA-D finetuning for both language directions when training with more than 30 hours of data. Pre-training and finetuning with 50-hr data can already outperform supervised systems trained with 100-hr data. However, when the amount of parallel data is decreased to 10 hrs, both multitask learning and pre-training cannot work.
5.2 Model variations
First, we examine Conformer vs. Transformer wav2vec 2.0. We perform finetuning with the pre-trained encoder and a randomly initialized unit decoder using Es-En “S2ST-syn” data and see that Conformer wav2vec 2.0-LARGE model gives an average 4.6 BLEU gain compared with Transformer LARGE model.
Next, we study how and affect unit mBART by training the model for 300k updates and finetuning on a unit-to-unit translation task, where both source and target speech are converted to reduced discrete unit sequences. From Table 5, we do not see large difference except when and .
Conclusions
In this work, we study self-supervised pre-training and data augmentation for direct S2ST models. We take advantage of an S2UT framework that encodes target speech into discrete representations, apply wav2vec 2.0 speech encoder and unit mBART decoder pre-training and perform partial finetuning. Experiments under various setups including synthetic and real target speech and low-resource all verify the effectiveness of the approach. We also show that applying MT to create weakly supervised data from speech in the source language can be further combined with pre-training to improve model performance.
Acknowledgements
We would like to thank Justine Kao and Brian Bui for the help on MOS evaluation.
References
Appendix A Model Training Hyper-parameters
To train the wav2vec 2.0 models, we use the wav2vec2_conformer_large_librivox configuration defined in fairseq. The model contains 24 Conformer blocks with model dimension 1024, inner dimension 4096, 16 attention heads and a total of 620.5M parameters. We use a total batch size of 14.7-hr for En and 19.4-hr for Es, and dropout 0.1. We use Adam with , weight decay of 0.1 and learning rate 0.005, and apply polynomial decay learning rate schedule with 32k warmup steps. We started the training with fp16 initially but switched to fp32 when we encountered NaNs in the back propagation. Rest of the hyper-parameters are listed in Table 6.
A.2 Unit mBART
We use the mbart_large architecture defined in fairseq to train the unit mBART. The model contains 12 Transformer layers in both encoder and decoder, embedding size 1024, feed-forward network (FFN) dimension 4096, 16 attention heads and a total of 353M parameters. Different from original text mBART, we do not apply sentence permutation noise in the denoising task, and we only apply masking noise. We use Adam with and learning rate 0.0003, and apply polynomial decay learning rate schedule with 10000 warmup steps. The model is trained with dropout 0.1, mask probability 0.3, for 500k steps. We use 1024 max tokens with 64 GPUs and update frequency 9.
A.3 Supervised S2UT baselines
To train the supervised S2UT models with multitasks, we use the s2ut_transformer_fisher configuration defined in fairseq. We use Adam with , weight decay of 0.1 and learning rate 0.0005, and apply inverse square root learning rate schedule with 10k warmup steps. For En-Es, we use a max tokens value of 3000, dropout value of 0.1 and 16 GPUs. For Es-En, we use 4500 max tokens value, dropout 0.3 and 8 GPUs.
For auxiliary tasks, we have a Transformer decoder on the sixth layer of the encoder for source character prediction and a Transformer decoder on the eighth layer of the encoder for target character prediction. Each of the multitasks has a loss weight of 8. The decoders have 2 Transformer layers with 256 embedding size and FFN dimension of 2048.
A.4 Finetuning
We use a dropout value of 0.1 and label smoothing value of 0.2 for both language directions. We use encoder layerdrop value of 0.1 for En-Es and 0.2 for Es-En. For the decoder, we use the same embedding parameters for the input and output. We use Adam with and apply inverse square root learning rate schedule with 10k warmup steps. We finetune the models for 15k updates with “S2ST-syn” data and 25k updates when incorporating the weakly supervised data. We use max tokens of 3500 for En-Es and 4000 for Es-En, set the update frequency to 15 and train on 16 GPUs. The learning rate and the number of updates we freeze the wav2vec 2.0 encoder in the beginning for each of the setups is listed in the Table 7.
A.5 Low-resource setup
Table 8 lists the statistics of the training data used in the low-resource setup. We use the same set of hyper-parameters for the wav2vec 2.0 and mBART finetuning on the low-resource setup as described in Sec. A.4 except for the ones listed in Table 9. Finally, Table 10 lists the BLEU scores on each dataset under the low-resource setup, whose average values are presented in Table 4.