CVSS Corpus and Massively Multilingual Speech-to-Speech Translation

Ye Jia, Michelle Tadmor Ramanovich, Quan Wang, Heiga Zen

Introduction

Speech-to-speech translation (S2ST) is an important means for breaking down the communication barriers between people speaking different languages. Conventionally, S2ST systems are built with a cascade of automatic speech recognition (ASR), text-to-text machine translation (MT), and text-to-speech (TTS) synthesis sub-systems, which are text-centric. Recently, work on S2ST without relying on intermediate text representation are emerging, such as end-to-end direct S2ST (Jia et al., 2019b; Kano et al., 2021; Jia et al., 2022) and cascade S2ST based on discrete speech representation (Tjandra et al., 2019; Zhang et al., 2021; Lee et al., 2022; Ma et al., 2021; Lee et al., 2021). However, as of today, publicly available corpora directly suitable for such research are extremely limited (see Table 1).

In this paper, we introduce CVSS, a Common Voice-based Speech-to-Speech translation corpus. CVSS is directly derived from the CoVoST 2 ST corpus, which is further derived from the Common Voice speech corpus. CVSS provides sentence-level parallel speech-to-speech translation pairs from 21 languages into English, namely from Arabic (ar), Catalan (ca), Welsh (cy), German (de), Estonian (et), Spanish (es), Persian (fa), French (fr), Indonesian (id), Italian (it), Japanese (ja), Latvian (lv), Mongolian (mn), Dutch (nl), Portuguese (pt), Russian (ru), Slovenian (sl), Swedish (sv), Tamil (ta), Turkish (tr), and Chinese (zh). The source speech in these 21 languages is crowd-sourced human volunteer recordings from the Common Voice project, totalling 1153 hours. Two versions of translation speech in English are provided for all the source speech, both are synthesized using state-of-the-art TTS systems, with each version providing unique values not existing in other public S2ST corpora:

CVSS-C: All the translation speech is in a single canonical speaker’s voice, totalling 719 hours. Despite being synthetic, the speech is highly natural, clean, and consistent in speaking style. These properties ease the modeling of the target speech and enable trained models to produce high quality translation speech suitable for general user-facing applications.

CVSS-T: The translation speech is in voices transferred from the corresponding source speech, totalling 784 hours. Each S2ST pair has a similar voice on the two sides despite being in different languages, making this dataset suitable for building models where voice preservation during speech translation (Jia et al., 2022) is desired.

Together with the source speech, the two S2ST datasets contain 1,872 and 1,937 hours of speech, respectively. In addition to translation speech, CVSS also provides normalized translation text matching the pronunciation in the translation speech (e.g. on numbers, currencies, acronyms, etc.), which can benefit both model training as well as evaluation.

Unlike existing corpora of simultaneous interpretation, e.g. VoxPopuli (Wang et al., 2021a) and STC (Shimizu et al., 2014), the target speech in CVSS is translation instead of interpretation. As a comparison, translation is typically verbatim and exact, while interpretation is typically summarizing and often drops less important details; there is also more linguistic variation and disfluencies in interpretation (He et al., 2016; Shimizu et al., 2013; Wang et al., 2021a).

CVSS is released under the very permissive Creative Commons Attribution 4.0 International (CC BY 4.0) license. It can be freely downloaded online.https://github.com/google-research-datasets/cvss

On each version of CVSS, we built two baseline direct S2ST models (Translatotron (Jia et al., 2019b) and Translatotron 2 (Jia et al., 2022)) and a baseline cascade S2ST model (ST →\to TTS). To build strong cascade S2ST baselines, we trained an ST model on CoVoST 2, which outperforms the previous state-of-the-art trained on the corpus without using extra data by ++5.8 average BLEU on all 21 language pairs, or ++6.9 average BLEU on the 4 high-resource language pairs. Nevertheless, the performance of the Translatotron 2 direct S2ST model approaches the strong cascade baseline when trained from scratch, and with only 0.1 or 0.7 BLEU difference on ASR transcribed translation when initialized from matching ST models. These results verified the effectiveness of both the CVSS corpus as well as the approach of direct S2ST. We hope the release of the CVSS corpus and the baselines we provide can help accelerate the research on direct S2ST.

Related works

Research on S2ST has progressed for over three decades since early efforts such as Waibel et al. (1991). However, publicly available corpora with parallel S2ST pairs are still extremely limited as of today. This is largely because until very recently, S2ST research has focused on the cascade approach, thus requiring separate ASR, MT, and TTS corpora. However, such corpora are not directly usable for building S2ST without relying on text representation.

Fisher Spanish-English ST corpus (Post et al., 2013) is the most widely used public corpus in recent S2ST works (Jia et al., 2019b; Zhang et al., 2021; Lee et al., 2022; Ma et al., 2021). It contains 127 hours of Spanish telephone conversations and corresponding English translation text. However, this corpus does not include translation speech, and all these works used their own versions of synthetic translation speech. The low sample rate (8 kHz) of the source speech also makes it less ideal for modern S2ST research.

VoxPopuli (Wang et al., 2021a) is a recent large speech corpus originated from European Parliament event recordings. It includes 17.3k hours simultaneous oral interpretation in 15 languages, which is by far the largest S2ST corpus publicly available. Because of the nature of oral interpretation, important information in the source speech can be missing in the interpretation. The variation in speakers’ voices, recording conditions, and disfluencies in the interpretation pose additional challenges for S2ST modeling on this corpus.

MaSS (Boito et al., 2020) is a small corpus of Bible reading in 8 languages, with about 20 hours of speech per language. STC (Shimizu et al., 2014) includes a small publicly available simultaneous interpretation corpus that interprets English TED Talks recordings into Japanese.The STC corpus includes en ↔\leftrightarrow ja simultaneous interpretation from multiple sources, but only a portion of the en →\to ja direction has both the source and target speech available.

A few recent works (Tjandra et al., 2019; Kano et al., 2021) used the BTEC corpus (Kikui et al., 2003, 2006), which is derived from a small hand-crafted MT corpus of phrases in the travel domain. This corpus is currently not available to be downloaded. Similarly, a few other corpora with S2ST pairs are no longer publicly available, such as: EPIC (Bendazzoli et al., 2005), containing 18 hours simultaneous interpretation among Italian, English and Spanish, originated from the European Parliament speech; CIAIR (Tohyama et al., 2004), containing 182 hours simultaneous interpretation between Japanese and English.

Most of the above mentioned recent S2ST works use synthetic translation speech as training targets (Jia et al., 2019b; Tjandra et al., 2019; Zhang et al., 2021; Kano et al., 2021; Lee et al., 2022; Jia et al., 2022; Ma et al., 2021). There are two primary motivations for doing so: 1) Collecting a large amount of synthetic speech is of much lower cost than collecting human recordings, in absence of a directly usable S2ST corpus; 2) Synthetic speech can be easier to model because of consistent voice, speaking style, and high cleanness. Jia et al. (2019b, 2022) showed that despite being training on synthetic speech, the trained S2ST models can produce translation speech in high naturalness.

A few works built S2ST models with real-world human recordings as training targets. Because large-scale human recordings usually have to be collected with multiple speakers in different recording conditions, these works have to introduce additional components for tackling the variation in speakers’ voices, speaking styles, and recording conditions, etc. Such components are often trained with additional corpora. Jia et al. (2019b) used a speaker encoder separately trained with a speaker verification corpus, in order to model such variation. Lee et al. (2021) used a speech normalizer separately trained with a synthetic speech normalization corpus and a speech quantizer separately trained with an unsupervised speech corpus, in order to eliminate such variation.

Besides collecting human speech or using synthetic speech to construct S2ST datasets, it is possible to mine S2ST data from existing multilingual untranscribed speech corpora. Duquenne et al. (2021) showed a proof-of-concept of such an approach.

Text normalization

The translation quality of S2ST is typically evaluated by measuring BLEU (Papineni et al., 2002) between reference translation text and ASR transcription of the translation speech (Jia et al., 2019b). Because ASR usually outputs with minimal punctuation and case supportParticularly, most of recent S2ST works used ASR models trained on the LibriSpeech corpus (Panayotov et al., 2015) for such evaluation. LibriSpeech corpus provides text in uppercase without punctuation marks., such evaluation typically computes BLEU case-insensitively and ignores punctuation marks. In addition, a few works, e.g. Lee et al. (2021), further normalize special tokens such as numbers in reference text before computing BLEU. Such text normalization is not standardized, which makes the result comparison among different works difficult. In CVSS, we provide normalized translation text matching the pronunciation in the target speech, which can be used for model training as well as help standardize the evaluation on this corpus.

Source corpora

CVSS is directly derived from the CoVoST 2 ST corpus, which is further derived from the Common Voice speech corpus.

(Ardila et al., 2020) is a massively multilingual transcribed speech corpus designed for ASR. The speech in the corpus is crowdsourcing collected by volunteer contributors reading text content from Wikipedia and other text corpora. The size and the language coverage of the corpus keeps growing. The current release (version 7) consists of 11,192 hours of validated speech in 76 languages.

CoVoST 2

(Wang et al., 2021b) is a large-scale multilingual ST corpus derived from Common Voice. It covers translation from 21 languages into English and from English into 15 languages. The source speech is directly from Common Voice version 4. The translation was collected from professional translators on the scripts from the Common Voice. The 21 X-En language pairs consist of 1,154 hours of speech in total.

TTS models

CVSS is constructed by synthesizing the translation text from CoVoST 2 into speech using two state-of-the-art TTS models. This section describes the two TTS models being used, both of which were trained on the LibriTTS corpus (Zen et al., 2019).

PnG NAT (Figure 1) is a combination of PnG BERT (Jia et al., 2021) and Non-Attentive Tacotron (NAT) (Shen et al., 2020). It synthesizes speech as natural as professional human speakers (Jia et al., 2021).

PnG BERT is an encoder model specifically designed for neural TTS. It takes both phoneme and grapheme representations of text as input, as well as the word-level alignment between them. Similar to BERT (Devlin et al., 2019), PnG BERT can be pre-trained on a large text corpus in a self-supervised manner. Experimental results show that PnG NAT using a pre-trained PnG BERT yields more natural prosody and more accurate pronunciation than a baseline NAT model using only phoneme input with no pre-training. Subjective side-by-side (SxS) preference evaluations show that raters have no statistically significant preference between the speech synthesized using PnG NAT and ground truth studio recordings from professional speakers (Jia et al., 2021).

We pre-trained PnG BERT on a plain text corpus mined from Wikipedia, containing 131M English sentences, and fine-tuned it in PnG NAT on the entire LibriTTS corpus. We followed the hyperparameters in Jia et al. (2021); Shen et al. (2020).

The performance of the trained PnG NAT model is evaluated by subjective Mean Opinion Score (MOS, more details in Sec. 7) on text from LibriTTS test sets in Table 2. As can be seen, the synthesized speech obtained about the same naturalness and speaker similarity as the ground truth recordings. The self-similarity between different ground truth recordings from this particular speaker is lower than the average on the corpus, reflecting higher expressiveness and more style variation in her recordings. Raters often commented “lower/higher voice (than the other)” in the similarity evaluation on the ground truth.

2. PnG NAT with voice cloning

To transfer the voices from the source speech to the translation speech, we modified PnG NAT to support zero-shot cross-lingual voice cloning (VC) by incorporating a speaker encoder in the same way as in Jia et al. (2018). The augmented TTS model is illustrated in Figure 2. The speaker encoder is separately trained in a speaker verification task and is frozen during TTS training. At training time, the paired target speech is used as the reference speech for the speaker encoder. At synthesis time, the phonemes and graphemes in the target language and the reference speech in the source language are fed into the model as inputs, and the model produces speech in the target language with the voice from the source speech transferred.

Compared to the speaker encoder (Wan et al., 2018) used in Jia et al. (2018), we used an improved model with better performance. The key improvements include: 1) The model is trained with the generalized end-to-end extended-set softmax loss (Pelecanos et al., 2021); 2) Instead of LSTM, the model is based on a 256×\times12 Conformer stack (Gulati et al., 2020); 3) We introduce an attentive temporal pooling layer (Wang et al., 2022; Pelecanos et al., 2022) to aggregate the Conformer output over time, then concatenate the weighted mean and standard deviation, and finally produce the 256-dim speaker embedding with two feed-forward layers. This speaker encoder has 21.2M parameters, and is trained on a mixture of a proprietary multilingual speech query dataset covering 37 locales collected by vendors, plus public corpora including LibriVox, CN-Celeb (Fan et al., 2020), TIMIT (Garofolo et al., 1993), Fisher (Cieri et al., 2004), and Mixer 4 and 5 (Cieri et al., 2007; Brandschain et al., 2008). The training data contain 122M utterances from 240K speakers in total. Compared to the speaker encoder used in Jia et al. (2018), the speaker verification Equal Error Rate (EER) on LibriSpeech is reduced from 2.5% to 0.9%.

Performance

The performance of the trained model is evaluated on both seen and unseen speakers from LibriTTS in Table 2. The synthesized speech obtained high naturalness and speaker similarity, although lower than ground truth recordings, due to the challenge of zero-shot voice transferring, especially when the reference audios are noisy.

Data generation

The CoVoST 2 corpus includes a few empty audio files (0 byte) originating from Common Voice version 4. We excluded these audios from CVSS. In addition, we used a proprietary voice activity detector (VAD) to filter out audios without any human voice. They in total filtered out 133 recordings from CoVoST 2.

2. Text normalization

We normalize the translation text from CoVoST 2 using a proprietary weighted finite state transducer (WFST)-based text normalizer (Ebden and Sproat, 2015). Non-standard words (Sproat et al., 2001), such as numbers, currency expressions, dates, common abbreviations, acronyms, etc., are detected and verbalized. Such normalized text is used as the input for TTS synthesis.

For S2ST model training and evaluation, we further converted the normalized text into lowercase, and removed punctuation marks except for apostrophes. This version of the normalized translation text is released in CVSS. Appendix A includes examples of such text normalization.

3. TTS synthesis

is synthesized using the PnG NAT model described in Sec. 4.1. A female speaker “lavocedorata” (ID 3983) from LibriTTS is used as the canonical speaker. Although this speaker has merely 6.7 minutes recordings in the training set, these recordings are highly fluent, clean and natural.

CVSS-T

is synthesized using the augmented PnG NAT model described in Sec. 4.2 for cross-lingual voice cloning. The speaker embedding computed on the source non-English speech is used for synthesizing the English translation speech.

Vocoder

A neural vocoder based on WaveRNN (Kalchbrenner et al., 2018) is used for converting the mel-spectrograms synthesized by the TTS models into waveforms. This neural vocoder is trained on a proprietary dataset of 420 hours studio recordings from 98 professional speakers in 6 English accents.

Data format

The synthesized speech is stored as monophonic WAV files at 24 kHz sample rate and in 16-bit linear PCM format.

4. Dataset splitting

Both CVSS-C and CVSS-T are split into train, dev and test subsets consistently with CoVoST 2. CoVoST 2 uses an extended CoVoST split in order to increase data utilization from the raw Common Voice dataset, by allowing multiple versions of recordings on the same sentences (likely from different speakers). This extended split is used for the train set of CoVoST 2, while the original Common Voice split is used for the dev and test sets to avoid skew to duplicate sentences in evaluation. We follow the same data split settings in CVSS.

Statistics

Basic statistics on both versions of CVSS are shown in Table 3. As can be seen, the synthesized translation speech is significantly shorter than the source speech, which is the result of better fluency and the absence of long silences. The duration of CVSS-C is slightly shorter than CVSS-T, indicating faster speaking pace.

The quality of the produced corpus is evaluated as “Targets” rows in Table 4 and Appendix B. CVSS-C obtained very high naturalness, while the naturalness and speaker similarity from CVSS-T is lower. Rater comments revealed that the naturalness of CVSS-T is primarily impacted by “noise” and “distortion”, which is likely the result of noisy reference speech from CoVoST 2 used for voice transferring; the speaker similarity is largely impacted by “different languages”, which does not necessarily reflect voice difference (same as observed in Zhang et al. (2019); Jia et al. (2022)). Objective d-vector similarity on CVSS-T obtained a very high 0.65 despite of the language difference (compared to 0.64 from LibriTTS unseen speakers in a same language in Table 2), suggesting high speaker similarity estimated for speaker verification. We further break down the evaluation on CVSS-T by speech duration in Figure 3, to reflect the impact of the amount of reference audio used for voice cloning.

Despite the naturalness difference between CVSS-T and CVSS-C, they both obtain high intelligibility, as reflected by ASR BLEU. The ASR BLEU is significantly lower on certain languages (e.g. zh) than others, because those data include a lot of non-English names and proper nouns, which cannot be recognized correctly by the English ASR model used in evaluation.

Baseline models

On each version of CVSS, we trained two baseline direct S2ST models (Translatotron and Translatotron 2) as well as a baseline cascade S2ST model (ST→\toTTS). All models are implemented using the Lingvo framework (Shen et al., 2019).

Following Jia et al. (2019b), we evaluated the translation quality and speech generation quality of S2ST models. The translation quality is measured by BLEU on ASR transcription from the translation speech (in lowercase, excluding punctuation marks) against the normalized reference translation. Because ASR makes errors, such BLEU can be thought a lower bound of the translation quality. We used an ASR model from Park et al. (2020) trained on LibriSpeech and LibriLight (Kahn et al., 2020), and computed BLEU using SacreBLEU (Post, 2018) with its default configuration. The speech generation quality is measured subjectively by 5-point mean opinion score (MOS) on naturalness and speaker similarity (Jia et al., 2018). Each MOS evaluation was conducted with 1,000 or more ratings by native North American English speakers. Each rater was limited to rate no more than 6 items per evaluation.

We group the evaluation results on high-resource source languages (French, German, Catalan and Spanish) and low-resource ones (all the rest). MOS evaluation was only conducted on the high-resource language pairs, because otherwise the low translation quality on low-resource languages would negatively impact the subjective assessment of the speech generation quality.

On each version of CVSS, we trained two baseline end-to-end direct S2ST models following Translatotron (Jia et al., 2019b) and Translatotron 2 (Jia et al., 2022). For both models, we followed the hyper-parameters from Sec. 5.5 in Jia et al. (2022) except for a few changes. Notably, we used a wider Conformer encoder (256×\times16) for the larger and more diverse training data. The detailed hyper-parameters are available in Table 7. All models were trained with a batch size of 768 for 240K steps. We picked checkpoints by the best average BLEU on the dev sets, and report the performance on the test sets in Table 4 (detailed in Appendix B).

2. Cascade S2ST baselines

To construct cascade S2ST baselines, we trained an ST model on the original CoVoST 2 corpus, and connected it to the same two TTS models used for constructing CVSS. Note that these cascade models have a data advantage over the direct models at training time (i.e. access to high quality TTS data).

We trained an ST model on the original CoVoST 2 corpus, using the same encoder and decoder architecture and hyper-parameters as in Translatotron 2, except that it predicts 8,192 SentencePiece (Kudo and Richardson, 2018) tokens with a beam size of 8, and was trained with a larger batch size and a higher learning rate (Table 7). This ST model outperforms the previous state-of-the-art ST models trained on CoVoST 2 without extra data by 5.8 or 6.9 BLEU, as average on all 21 or the 4 high-resource language pairs. It even outperforms a few previous works using models more than 15×\times larger, and pre-trained with extra large-scale speech, text, and MT data (although behind even larger ones). See Table 5 for the performance of this ST model and Appendix C for more details. Such improvements over the previous works partially come from the using of a deeper Conformer encoder which learns better speech representation, and we also noted that the extra regularization on the decoder was crucial for avoiding overfitting.

3. Pre-training

We explored utilizing pre-training in ASR and ST tasks to improve the performance of both direct and cascade S2ST models. Such pre-training was conducted within the CoVoST 2 corpus without using extra datasets.

Following Wang et al. (2021b), we pre-trained a multilingual ASR model on all the 22 languages in CoVoST 2, and used it for initializing the ST models for cascade S2ST. These ASR and ST models used the same model architecture and hyper-parameters as in Sec. 7.2 except for using a larger 16k multilingual SentencePiece vocabulary. Similarly, we used the trained ST models to initialize the encoder and decoder of the Translatotron 2 direct S2ST models.

For the simplicity and self-containedness as baselines, we did not explore self-supervised pre-training with extra data in this work. However, such an approach remains promising for improving the performance of S2ST.

4. Results

As can be seen from Table 4, both the cascade model and the Translatotron 2 direct S2ST model produced translation speech as natural as the reference targets, all of which were as natural as human recordings (Table 2) – Thanks to the duration-based autoregressive speech generation (Shen et al., 2020) used in both the PnG NAT TTS model and the Translatotron 2 S2ST model. Both of them also obtained translation quality comparable to the ST evaluation (Table 5), with the cascade model performed slightly better, indicating the effectiveness of both cascade and direct S2ST. The performance of the original Translatotron was behind Translatotron 2 and the cascade S2ST model.

Pre-training (CVSS-C)

Similarly to observed in ST tasks (Weiss et al., 2017; Bansal et al., 2019; Jia et al., 2019a; Wang et al., 2021b), pre-training with weakly supervised data can benefit the performance of the more difficult task. ASR pre-training further improved the performance of our very strong ST model (Table 5), which in turn led to better performance of the cascade S2ST (Table 6). ST pre-training improved the performance of the Translatotron 2 direct S2ST models, to be very close to the cascade S2ST models (with 0.1 / 0.7 BLEU differences as average on all language pairs, when initialized from matching ST models without / with ASR pre-training, Table 6).

CVSS-T

All three models trained on CVSS-T were able to preserve source speakers’ voices during speech translation, with about the same speaker similarity to the source speech as the reference targets. Similar to the results on CVSS-C, both Translatotron 2 and the cascade model obtained about the same naturalness as the reference targets, with the original Translatotron behind them. Both Translatotron 2 and the cascade model also obtained ASR BLEU similar to the same on CVSS-C, indicating the effectiveness of Translatotron 2 as a direct S2ST model capable of voice preservation (the performance of the cascade model is expected since it is consistent with the CVSS-T data construction), as well as the high intelligibility of the translation speech in CVSS-T despite of lower naturalness compared to CVSS-C. Interestingly, the translation quality from the original Translatotron was better on the apparently more difficult CVSS-T dataset than on CVSS-C. This may be explained by the extra task of voice transferring that encouraged its decoder to utilize the attention output. As a matter of fact, inability to pick up attention output is one of the challenges in the original Translatotron tuning (Jia et al., 2019b).

5. Discussion

Although the translation quality from the direct S2ST models did not surpass the cascade models in our experiments, we observed cases where direct S2ST demonstrated advantages over the latter, in terms of avoiding error propagation on rare words, which is a known challenge for ST (Gaido et al., 2021). For example, for a German source speech with content “Mogadischu ist die Hauptstadt von Somalia”, the ST model in the cascade S2ST mistakenly translated the speech corresponding to “Mogadischu” into English text as “UgoDIShu”, which turned into being considered as four words by the downstream TTS model because of text normalization, and finally produced translation speech unable to be understood (ASR transcribed it into “hugo d i shoo”, with “d” and “i” pronounced as individual letters). As a comparison, the Translatotron 2 direct S2ST model mostly copied the pronunciation from the source speech into the translation speech. Although it was not able to be recognized correctly by the ASR model for evaluation (transcribed as “bogodisu”), it was able to be understood by humans. Similar examples were reported in Jia et al. (2019b). This can be a potential advantage of direct S2ST worth further exploration.

Conclusion

We described two massively multilingual-to-English S2ST datasets, CVSS-C and CVSS-T, each with about 1.9K hours of sentence-level parallel S2ST pairs, covering 21 source languages. The translation speech in CVSS-C is in a single canonical speaker’s voice, while the same in CVSS-T is in voices transferred from the source speech. Each dataset provides unique values not existing in other public S2ST corpora.

We built baseline multilingual direct S2ST models and cascade S2ST models on both datasets, verifying the effectiveness of the corpus. To build strong cascade S2ST baselines, we trained an ST model on CoVoST 2, which outperforms the previous state-of-the-art by 5.8 BLEU. Nevertheless, the performance of the direct S2ST models approaches the strong cascade baselines when trained from scratch, and with only 0.1 or 0.7 BLEU difference on ASR transcribed translation when initialized from matching ST models.

Future work includes expanding the corpus coverage to En→\toX directions.

Acknowledgements

We acknowledge the volunteer contributors and the organizers of the Common Voicehttps://commonvoice.mozilla.org/ and LibriVox https://librivox.org/ projects for their contribution and collection of recordings, the creators of Common Voice (Ardila et al., 2020), CoVoST (Wang et al., 2020), CoVoST 2 (Wang et al., 2021b), Librispeech (Panayotov et al., 2015) and LibriTTS (Zen et al., 2019) corpora for their previous works. We would like to thank Ankur Bapna for helps on data processing, Yiling Huang and Jason Pelecanos for improving the speaker encoder model, and Colin Cherry and Alexis Conneau for helpful feedback.

Bibliographical References

Appendix A Examples of normalized text

Appendix B Detailed performance of the S2ST models

Appendix C Detailed performance of the ST models