Direct Simultaneous Speech-to-Text Translation Assisted by Synchronized Streaming ASR
Junkun Chen, Mingbo Ma, Renjie Zheng, Liang Huang
Introduction
Simultaneous speech-to-text translation incrementally translates source-language speech into target-language text, and is widely useful in many cross-lingual communication scenarios such as international travels and multinational conferences. The conventional approach to this problem is a cascaded one Arivazhagan et al. (2020); Xiong et al. (2019); Zheng et al. (2020b), involving a pipeline of two steps. First, the streaming automatic speech recognition (ASR) module transcribes the input speech on the fly Moritz et al. (2020); Wang et al. (2020), and then a simultaneous text-to-text translation module translates the partial transcription into target-language text Oda et al. (2014); Dalvi et al. (2018); Ma et al. (2019); Zheng et al. (2019a, b, 2020a); Arivazhagan et al. (2019).
However, the cascaded approach inevitably suffers from two limitations: (a) error propagation, where streaming ASR’s mistakes confuse the translation module (which are trained on clean text), and this problem worsens with noisy environments and accented speech; and (b) extra latency, where the translation module has to wait until streaming ASR’s output stabilizes, as ASR by default can repeatedly revise its output (see Fig. 1).
To overcome the above issues, some recent efforts Ren et al. (2020); Ma et al. (2020b, a) attempt to directly translate the source speech into target text simultaneously by adapting text-based wait- strategy Ma et al. (2019). However, unlike simultaneous translation whose input is already segmented into words or subwords, in speech translation, the key challenge is to figure out the number of valid tokens within a given source speech segment in order to apply the wait- policy. Ma et al. (2020b, a) simply assume a fixed number of words within a certain number of speech frames, which does not consider various aspects of speech such as different speech rate, duration, pauses and silences, all of which are common in realistic speech. Ren et al. (2020) design an extra Connectionist Temporal Classification (CTC)-based speech segmenter to detect the word boundaries in speech. However, the CTC-based segmenter inherits the same shortcoming of CTC, which only makes local predictions, thus limiting its segmentation accuracy. On the other hand, to alleviate the error propagation, Ren et al. (2020) employ several different knowledge distillation techniques to learn the attentions of ASR and MT jointly. These knowledge distillation techniques are complicated to train and it is an indirect solution for the error propagation problem.
We instead present a simple but effective solution (see Fig. 2) by employing two separate, but synchronized, decoders, one for streaming ASR and the other for End-to-End Speech-to-text Translation (E2E-ST). Our key idea is to use the intermediate results of streaming ASR to guide the decoding policy of, but not feed as input to, the E2E-ST decoder. We look at the beam of streaming ASR to decide the number of tokens within the given source speech segment. Then it is straightforward for the E2E-ST decoder to apply the wait- policy and decide whether to commit a target word or to wait for more speech frames. During training time, we jointly train ASR and E2E-ST tasks with a shared speech encoder in a multi-task learning (MTL) fashion to further improve the translation accuracy. We also note that having streaming ASR as an auxiliary output is extremely useful in real application scenarios where the user often wants to see both the transcription and the translation. En-to-De and En-to-Es experiments on the MuST-C dataset demonstrate that our proposed technique achieves substantially better translation quality at similar level of latency.
Preliminaries
We formalize full-sentence tasks (ASR, MT and ST) using the sequence-to-sequence framework, and the streaming tasks (simultaneous MT and streaming ASR) using the test-time wait- method.
The encoder first encodes the entire source input into a sequence of hidden states; in NMT, the input is a sequence of words, , while in ASR and ST, we use to denote the input speech frames. A decoder sequentially predicts target language tokens in NMT and ST or transcription in ASR, conditioned on all encoder hidden states and previously committed tokens. For example, the NMT model and its parameters are defined as:
Similarly, we can obtain the definitions for ASR () and ST (). Our model was learned from scratch in this work, but it can be improved with pre-training methods Zheng et al. (2021); Chen et al. (2020).
Simultaneous MT and Streaming ASR
In streaming decoding scenarios, we have to predict target tokens conditioned on the partial source input that is available. For example, the test-time wait- method of Ma et al. (2019) predicts each target token after reading source tokens using a full-sentence NMT model: {fleqn}
Intuitively speaking, wait- only commits a new target word on receiving each new source word after an initial source words waiting. Similarly, in the case of streaming ASR, we could define with growing speech chunks that are fed gradually.
Direct Simultaneous Translation with Synchronized Streaming ASR
In text-to-text simultaneous translation, the input stream is already segmented. However, when we deal with speech frames as source inputs, it is not easy to determine the number of valid tokens within certain speech segments. Therefore, to better guide the translation policy, it is essential to detect the number of valid tokens accurately within low latency. Different from the sophisticated design of speech segmenter in Ren et al. (2020), we propose a simple but effective method by using a synchronized streaming ASR and using its beam to determine the number of words within certain speech segments. Note that we only use streaming ASR for source word counting, but the translation decoder does not condition on any of ASR’s output.
As shown in Fig. 2, at inference time, the speech signals are fed into the ST encoder by a series of fixed-size chunks , where can be chosen from 32, 48 and 64 frames of spectrogram. As a result of the CNN encoder, there is down sampling rate (e.g., we use ), from spectrogram to encoder hidden states. For example, when we receive a chunk of 32 frames, the encoder will generate 8 more hidden states. In conventional streaming ASR, the number of steps of beam search is the same as the number of hidden states.
We denote to be the beam at time step , which is an ordered list of size of , and it expands to the next beam with the same size:
where returns the top candidates, and expands the candidates from the previous step to the next step. Each candidate is a pair , where is the current prefix and is the accumulated probability from joint score between an external language model, CTC and ASR probabilities, . We denote the number of observable speech chunks at step as . And vice versa, for each new speech chunk, ASR beam search will advance for steps.
Note CTC often commits empty tokens due to empty speech frames, and the lengths of different hypotheses within beam of streaming ASR are quite different from each other. To take every hypothesis into consideration, we design two policies to decide the number of valid tokens.
Longest Common Prefix (LCP) uses the length of longest shared prefix in the streaming ASR beam as the number of valid tokens within given speech. This is the most conservative strategy, which has similar latency to cascaded methods.
Shortest Hypothesis (SH) uses the length of shortest hypothesis in the current streaming ASR beam as the number of valid tokens.
More formally, let denote the number of valid tokens in the beam under policy :
For example in Fig. 3, , . Also note that for any beam , and that both policies are monotonic, i.e. for and all .
Note we always feed the entire observable speech segments into ST for translation, and streaming ASR-generated transcription is not used for translation, so LCP might have similar latency with cascaded methods but the translation accuracy is much better because more information on the source side is revealed to the translation decoder.
As shown in Algorithm 1, during simultaneous ST, we monitor the value of while speech chunks are gradually fed into system. When we have where is the number of translated tokens, the ST decoder will be triggered to generate one new token as follows: {fleqn}
2 Joint Training between ST and ASR
Different from existing simultaneous translation solutions from Ren et al. (2020); Ma et al. (2020b, a), which make adaptations over vanilla E2E-ST architecture as shown in gray line of Fig. 4, we instead use simple MTL architecture which performs joint full-sentence training between ST and ASR:
For ASR training, we use hybrid CTC/Attention framework Watanabe et al. (2017). Note that we train ASR and ST MTL with full-sentence fashion for simplicity and training efficiency, and only perform wait- decoding policy at inference time. Also, and share the same speech encoder.
Experiments
We conduct experiments on English-to-German (EnDe) and English-Spanish (EnEs) translation on MuST-C Di Gangi et al. (2019). We employ Transformer Vaswani et al. (2017) as the basic architecture and LSTM Hochreiter and Schmidhuber (1997) for LM. For streaming ASR decoding we use a beam size of 5. Translation decoding is greedy due to incremental commitment.
Raw audios are processed with Kaldi Povey et al. (2011) to extract 80-dimensional log-Mel filterbanks stacked with 3-dimensional pitch features using a 10ms step size and a 25ms window size. Text is processed by SentencePiece Kudo and Richardson (2018) with a joint vocabulary size of 8K. We take Transformer Vaswani et al. (2017) as our base architecture, which follows 2 layers of 2D convolution of size 3 with stride size of 2. The Transformer model has 12 encoder layers and 6 decoder layers. Each layer has 4 attention head with a size of 256. Our streaming ASR decoding method follows Moritz et al. (2020). We employ 10 frames look ahead for all experiments. For LM, we use 2 layers stacked LSTM Hochreiter and Schmidhuber (1997) with 1024-dimensional hidden states, and set the embedding size as 1024. LM are trained on English transcription from the corresponding language pair in MuST-C corpus. For the cascaded model, we train ASR and MT models on Must-C dataset respectively, and they have the same Transformer architecture of our ST model. Our experiments are run on 8 1080Ti GPUs. And the we report the case-sensitive detokenized BLEU.
In order to clearly compare with related works, we evaluate the latency with AL defined in Ma et al. (2020b) and AP defined in Ren et al. (2020). As shown in Fig. 5, for EnDe, results are on the dev set to be consistent with Ma et al. (2020b). Compared with baseline models, our method achieves much better translation quality with similar latency. To validate the effectiveness of our method, we compare our method with Ren et al. (2020) on EnEs translation. Their method does not evaluate the plausibility of the detected tokens, so it has a more aggressive decoding policy which results in lower latency. However, our method can still achieve better results with slightly lower latency. Besides that, our model is trained in full-sentence mode, and only decodes with wait- at inference time, which is very efficient to train. Our test-time wait- could achieve similar quality with their genuine wait- (i.e., retrained) models which are very slow to train. When we compare with their test-time wait-, our model significantly outperforms theirs.
We further evaluate our method on the test set of EnDe and EnEs translation. As shown in Fig. 7, compared with the cascaded model, our model has notable successes in latency and translation quality. To verify the online usability of our model, we also show computational-aware latency. Because our chunk window is 480ms, and the latency caused by the computation is smaller than this window size, which means that we can finish decoding the previous speech chunk when the next speech chunk needs to be processed, so our model can be effectively used online.
Fig. 6 demonstrates that our method can effectively avoid the error propagation and obtain better latency compared to the cascaded model.
Effect of chunk size and joint decision
Table 1 shows that the results are relatively stable with various chunk sizes. It can be flexible to balance the response frequency and computational ability. We explore the effectiveness of ASR joint scoring, and observe that the translation quality drops a lot without LM. Without LM and AD, our token recognition approach is similar to the speech segmentation in Ren et al. (2020), which implies that their model is hard to segment the source speech accurately, leading to unreliable translation decisions for ST.
Conclusion
We proposed a simple but effective ASR-assisted simultaneous E2E-ST framework. The streaming ASR module can guide (but not give direct input to) the wait- policy for simultaneous translation. Our method improves ST accuracy with similar latency.
Acknowledgments
This work is supported in part by NSF IIS-1817231 and IIS-2009071.