Direct Simultaneous Speech-to-Text Translation Assisted by Synchronized Streaming ASR

Junkun Chen, Mingbo Ma, Renjie Zheng, Liang Huang

Introduction

Simultaneous speech-to-text translation incrementally translates source-language speech into target-language text, and is widely useful in many cross-lingual communication scenarios such as international travels and multinational conferences. The conventional approach to this problem is a cascaded one Arivazhagan et al. (2020); Xiong et al. (2019); Zheng et al. (2020b), involving a pipeline of two steps. First, the streaming automatic speech recognition (ASR) module transcribes the input speech on the fly Moritz et al. (2020); Wang et al. (2020), and then a simultaneous text-to-text translation module translates the partial transcription into target-language text Oda et al. (2014); Dalvi et al. (2018); Ma et al. (2019); Zheng et al. (2019a, b, 2020a); Arivazhagan et al. (2019).

However, the cascaded approach inevitably suffers from two limitations: (a) error propagation, where streaming ASR’s mistakes confuse the translation module (which are trained on clean text), and this problem worsens with noisy environments and accented speech; and (b) extra latency, where the translation module has to wait until streaming ASR’s output stabilizes, as ASR by default can repeatedly revise its output (see Fig. 1).

To overcome the above issues, some recent efforts Ren et al. (2020); Ma et al. (2020b, a) attempt to directly translate the source speech into target text simultaneously by adapting text-based wait-kk strategy Ma et al. (2019). However, unlike simultaneous translation whose input is already segmented into words or subwords, in speech translation, the key challenge is to figure out the number of valid tokens within a given source speech segment in order to apply the wait-kk policy. Ma et al. (2020b, a) simply assume a fixed number of words within a certain number of speech frames, which does not consider various aspects of speech such as different speech rate, duration, pauses and silences, all of which are common in realistic speech. Ren et al. (2020) design an extra Connectionist Temporal Classification (CTC)-based speech segmenter to detect the word boundaries in speech. However, the CTC-based segmenter inherits the same shortcoming of CTC, which only makes local predictions, thus limiting its segmentation accuracy. On the other hand, to alleviate the error propagation, Ren et al. (2020) employ several different knowledge distillation techniques to learn the attentions of ASR and MT jointly. These knowledge distillation techniques are complicated to train and it is an indirect solution for the error propagation problem.

We instead present a simple but effective solution (see Fig. 2) by employing two separate, but synchronized, decoders, one for streaming ASR and the other for End-to-End Speech-to-text Translation (E2E-ST). Our key idea is to use the intermediate results of streaming ASR to guide the decoding policy of, but not feed as input to, the E2E-ST decoder. We look at the beam of streaming ASR to decide the number of tokens within the given source speech segment. Then it is straightforward for the E2E-ST decoder to apply the wait-kk policy and decide whether to commit a target word or to wait for more speech frames. During training time, we jointly train ASR and E2E-ST tasks with a shared speech encoder in a multi-task learning (MTL) fashion to further improve the translation accuracy. We also note that having streaming ASR as an auxiliary output is extremely useful in real application scenarios where the user often wants to see both the transcription and the translation. En-to-De and En-to-Es experiments on the MuST-C dataset demonstrate that our proposed technique achieves substantially better translation quality at similar level of latency.

Preliminaries

We formalize full-sentence tasks (ASR, MT and ST) using the sequence-to-sequence framework, and the streaming tasks (simultaneous MT and streaming ASR) using the test-time wait-kk method.

The encoder first encodes the entire source input into a sequence of hidden states; in NMT, the input is a sequence of words, x=(x1,x2,...,xm){\mathbf{x}}=(x_{1},x_{2},...,x_{m}), while in ASR and ST, we use s\mathbf{s} to denote the input speech frames. A decoder sequentially predicts target language tokens y=(y1,y2,...,yn){\mathbf{y}}=(y_{1},y_{2},...,y_{n}) in NMT and ST or transcription z\mathbf{z} in ASR, conditioned on all encoder hidden states and previously committed tokens. For example, the NMT model and its parameters θfullMT\bm{\theta}_{\text{full}}^{\text{MT}} are defined as:

Similarly, we can obtain the definitions for ASR (pfull(z∣s;θfullASR)p_{\text{full}}(\mathbf{z}\mid\mathbf{s};{\bm{\theta}}_{\text{full}}^{\text{ASR}})) and ST (pfull(y∣s;θfullST)p_{\text{full}}({\mathbf{y}}\mid\mathbf{s};\bm{\theta}_{\text{full}}^{\text{ST}})). Our model was learned from scratch in this work, but it can be improved with pre-training methods Zheng et al. (2021); Chen et al. (2020).

Simultaneous MT and Streaming ASR

In streaming decoding scenarios, we have to predict target tokens conditioned on the partial source input that is available. For example, the test-time wait-kk method of Ma et al. (2019) predicts each target token yty_{t} after reading source tokens x≤t+k{\mathbf{x}}_{\leq t+k} using a full-sentence NMT model: {fleqn}

Intuitively speaking, wait-kk only commits a new target word on receiving each new source word after an initial kk source words waiting. Similarly, in the case of streaming ASR, we could define z^t\hat{z}_{t} with growing speech chunks s‾i\overline{\mathbf{s}}_{i} that are fed gradually.

Direct Simultaneous Translation with Synchronized Streaming ASR

In text-to-text simultaneous translation, the input stream is already segmented. However, when we deal with speech frames as source inputs, it is not easy to determine the number of valid tokens within certain speech segments. Therefore, to better guide the translation policy, it is essential to detect the number of valid tokens accurately within low latency. Different from the sophisticated design of speech segmenter in Ren et al. (2020), we propose a simple but effective method by using a synchronized streaming ASR and using its beam to determine the number of words within certain speech segments. Note that we only use streaming ASR for source word counting, but the translation decoder does not condition on any of ASR’s output.

As shown in Fig. 2, at inference time, the speech signals are fed into the ST encoder by a series of fixed-size chunks s‾[1:i]=[s‾1,...,s‾i]\overline{\mathbf{s}}_{[1:i]}=[\overline{\mathbf{s}}_{1},...,\overline{\mathbf{s}}_{i}], where w=∣s‾i∣w=|\overline{\mathbf{s}}_{i}| can be chosen from 32, 48 and 64 frames of spectrogram. As a result of the CNN encoder, there is down sampling rate rr (e.g., we use r=4r=4), from spectrogram to encoder hidden states. For example, when we receive a chunk of 32 frames, the encoder will generate 8 more hidden states. In conventional streaming ASR, the number of steps of beam search is the same as the number of hidden states.

We denote BjB_{j} to be the beam at time step jj, which is an ordered list of size of bb, and it expands to the next beam Bj+1B_{j+1} with the same size:

where top⁡b(⋅)\operatornamewithlimits{\mathbf{top}}^{b}(\cdot) returns the top bb candidates, and next(B,j)\mathit{next}(B,j) expands the candidates from the previous step to the next step. Each candidate is a pair ⟨z,p⟩\langle{\mathbf{z},p}\rangle, where z\mathbf{z} is the current prefix and pp is the accumulated probability from joint score between an external language model, CTC and ASR probabilities, p^fullASR\hat{p}_{\text{full}}^{\text{ASR}}. We denote the number of observable speech chunks at jj step as τ(j)=⌈j∗r/w⌉\tau(j)=\left\lceil j*r/w\right\rceil. And vice versa, for each new speech chunk, ASR beam search will advance for w/rw/r steps.

Note CTC often commits empty tokens ϵ\epsilon due to empty speech frames, and the lengths of different hypotheses within beam of streaming ASR are quite different from each other. To take every hypothesis into consideration, we design two policies to decide the number of valid tokens.

Longest Common Prefix (LCP) uses the length of longest shared prefix in the streaming ASR beam as the number of valid tokens within given speech. This is the most conservative strategy, which has similar latency to cascaded methods.

Shortest Hypothesis (SH) uses the length of shortest hypothesis in the current streaming ASR beam as the number of valid tokens.

More formally, let ϕπ(B)\phi_{\pi}(B) denote the number of valid tokens in the beam BB under policy π\pi:

For example in Fig. 3, ϕLCP(B7) ⁣= ⁣3\phi_{\text{LCP}}(B_{7})\!=\!3, ϕSH(B7) ⁣= ⁣5\phi_{\text{SH}}(B_{7})\!=\!5. Also note that ϕLCP(B)≤ϕSH(B)\phi_{\text{LCP}}(B)\leq\phi_{\text{SH}}(B) for any beam BB, and that both policies are monotonic, i.e. ϕπ(Bj)≤ϕπ(Bj+1)\phi_{\pi}(B_{j})\leq\phi_{\pi}(B_{j+1}) for π∈{LCP,SH}\pi\in\{\text{LCP},\text{SH}\} and all jj.

Note we always feed the entire observable speech segments into ST for translation, and streaming ASR-generated transcription is not used for translation, so LCP might have similar latency with cascaded methods but the translation accuracy is much better because more information on the source side is revealed to the translation decoder.

As shown in Algorithm 1, during simultaneous ST, we monitor the value of ϕπ(Bj)\phi_{\pi}(B_{j}) while speech chunks are gradually fed into system. When we have ϕπ(B)−k⩾t\phi_{\pi}(B)-k\geqslant t where tt is the number of translated tokens, the ST decoder will be triggered to generate one new token as follows: {fleqn}

2 Joint Training between ST and ASR

Different from existing simultaneous translation solutions from Ren et al. (2020); Ma et al. (2020b, a), which make adaptations over vanilla E2E-ST architecture as shown in gray line of Fig. 4, we instead use simple MTL architecture which performs joint full-sentence training between ST and ASR:

For ASR training, we use hybrid CTC/Attention framework Watanabe et al. (2017). Note that we train ASR and ST MTL with full-sentence fashion for simplicity and training efficiency, and only perform wait-kk decoding policy at inference time. Also, θfullST\bm{\theta}_{\text{full}}^{\text{ST}} and θfullASR{\bm{\theta}}_{\text{full}}^{\text{ASR}} share the same speech encoder.

Experiments

We conduct experiments on English-to-German (En→\toDe) and English-Spanish (En→\toEs) translation on MuST-C Di Gangi et al. (2019). We employ Transformer Vaswani et al. (2017) as the basic architecture and LSTM Hochreiter and Schmidhuber (1997) for LM. For streaming ASR decoding we use a beam size of 5. Translation decoding is greedy due to incremental commitment.

Raw audios are processed with Kaldi Povey et al. (2011) to extract 80-dimensional log-Mel filterbanks stacked with 3-dimensional pitch features using a 10ms step size and a 25ms window size. Text is processed by SentencePiece Kudo and Richardson (2018) with a joint vocabulary size of 8K. We take Transformer Vaswani et al. (2017) as our base architecture, which follows 2 layers of 2D convolution of size 3 with stride size of 2. The Transformer model has 12 encoder layers and 6 decoder layers. Each layer has 4 attention head with a size of 256. Our streaming ASR decoding method follows Moritz et al. (2020). We employ 10 frames look ahead for all experiments. For LM, we use 2 layers stacked LSTM Hochreiter and Schmidhuber (1997) with 1024-dimensional hidden states, and set the embedding size as 1024. LM are trained on English transcription from the corresponding language pair in MuST-C corpus. For the cascaded model, we train ASR and MT models on Must-C dataset respectively, and they have the same Transformer architecture of our ST model. Our experiments are run on 8 1080Ti GPUs. And the we report the case-sensitive detokenized BLEU.

In order to clearly compare with related works, we evaluate the latency with AL defined in Ma et al. (2020b) and AP defined in Ren et al. (2020). As shown in Fig. 5, for En→\toDe, results are on the dev set to be consistent with Ma et al. (2020b). Compared with baseline models, our method achieves much better translation quality with similar latency. To validate the effectiveness of our method, we compare our method with Ren et al. (2020) on En→\toEs translation. Their method does not evaluate the plausibility of the detected tokens, so it has a more aggressive decoding policy which results in lower latency. However, our method can still achieve better results with slightly lower latency. Besides that, our model is trained in full-sentence mode, and only decodes with wait-kk at inference time, which is very efficient to train. Our test-time wait-kk could achieve similar quality with their genuine wait-kk (i.e., retrained) models which are very slow to train. When we compare with their test-time wait-kk, our model significantly outperforms theirs.

We further evaluate our method on the test set of En→\toDe and En→\toEs translation. As shown in Fig. 7, compared with the cascaded model, our model has notable successes in latency and translation quality. To verify the online usability of our model, we also show computational-aware latency. Because our chunk window is 480ms, and the latency caused by the computation is smaller than this window size, which means that we can finish decoding the previous speech chunk when the next speech chunk needs to be processed, so our model can be effectively used online.

Fig. 6 demonstrates that our method can effectively avoid the error propagation and obtain better latency compared to the cascaded model.

Effect of chunk size and joint decision

Table 1 shows that the results are relatively stable with various chunk sizes. It can be flexible to balance the response frequency and computational ability. We explore the effectiveness of ASR joint scoring, and observe that the translation quality drops a lot without LM. Without LM and AD, our token recognition approach is similar to the speech segmentation in Ren et al. (2020), which implies that their model is hard to segment the source speech accurately, leading to unreliable translation decisions for ST.

Conclusion

We proposed a simple but effective ASR-assisted simultaneous E2E-ST framework. The streaming ASR module can guide (but not give direct input to) the wait-kk policy for simultaneous translation. Our method improves ST accuracy with similar latency.

Acknowledgments

This work is supported in part by NSF IIS-1817231 and IIS-2009071.

References