Streaming Simultaneous Speech Translation with Augmented Memory Transformer

Xutai Ma, Yongqiang Wang, Mohammad Javad Dousti, Philipp Koehn, Juan Pino

Introduction

Streaming speech translation targets low latency scenarios such as simultaneous interpretation. Unlike the streaming automatic speech recognition (ASR) task where input and output are monotonically aligned, translation needs a larger future context due to reordering. Generally, simultaneous translation models start to translate with partial input, and then alternate between generating predictions and consuming additional input. While most previous work on simultaneous translation focus on text input, the end-to-end approach for simultaneous speech translation has also very recently attracted interest from the community due to potentially lower latency compared with cascade models . However, most studies tend to focus on an ideal setup, where the computation time to generate the translation is neglected. This assumption may be reasonable for text to text but not for speech to text translation since the latter has much longer input sequences. A simultaneous speech translation model may have the ability to generate translations with partial input but may not be useful for real-time applications because of slow computation in generating output tokens.

While the majority of previous work on streaming speech to text tasks has focused on ASR, most prior work on ASR is not directly applicable to translation. Encoder-only or transducer structures are widely implemented, since ASR assumes the output is monotonically aligned to the input. In order to achieve an efficient streaming simultaneous speech translation model, we combine streaming ASR and simultaneous translation techniques and introduce an end-to-end transformer-based speech translation model with an augmented memory encoder .

The augmented memory encoder has shown considerable improvements on latency with little sacrifice on quality with hybrid or transducer-based models on the ASR task. It incrementally encodes fixed-length sub-sentence level segments and stores the history information with a memory bank, which summarizes each segment. The self-attention is only performed on the current segment and memory banks. A decoder with simultaneous policies is then introduced on top of the encoder. We apply the proposed model to the simultaneous speech translation task on the MuST-C dataset .

This paper is organized as follows. We first define the evaluation method for simultaneous speech translation. We then introduce the model based on the augmented memory transformer. Finally, we conduct a series of experiments to demonstrate the effectiveness of the proposed approach.

Evaluation

This paper focuses on streaming and simultaneous speech translation, which features two additional capabilities compared to a traditional offline model. The first one is the efficient computation needed to handle streaming input, and the second is the ability to start a translation with partial input, then dynamically generate additional tokens or read additional input. Both factors need to be considered for evaluation.

A simultaneous system is evaluated with respect to quality, usually with BLEU, and latency. Latency is evaluated with computation-aware and non computation-aware Average Lagging (AL) . Denote the input sequence as X=[x1,...]\mathbf{X}=[\mathbf{x}_{1},...], where each element is a feature vector extracted from a sliding window of size TT and the translation of the system Y=[y1,...]\mathbf{Y}=[y_{1},...] and a reference translation Y∗\mathbf{Y}^{*}. Non computation-aware (NCA) latency is defined as follows

where τ(∣X∣)\tau(|X|) is the index of the first target token when the system has read the entire source input, and d(yi)d(y_{i}) is the duration of the speech that has been read when generating word yiy_{i}. Additionally, a computation-aware (CA) version of AL is also considered, by replace d(yi)d(y_{i}) with the time needed to generate yiy_{i}.

Model

The proposed streaming speech translation model, illustrated in Fig. 1, consists of two components, an augmented memory encoder and a simultaneous decoder. The encoder incrementally and efficiently encodes streaming input, while the decoder starts translation with partial input, then interleaves reading new input and predicting target token under the guidance of a simultaneous translation policy.

The self-attention module in the original transformer model attends to the entire input sequence, which precludes streaming capability. Denote H=[h1,..]\mathbf{H}=[\mathbf{h}_{1},..] the input of a certain encoder layer. Each self-attention projects the input into query, key and value.

At each position jj, a weight is calculated as follows

The self-attention at position jj can then be calculated as

The calculation of self-attention make it inefficient for streaming applications. proposes an augmented memory transformer encoder to address this issue. Instead of attending to entire input sequence X\mathbf{X}, the self-attentions are applied on a sequence of sub-utterance level segments S=[s1,...]\mathbf{S}=[\mathbf{s}_{1},...]. A segment sn\mathbf{s}_{n}, which contains a span of input features, consists of three parts: left context ln\mathbf{l}_{n} of size LL, main context cn\mathbf{c}_{n} of size CC and right context rn\mathbf{r}_{n} of size RR. Each segment overlaps with adjacent segments — the overlap between current and previous segment is ln\mathbf{l}_{n}, and between current and the next segment is rn\mathbf{r}_{n}. Self-attention is computed at the segment level, which reduces the amount of computation. The new query, key and value for each segment are

Where σn=∑xk∈snxk\sigma_{n}=\sum_{\mathbf{x}_{k}\in\mathbf{s}_{n}}\mathbf{x}_{k} is a summarization of the segment sn\mathbf{s}_{n}, and Mn−N:n=[mn−N,...,mn−1]\mathbf{M}_{n-N:n}=[\mathbf{m}_{n-N},...,\mathbf{m}_{n-1}] are the memory banks. Each memory bank is calculated as follows:

which is introduced to represent history information. A hyperparameter NN controls how many memory banks are retained. The self-attention is then calculated as follows:

Then only the central encoder states are kept and a the concatenation of the segment states Z=[z1,...]\mathbf{Z}=[\mathbf{z}_{1},...] is passed to decoder. Because of the left and right contexts, an arbitrary encoder can run on the segments without boundary mismatch. In this paper, we adapt the encoder of the convtransformer architecture . The encoder first consists of convolutional layers with stride 2 that subsample the input. Full self-attention layers can then be calculated.

2 Simultaneous Decoder

A simultaneous decoder starts translation with partial input based on a policy. Simultaneous policies decide whether the model should read new inputs or generate a new prediction at a given time. However, different from text translation, our preliminary experiments show that for simultaneous speech translation, encoder states are too granular for policy learning. Thus, we adopt the idea of pre-decision for better efficiency by making simultaneous read and write decision on chunks of encoder states. Here, the simpler fixed pre-decision strategy is used where the decision is made every fixed number of encoder states. Denote the sequence of chunks are W=[w1,...]\mathbf{W}=[\mathbf{w}_{1},...] and the start and end encoder state index of wk\mathbf{w}_{k} is Ws(k),We(k)W_{s}(k),W_{e}(k). Denote the prediction of model Y=[y1,...]\mathbf{Y}=[y_{1},...], the general decoding algorithm of a simultaneous policy P\mathcal{P} with augmented memory transformer is described in Algorithm 1.

In theory, Algorithm 1 supports arbitrary simultaneous translation policies. In this paper, for simplicity, wait-kk is used. It waits for kk source tokens and then operating then reading and writing alternatively. Notice that our method is compatible with an arbitrary simultaneous translation policy.

Note that the decoder self-attention still has access to all previous decoder hidden states; in order to preserve streaming capability for the decoder, decoder states are reset every time an end-of-sentence token is predicted. The augmented memory is not introduced in the decoder because the target sequence is dramatically smaller than the source speech sequence. The decoder can still predict a token in a negligible time compared with encoding source with the input becoming longer.

Experiments

Experiments were conducted on the English-German MuST-C dataset . The training data consists of 408 hours of speech and 234k sentences of text. We use Kaldi to extract 80 dimensional log-mel filter bank features. The features are computed with a 25msms window size and a 10msms window shift and normalized with global cepstral mean and variance. Text is tokenized with a SentencePiecehttps://github.com/google/sentencepiece 10k unigram vocabulary. Translation quality is evaluated with case-sensitive detokenized BLEU with SacreBLEUhttps://github.com/mjpost/sacrebleu. The latency is evaluated by Average Lagging , with the SimulEval toolkithttps://github.com/facebookresearch/SimulEval.

The speech translation model is based on the convtransformer architecture . It first contains two convolutional layers with subsampling ratio of 4. Both encoder and decoder have a hidden size of 256 and 4 attention heads. There are 12 encoder layers and 6 decoder layers. The model is trained with label smoothed (0.1) cross entropy. We use the Adam optimizer , with a learning rate of 0.0001 and an inverse square root schedule.

We use a simplified version of as our baseline model. A unidirectional mask is introduced to prevent the encoder from looking into future information. For baseline models, we follow common practice for simultaneous text translation where the entire encoder is updated once there is new input. Instead of a multi-task setting and a decision making process depending on word boundaries obtained from the auxiliary ASR task, we make decisions on a fixed size chunk of encoder states, following . Our choice is motivated by the fact that in , a fixed chunk size gave similar quality-latency trade-offs as word boundaries.

All transformer-based speech translation models are first pre-trained on the ASR task, in order to initialize the encoder. Each experiment is run on 8 Tesla V100 GPUs with 32 GB memory. The code is developed based on Fairseqhttps://github.com/pytorch/fairseq, and it will be published upon acceptance. All scores are reported on the dev set.

Results

We first analyze the effect of the segment and context sizes, and use the resulting optimal settings for further analysis on the maximum number of memory banks and for comparison with the baseline.

We analyze the effect of different segment, left and right context sizes. For all experiments, we use the wait-kk (k=1,3,5,7k=1,3,5,7) policy on a chunk of 8 encoder states . The latency-quality trade-offs with different sizes are shown in Figure 2. We first observe that, increasing the left and right context size will improve the quality with very small trade-off on latency, for instance, from curve “S64 L16 R16” to “S64 L32 R32”. This indicates that context of both sides can alleviate the boundary effect. We also notice that when we reduce the segment size from 64 to 32, the BLEU score decreases dramatically. Similar observations are made in but the ASR models are more robust to decreasing the segment and context sizes. We hypothesize that reordering in translation makes the model more sensitive to these sizes.

2 Number of Memory Banks

In streaming translation, the input is theoretically infinite. In order to prevent memory explosion, we explore the effect of reducing the number of the memory banks. Fig. 3 shows the effect of different numbers of memory banks.

We can see that the model is very robust to the size of the memory banks. Similar to , when the maximum number of memory banks is large, for instance, larger than 3, there is little or no performance drop. However, we still observe a drop in performance with a maximum number of one memory bank. Finally, we found that training with different maximum numbers of memory banks was necessary as limiting the number of memory banks only at inference time degraded performance.

3 Comparison with Baseline

In Fig. 4, we compared our model with the baseline model described in Section 4. The proposed model achieves better quality with an increase in computation aware and non computation aware latency.

The baseline achieves competitive latency because it only updates encoder states every 8 steps. However, there may be instances where recomputing encoder states every step may be needed, for example in the case of a flexible pre-decision module or when a the model includes a boundary detector . We can see in Fig. 4 that the computation aware AL for the baseline increases substantially with a chunk of size 1.

Conclusion

In this paper, we tackle the real-life application of streaming simultaneous speech translation. We propose a transformer-based model, equipped with an augmented memory, in order to handle long or streaming input. We study the effect of segment and context sizes, and the maximum number of memory banks. We show that our model has better quality with an acceptable latency increase compared with a transformer with unidirectional mask baseline and presents better quality-latency trade-offs than that baseline where encoder states are recomputed at every step.

References