SimulMT to SimulST: Adapting Simultaneous Text Translation to End-to-End Simultaneous Speech Translation

Xutai Ma, Juan Pino, Philipp Koehn

Introduction

Simultaneous speech translation (SimulST) generates a translation from an input speech utterance before the end of the utterance has been heard. SimulST systems aim at generating translations with maximum quality and minimum latency, targeting applications such as video caption translations and real-time language interpretation. While great progress has recently been achieved on both end-to-end speech translation Ansari et al. (2020) and simultaneous text translation (SimulMT) Grissom II et al. (2014); Gu et al. (2017); Luo et al. (2017); Lawson et al. (2018); Alinejad et al. (2018); Zheng et al. (2019b, a); Ma et al. (2020); Arivazhagan et al. (2019, 2020), little work has combined the two tasks together Ren et al. (2020).

End-to-end SimulST models feature a smaller model size, greater inference speed and fewer compounding errors compared to their cascade counterpart, which perform streaming speech recognition followed by simultaneous machine translation. In addition, it has been demonstrated that end-to-end SimulST systems can have lower latency than cascade systems Ren et al. (2020).

In this paper, we study how to adapt methods developed for SimulMT to end-to-end SimulST. To this end, we introduce the concept of pre-decision module. Such module guides how to group encoder states into meaningful units prior to making a READ/WRITE decision. A detailed analysis of the latency-quality trade-offs when combining a fixed or flexible pre-decision module with a fixed or flexible policy is provided. We also introduce a novel computation-aware latency metric, adapted from Average Lagging (AL) Ma et al. (2019).

Task formalization

A SimulST model takes as input a sequence of acoustic features X=[x1,...x∣X∣]{\bm{X}}=[{\bm{x}}_{1},...{\bm{x}}_{|{\bm{X}}|}] extracted from speech samples every TsT_{s} ms, and generates a sequence of text tokens Y=[y1,...,y∣Y∣]{\bm{Y}}=[y_{1},...,y_{|{\bm{Y}}|}] in a target language. Additionally, it is able to generate yiy_{i} with only partial input X1:n(yi)=[x1,...xn(yi)]{\bm{X}}_{1:n(y_{i})}=[{\bm{x}}_{1},...{\bm{x}}_{n(y_{i})}], where n(yi)≤∣X∣n(y_{i})\leq|{\bm{X}}| is the number of frames needed to generate the ii-th target token yiy_{i}. Note that nn is a monotonic function, i.e. n(yi−1)≤n(yi)n(y_{i-1})\leq n(y_{i}).

A SimulST model is evaluated with respect to quality, using BLEU Papineni et al. (2002), and latency. We introduce two latency evaluation methods for SimulST that are adapted from SimulMT. We first define two types of delays to generate the word yiy_{i}, a computation-aware (CA) and a non computation-aware (NCA) delay. The CA delay of yiy_{i}, dCA(yi)d_{\text{CA}}(y_{i}), is defined as the time that elapses (speech duration) from the beginning of the process to the prediction of yiy_{i}, while the NCA delay for yiy_{i} dCA(yi)d_{\text{CA}}(y_{i}) is defined by dNCA(yi)=Ts⋅n(yi)d_{\text{NCA}}(y_{i})=T_{\text{s}}\cdot n(y_{i}). Note that dNCAd_{\text{NCA}} is an ideal case for dCAd_{\text{CA}} where the computational time for the model is ignored. Both delays are measured in milliseconds. Two types of latency measurement, LCAL_{CA} and LNCAL_{NCA}, are calculated accordingly: L=C(D)L=\mathcal{C}({\bm{D}}) where C\mathcal{C} is a latency metric and D=[d(y1),...,d(y∣Y∣)]\mathbf{D}=[d(y_{1}),...,d(y_{|{\bm{Y}}|})].

To better evaluate the latency for SimulST, we introduce a modification to AL. We assume an oracle system that can perform perfect simultaneous translation for both latency and quality, while in Ma et al. (2019) the oracle is ideal only from the latency perspective. We evaluate the lagging based on time rather than steps. The modified AL metric is defined in Eq. 1:

where ∣Y∗∣|{\bm{Y}}^{*}| is the length of the reference translation, τ(∣X∣)\tau(|{\bm{X}}|) is the index of the first target token generated when the model has read the full input. There are two benefits from this modification. The first is that latency is measured using time instead of steps, which makes it agnostic to preprocessing and segmentation. The second is that it is more robust and can prevent an extremely low and trivial value when the prediction is significantly shorter than the reference.

Method

End-to-end ST models directly map a source speech utterance into a sequence of target tokens. We use the S-Transformer architecture proposed by (Di Gangi et al., 2019b), which achieves competitive performance on the MuST-C dataset Di Gangi et al. (2019a). In the encoder, a two-dimensional attention is applied after the CNN layers and a distance penalty is introduced to bias the attention towards short-range dependencies.

We investigate two types of simultaneous translation mechanisms, flexible and fixed policy. In particular, we investigate monotonic multihead attention Ma et al. (2020), which is an instance of flexible policy and the prefix-to-prefix model Ma et al. (2019), an instance of fixed policy, designated by wait-kk from now on.

(MMA) Ma et al. (2020) extends monotonic attention Raffel et al. (2017); Arivazhagan et al. (2019) to Transformer-based models. Each head in each layer has an independent step probability pijp_{ij} for the iith target and jjth source step, and then uses a closed form expected attention for training. A weighted average and variance loss were proposed to control the behavior of the attention heads and thus the trade-offs between quality and latency.

(Ma et al., 2019) is a fixed policy that waits for kk source tokens, and then reads and writes alternatively. Wait-kk can be a special case of Monotonic Infinite-Lookback Attention (MILk) Arivazhagan et al. (2019) or MMA where the step-wise probability pij=0p_{ij}=0 if j−i<kj-i<k else pij=1p_{ij}=1.

2 Pre-Decision Module

In SimulMT, READ or WRITE decisions are made at the token (word or BPE) level. However, with speech input, it is unclear when to make such decisions. For example, one could choose to read or write after each frame or after generating each encoder state. Meanwhile, a frame typically only covers 10ms of the input while an encoder state generally covers 40ms of the input (assuming a subsampling factor of 4), while the average length of a word in our dataset is 270ms. Intuitively, a policy like wait-kk will not have enough information to write a token after reading a frame or generating an encoder state. In principle, a flexible or model-based policy such as MMA should be able to handle granular input. Our analysis will show, however, that while MMA is more robust to the granularity of the input, it also performs poorly when the input is too fine-grained.

In order to overcome these issues, we introduce the notion of pre-decision module, which groups frames or encoder states, prior to making a decision. A pre-decision module generates a series of trigger probabilities ptrp_{tr} on each encoder states to indicate whether a simultaneous decision should be made. If ptr>0.5p_{tr}>0.5, the model triggers the simultaneous decision making, otherwise keeps reading new frames. We propose two types of pre-decision module.

A straightforward policy for a fixed pre-decision module is to trigger simultaneous decision making every fixed number of frames. Let Δt\Delta t be the time corresponding to this fixed number of frames, with Δt\Delta t a multiple of TsT_{s}, and re=int(∣X∣/∣H∣)r_{e}=\text{int}(|{\bm{X}}|/|{\bm{H}}|). ptrp_{tr} at encoder step jj is defined in Eq. 2:

We use an oracle flexible pre-decision module that uses the source boundaries either at the word or phoneme level. Let A{\bm{A}} be the alignment between encoder states and source labels (word or phoneme). A(hi){\bm{A}}(h_{i}) represents the token that hih_{i} aligns to. The trigger probability can then be defined in Eq. 3:

Experiments

We conduct experiments on the English-German portion of the MuST-C dataset Di Gangi et al. (2019a), where source audio, source transcript and target translation are available. We train on 408 hours of speech and 234k sentences of text data. We use Kaldi Povey et al. (2011) to extract 80 dimensional log-mel filter bank features, computed with a 25msms window size and a 10msms window shift. For text, we use SentencePiece Kudo and Richardson (2018) to generate a unigram vocabulary of size 10,000. We use Gentlehttps://lowerquality.com/gentle/ to generate the alignment between source text and speech as the label to generate the oracle flexible pre-decision module. Translation quality is evaluated with case-sensitive detokenized BLEU with SacreBLEU Post (2018). The latency is evaluated with our proposed modification of AL Ma et al. (2019). All results are reported on the MuST-C dev set.

All speech translation models are first pre-trained on the ASR task where the target vocabulary is character-based, in order to initialize the encoder. We follow the same hyperparameter settings from Di Gangi et al. (2019b). We follow the latency regularization method introduced by Ma et al. (2020); Arivazhagan et al. (2019). The objective function to optimize is

Where C\mathcal{C} is a latency metric (AL in this case) and D{\bm{D}} is described in Section 2. Only samples with AL>0\text{AL}>0 are regularized to avoid overfitting. For the models with monotonic multihead attention, we first train a model without latency with λlatency=0\lambda_{\text{latency}}=0. After the model converges, λlatency\lambda_{\text{latency}} is set to a desired value and we continue trining the model until convergence.

The latency-quality trade-offs of the 4 types of model from the combination of fixed or flexible pre-decision with fixed or flexible policy are presented in Fig. 2. The non computation-aware delays are used to calculate the latency metric in order to evaluate those trade-offs from a purely algorithmic perspective.

(Fig. 2(a)). As expected, both quality and latency increase with step size and lagging. In addition, the latency-quality trade-offs are highly dependent on the step size of the pre-decision module. For example, with step size 120ms, the performance is very poor even with large kk because of very limited information being read before writing a target token. Large step sizes improve the quality but introduce a lower bound on the latency. Note that step size 280ms, which provides an effective latency-quality trade-off compared to other step sizes, also matches the average word length of 271ms. This motivates the study of a flexible pre-decision module based on word boundaries.

(Fig. 2(b)) Similar to wait-kk, MMA obtains very poor performance with a small step size of 120ms. For other step sizes, MMA obtains similar latency-quality trade-offs, demonstrating some form of robustness to the step size.

Curve ⋆\star and in figure Fig. 2 show latency-quality trade-offs when the pre-decision module is determined by oracle word or phoneme boundaries. Note that a SimulST model would not normally have access to this information and that the purpose of this experiment is to guide future design of a flexible pre-decision model. First, as previously observed, the granularity of the pre-decision greatly influences the latency-quality trade-offs. Models using phoneme boundaries obtain very poor translation quality because those boundaries are too granular, with an average phoneme duration of 77ms. In addition, comparing MMA and wait-kk with phoneme boundaries, MMA is found to be more robust to the granularity of the pre-decision.

The best settings for each approach are compared in Fig. 3. For fixed pre-decision, we choose the setting that has the best quality for each latency bucket of 500ms, while for the flexible pre-decision we use oracle word boundaries. For both wait-kk and MMA, the flexible pre-decision module outperforms the fixed pre-decision module. This is expected since the flexible pre-decision module uses oracle information in the form of pre-computed word boundaries but provides a direction for future research. The best latency-quality trade-offs are obtained with MMA and flexible pre-decision from word boundaries.

We also consider the computation-aware latency described in Section 2, shown in Fig. 4. The focus is on fixed pre-decision approaches in order to understand the relation between the granularity of the pre-decision and the computation time. Fig. 4 shows that as the step size increases, the difference between the NCA and the CA latency shrinks. This is because with larger step sizes, there is less overhead of recomputing the bidirectional encoder states This is a common practice in SimulMT where the input length is significantly shorter than in SimulST Arivazhagan et al. (2019); Ma et al. (2019); Arivazhagan et al. (2020). We recommend future work on SimulST to make use of CA latency as it reflects a more realistic evaluation, especially in low-latency regimes, and is able to distinguish streaming capable systems.

Conclusion

We investigated how to adapt SimulMT methods to end-to-end SimulST by introducing the concept of pre-decision module. We also adapted Average Lagging to be computation-aware. The effects of combining a fixed or flexible pre-decision module with a fixed or flexible policy were carefully analyzed. Future work includes building an incremental encoder to reduce the CA latency and design a learnable pre-decision module.

References