Towards Fast and Accurate Streaming End-to-End ASR

Bo Li, Shuo-yiin Chang, Tara N. Sainath, Ruoming Pang, Yanzhang He, Trevor Strohman, Yonghui Wu

Introduction

End-to-end (E2E) models have attracted large interest in both academia and industry. These models fold in components of the conventional automatic speech recognition (ASR) systems, namely an acoustic model (AM), pronunciation model (PM) and language model (LM), into a single neural network and optimize them jointly. E2E models simplify ASR system building and maintenance. They can have a much smaller model size than conventional ASR systems and are therefore more suitable for systems that perform the recognition on mobile devices. Among E2E variants, recurrent neural network transducer (RNN-T) has shown potential for on-device streaming ASR .

Besides recognition quality, latency is another critical metric for streaming ASR. In this paper, we define recognition latency as the time difference between when the user stops speaking and when the system produces its final text hypothesis. It is desirable for model latency to be low enough that the system responds to the user quickly, while still high enough that it does not cut off the user’s speech. Building models that have a better trade-off between word error rate (WER) and latency is crucial to achieving fast and accurate streaming speech recognition

The decision of whether a user has stopped speaking is usually generated by an endpointer (EP) model. A voice activity detector (VAD) that detects speech and filters out non-speech is one such model. It can be used to declare an end-of-query (EOQ) as soon as VAD observes speech followed by a fixed interval of silence. VAD is not optimized to distinguish within-speech and query-end silences and may generate many false positive endpointing decisions. EOQ-based models address these issues . They are directly optimized to distinguish speech and different types of silence including initial, intermediate and final. They have been shown to give better latency and WER trade-offs.

Even with EOQ, the endpointer model and the ASR model are still optimized independently. Information captured by ASR models is not shared to the endpointer, which may be useful for making endpointing decisions. It would be better to optimize the endpointer and ASR models together. E2E models make this joint optimization simpler than with conventional modeling approaches. does this by folding the EOQ detector into the RNN-T model by introducing a special token ( ), signaling the end of speech, into RNN-T’s output vocabulary. It is treated the same as all the other tokens during training. However, during inference it is used as one of the signals to end a search path. Premature prediction may cause not only substitution errors but also deletions.

To achieve better WER and latency trade-offs, we not only need the joint optimization of endpointer and ASR, the token should also be predicted as close to the end of the last word as possible. In this work we propose to extend the joint RNN-T endpointer (EP) model in a number of ways. First, we introduce penalties for emitting too early or late in training, to encourage the model to find a good WER and latency trade-off. These penalties are applied to the token, where the ground truth is obtained from a forced alignment between the transcript and audio signals.

Second, premature prediction causes a sequence level loss rather than a single token’s. This leads us to explore whether sequence training would address this problem. We hence investigated minimum word error rate training for RNN-T EP models, which is found to yield both a WER and latency improvements. Third, we rescore RNN-T EP’s hypotheses with a non-streaming model, namely Listen, Attend and Spell (LAS) . The direct modeling of in RNN-T makes the score combination with LAS, which emits already, more consistent. While the rescoring model does not directly change the latency of RNN-T, WER gains it brings gives us more room for potential WER and latency trade-offs. The final setup, RNN-T EP with late penalty, LAS rescoring and MWER training, achieves a 18.7% relative WER reduction and 40ms median latency and 160ms 90-percentile latency reductions on a Voice Search task comparing to the original RNN-T EP .

The rest of the paper is organized as follows. Section 2 explains the model architecture of the RNN-T EP and then details the proposed improvements of the RNN-T EP model using early and late penalties, MWER training and LAS rescoring. Section 3 and 4 presents the experimental setup, results and analysis.

RNN Transducer and Endpointer

Extending RNN-T’s output vocabulary with a special token helps improve its latency , as the endpointing decision is made jointly with the model rather than with a separate endpointer. However, there is no constraint on when should occur during training. A premature prediction can result in deletion errors, while late predictions of can increase latency as is used to inform the system when the speech ends. In this paper, we address these issues by applying additional early and late penalties on the token (Equation (1)). Specifically, during training for every input frame in {x1,…,xT}\{{\bf x}_{1},\dots,{\bf x}_{T}\} and every label {y1,…,yU}\{y_{1},\dots,y_{U}\}, RNN-T computes a U×TU\times T matrix PRNN-T(y∣x){\mathbf{P}}_{\text{RNN-T}}({\bf y}|{\bf x}), which is used in the training loss computation. The last label yUy_{U} is always . We denote t</s>t_{\texttt{</s>}} as the frame index after the last non-silence phoneme, obtained from the forced alignment of the audio with a conventional model. The RNN-T log-probability log⁡PRNN-T(yU∣x)\log{\mathbf{P}}_{\text{RNN-T}}(y_{U}|{\bf x}) is modified to include a penalty at each time step tt for predicting too early or too late. tbuffert_{\text{buffer}} gives a grace period after the reference t</s>t_{\texttt{</s>}} before the late penalty is applied. αearly\alpha_{\text{early}} and αlate\alpha_{\text{late}} are scales on the early and late penalties respectively. All hyper parameters are tuned experimentally.

2 MWER Training

Minimizing RNN-T loss corresponds to improving the log-likelihood of the training data. However, ASR system performance is measured in terms of WER, not log-likelihood. To address this mismatch, proposes to minimize expected WER of the RNN-T model by approximating the expectation with samples draw from the model. Minimum word error rate training (MWER) is later applied to attention based LAS E2E models .

During the beam search decoding of the RNN-T EP model, the inference is terminated when either a blank symbol is generated at the last input frame or an token is predicted. Premature prediction results in deletion of the remaining reference target sequence, leading to a large sequence loss. This makes it more suitable for sequence training techniques. In this work, we hence investigate MWER training with N-best hypotheses for the RNN-T EP model.

3 Listen, Attend and Spell Rescoring

Non-streaming E2E models such as Listen, Attend and Spell (LAS) has shown better performance than streaming ones such as RNN-T. LAS has been explored to serve as a second pass rescorer, that can still fit within the on-device latency constraints. As illustrated in Figure 1, the model first collects the output of the RNN-T encoder of all the frames e=[e1,…,eT]\bf{e}=[\bf{e}_{1},\dots,\bf{e}_{T}]. They are then forwarded through an extra LAS encoder to generate a new set of encoder features for the LAS decoder. The decoder then computes output yLAS\bf{y}_{\text{LAS}} accordingly. During inference, we first pick the top-K hypotheses from the RNN-T decoder. We then run the LAS model on each sequence in the teacher-forcing mode to compute a score, which combines log probability of the sequence and the attention coverage penalty . The sequence with the highest LAS score is picked as the output sequence.

One of the issues in is that RNN-T did not produce a score for , while LAS is indeed trained to produce a score for it. Thus, when rescoring RNN-T hypotheses, an “artificial” score for RNN-T was added to the from LAS. One can argue that including a score for generated from RNN-T based on the inputs should help recognition, as it gives more confidence as to if the sentence should actually be completed. In this work, we look at improving LAS rescoring with the RNN-T EP model by including a score for . The use of token in RNN-T makes the score combination with LAS more consistent acorss all the output units. It is important to note that LAS rescoring cannot make the RNN-T model emit faster; it can only improve the WER of RNN-T. However, the improvement of WER may provide additional room to trade WER for latency.

Experimental Setups

We use the same multidomain dataset as for training. Multistyle training (MTR) is used for noise robustness . During training, a noise configuration, which defines mixing conditions like the size of the room, reverberation time, position of the microphone, speech and noise sources, signal to noise ratio (SNR), etc, for each utterance is randomly sampled from a collection of 3 million pre-generated configurations. The detailed noise configuration can be found in . The test set we use consists of 14K Voice Search utterances with duration less than 5.5 seconds long. They are all anonymized and hand-transcribed, and are representative of Google traffic.

2 Modeling

The input waveforms are framed using a 32 msec window with 10 msec shift. Globally normalized 128 dimension logmel features extracted from frequencyies spanning from 125 Hz to 7.5kHz are used as inputs. The input window size is 4, consisting of 3 frames on the left and no future context. It is further subsampled by a factor of 3 making the system operate at 33 Hz .

Similar to , multidomain models are trained with domain id as an additional input for learning domain-dependent variations. Following , all LSTM layers in the model are unidirectional, with 2048 units and a projection layer with 640 units. The RNN-T encoder consists of 8 LSTM layers, with a time-reduction layer after the second layer. The RNN-T decoder consists of a prediction network with 2 LSTM layers, and a joint network with a single feed-forward layer with 640 units. The additional LAS encoder consists of 2 LSTM layers. The LAS decoder consists of multi-head attention with 4 attention heads, which is fed into 2 LSTM layers. All models are trained on 8x8 Cloud TPU using the Tensorflow Lingvo toolkit to predict 4,096 word pieces including the token.

3 Inference

Despite the use of multidomain training, this work focuses only on the Voice Search task. We append the token only to the Voice Search queries and keep the other data untouched. We report both the recognition performance in terms of word error rate (WER) and the latency of the models for Voice Search only. The latency metrics used in this paper includes median latency (EP50), 90 percentile latency (EP90) and the endpointing coverage (EOU) which represents the percentage of the test data actually receives an end-of-utterance signal from the endpointer model.

There is a trade-off between accuracy and latency, which is often depicted by ROC curves. For EOU EPs, it is obtained by adjusting the endpointing decision threshold. For RNN-T EPs, the endpointing decision is defined by:

α</s>\alpha_{\texttt{</s>}} is a penalty term for the posterior of that modifies the ordering for the hypothesis with . β\beta is a predefined threshold that determines if is allowed in the search beam . Sweeping α</s>\alpha_{\texttt{</s>}} and β\beta gives us a ROC curve of the WER and latency trade-off. For simplicity, we most of the time report a single trade-off point and only show the ROC curves at the end for the final comparisons.

Results

We first train a RNN-T model to predict 4,096 word pieces for the ASR task only (no ) as was done in past . This RNN-T cannot be used to output an endpointing decision and an external EOQ EP is used (B1 in Table 1). The endpointer and the RNN-T ASR model are trained independently and at the inference time, the information from RNN-T’s hypotheses cannot be used for endpointing decisions. To address this issue, we also trained a joint endpointing and recognition RNN-T EP model proposed in (B2 in Table 1). As suggested in , we also use the independnetly trained EOQ EP as a backup for the RNN-T EP model. From Table 1, the RNN-T EP (B2) shows good latency gains (130ms EP50 and 200ms EP90 latency reductions and a 5.4% absolute EOU coverage improvement) but has an increase of 0.3% WER. One assumption of this regression is that during training is treated the same as all the other tokens, with no constraint on how early or late should occur; however in inference, a path ends when a token is predicted. Predicting EOS prematurely brings in deletion errors.

2 Early and Late Penalties

To address the potential premature prediction, we adopt an early penalty term to the training. It is added only if is predicted at any frame earlier than its ground truth time. When adding the early penalty, we scale it by a factor of 0.1 which is found to work well. This (E1 in Table 2) reduces the WER from 7.5% to 7.2% but degrades on latency comparing to B2. The use of early penalty does help the model to address premature prediction but has the risk of the model learning to over-delay its predictions, which leads to worse latency. The regression on EP90 is more sever which is because many tail cases are not endpointed by RNN-T EP and they simply fall back to the EOQ EP.

We further introduce a late penalty term to penalize the prediction that happens too late comparing to the ground truth. During training the granularity of the time is frame (particularly 60ms in our setup). We experimented with tbuffer={3,5,7}t_{\text{buffer}}=\{3,5,7\} which corresponds to a grace period of 180ms, 300ms and 420ms after the reference label. The results are presented in Table 2. With 3 frames’ buffer, we obtain the best median latency but 5-frame gives the best 90-percentile latency which is still worse than B2. We take model E2, namely the RNN-T EP model with early penalty and 3-frame late penalty, as the setup for following experiments.

3 MWER training

For the RNN-T EP model, a wrong prediction of leads to not just a token error but a sequence level loss as it is used to terminate a path in beam search. MWER training optimizes sequence level loss and penalizes WER when is emitted too early, thus prompting our investigation in this section.

We conducted MWER training for the RNN-T model without (B1) and the best RNN-T EP (E2). Both the pre- and post-MWER results are reported in Table 3. For B1, the latency is controlled by a separate EOU EP and hence remains the same after MWER training (B3). But the WER reduces from 7.2% to 6.9%. While for E2, MWER training (E4) maintains the same 7.2% WER as E2 but achieves 220ms EP90 reduction with 50ms regression on EP50. Because optimizing MWER already penalizes premature predictions, we turn off the early penalty for MWER training. This (E5 in Table 3) reduces WER from 7.2% to 6.9% WER and more importantly it still yields a 270ms EP90 latency reduction while maintaining the same EP50 latency as E2. MWER training of the RNN-T EP model with only late penalty can bring in both WER and latency improvements. Comparing to B2, E5 gives 8.0% relative WER reduction and 30ms EP50 and 130ms EP90 latency reductions.

4 LAS Rescoring

So far we see good latency reductions, but the WER gains are small. In the literature, two-pass model that runs RNN-T as the first pass streaming model for fast response and LAS as the rescorer has been shown to be effective in WER reductions. We hence investigate the effect of LAS rescoring on RNN-T EP model. We took the pre-MWER model E2 and added an additional encoder with two LSTM layers and an extra LAS decoder (Figure 1). They are trained with cross entropy (CE) loss with the RNN-T weights frozen. The results are presented as E6 in Table 4. In this work, the latency is only measured for the first pass model. With the same decoding configuration as E2, LAS rescoring reduces the WER by 11.1% relative from 7.2% to 6.4%. Although LAS rescoring cannot directly affect first pass latency, with the WER gains, we may be able to trade WER for latency. We further swept the penalty scale for and obtained an operation point with the same 6.4% WER but 10ms EP50 and 130ms EP90 reductions. As mentioned in Section 2.3, one problem for LAS rescoring of RNN-T without as done in is that RNN-T does not generate an explicit score to combine with that from LAS. To simulate that effect, we zeroed out the score from RNN-T EP and swept a global value to be combined with LAS score. The result (E6 + ignore RNN-T score in Table 4) shows an increase in WER from 6.4% to 6.6%, highlighting the benefit of using in RNN-T EP for LAS rescoring.

Instead of CE loss, MWER loss of RNN-T outputs can be used to update the LAS rescorer (E7 in Table 4). It further reduces the WER down to 6.2% and obtains 100ms EP90 reductions. Moreover, when we update both RNN-T and LAS during MWER (E8), we can obtain another 70ms EP90 reduction. Comparing to B2, the RNN-T EP + LAS with MWER training gives a 18.7% relative WER reduction and 40ms EP50 and 160ms EP90 reductions.

5 Analysis

The proposed RNN-T EP with late penalty and MWER training (E5) gives us both WER and latency improvement over the RNN-T (B1) and the original RNN-T EP (B2). Further WER improvement is achieved via a second pass LAS rescorer (E8). In this section, we compare these systems across different operating points. We plotted the WER vs latency (EP90) curve for these four models (B1, B2, E5, E8) in Figure 2 by varying the penalty scale α</s>\alpha_{\texttt{</s>}} and threshold β\beta. Lower curves are better. RNN-T (B1) tends to delay outputs and has worse latency. With , RNN-T EP (B2) addresses the latency problem but with some WER degradations. With the modifications proposed in this work, namely late penalty, MWER and LAS rescoring, both E5 and E8 have much better WER and latency trade-offs.

References