FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization
Jiahui Yu, Chung-Cheng Chiu, Bo Li, Shuo-yiin Chang, Tara N. Sainath, Yanzhang He, Arun Narayanan, Wei Han, Anmol Gulati, Yonghui Wu, Ruoming Pang
INTRODUCTION
End-to-end (E2E) recurrent neural network transducer (RNN-T) models have gained enormous popularity for streaming ASR applications, as they are naturally streamable . However, naive training with a sequence transduction objective to maximize the log-probability of target sequence is unregularized and these streaming models learn to predict better by using more context, causing significant emission delay (i.e., the delay between the user speaking and the text appearing). Recently there are some approaches trying to regularize or penalize the emission delay. For example, Li et al. proposed Early and Late Penalties to enforce the prediction of (end of sentence) within a reasonable time window given by a voice activity detector (VAD). Constrained Alignments were also proposed by extending the penalty terms to each word, based on speech-text alignment information generated from an existing speech model.
While being successful in terms of reducing latency of streaming RNN-T models, these two regularization approaches suffer from accuracy regression . One important reason is because both regularization techniques penalize the per-token or per-frame prediction probability independently, which is inconsistent with the sequence-level transducer optimization of per-sequence probability calculated by the transducer forward-backward algorithm . Although some remedies like second-pass Listen, Attend and Spell (LAS) rescorer and minimum word error rate (MWER) training technique have been used to reduce the accuracy regression, these approaches come at a non-negligible compute cost in both training and serving.
In this work, we propose a novel sequence-level emission regularization method for streaming models based on transducers, which we call FastEmit. FastEmit is designed to be directly applied on the transducer forward-backward per-sequence probability, rather than individual per-token or per-frame prediction of probability independently. In breif, in RNN-T it first extends the output vocabulary space with a ‘blank token’ , meaning ‘output nothing’. Then the transducer forward-backward algorithm calculates the probability of each lattice (speech-text alignment) in the matrix, where and is the length of input and output sequence respectively. Finally the optimal lattice in this matrix can be automatically learned by maximizing log-probability of the target sequence. It is noteworthy that in this transducer optimization, emitting a vocabulary token and the blank token are treated equally, as long as the log-probability of the target sequence can be maximized. However, in streaming ASR systems the blank token ‘output nothing’ should be discouraged as it leads to higher emission latency. We will show in detail that FastEmit, as a sequence-level regularization method, encourages emitting vocabulary tokens and suppresses blank tokens across the entire sequence based on transducer forward-backward probabilities, leading to significantly lower emission latency while retaining recognition accuracy.
FastEmit has many advantages over other regularization methods to reduce emission latency in end-to-end streaming ASR models: (1) FastEmit is a sequence-level regularization based on transducer forward-backward probabilities, thus is more suitable when applied jointly with the sequence-level transducer objective. (2) FastEmit does not require any speech-word alignment information either by labeling or generated from an existing speech model. Thus it is easy to ‘plug and play’ in any transducer model on any dataset without any extra effort. (3) FastEmit has minimal hyper-parameters to tune. It only introduces one hyper-parameter to balance the transducer loss and regularization loss. (4) There is no additional training or serving cost to apply FastEmit.
We apply FastEmit on various end-to-end streaming ASR networks including RNN-Transducer , Transformer-Transducer , ConvNet-Transducer and Conformer-Transducer . We achieve latency reduction with significantly better accuracy over previous methods on a Voice Search test set. FastEmit also improves streaming ASR accuracy from to WER, meanwhile reduces 90th percentile latency from ms to only on LibriSpeech.
TRANSDUCER WITH FastEmit
In this section, we first delve into transducer and show why naively optimizing the transducer objective is unregularized thus unsuitable for low-latency streaming ASR models. We then propose FastEmit as a sequence-level emission regularization method to regularize the emission latency.
Transducer optimization automatically learns probabilistic alignments between an input sequence and an output sequence , where and denote the length of input and output sequences respectively. To learn the probabilistic alignments, it first extends the output space with a ‘blank token’ (meaning ‘output nothing’, visually denoted as right arrows in Figure 1 and 2): . The allocation of these blank tokens then determines an alignment between the input and output sequences. Given an input sequence , the transducer aims to maximize the log-probability of a conditional distribution:
where is a function that removes the tokens from each alignment lattice , and is the ground truth output sequence tokenized from text label.
As shown in Figure 1, we denote each node as the probability of emitting the first elements of the output sequence by the first frames of the input sequence. We further denote the prediction from a neural network and as the probability of label token (up arrows in figures) and blank token (right arrows in figures) at node . To optimize the transducer objective, an efficient forward-backward algorithm is used to calculate the probability of each alignment and aggregate all possible alignments before propagating gradients back to and . It is achieved by defining forward variable as the probability of emitting during , and backward variable as the probability of emitting during , using an efficient forward-backward propagation algorithm:
where the initial conditions are , . It is noteworthy that defines the probability of all complete alignments in :
By diffusion analysis of the probability of all alignments, we know that is equal to the sum of over any top-left to bottom-right diagonal nodes (i.e., all complete alignments will pass through any diagonal cut in the matrix in Figure 1) :
Finally, gradients of transducer loss function w.r.t. neural network prediction of probability and can be calculated according to Equations 1, 2, 3, 4 and 5.
2 FastEmit
Now let us consider any node in the matrix, for example, the blue node at , as shown in Figure 2. First we know that the probability of emitting during is . At the next step, the alignment can either ‘go up’ by predicting label to the green node with probability , or ‘turn right’ by predicting blank to the red node with probability . Finally together with backward probability of the new node, the probability of all complete alignments passing through node in Equation 4 can be decomposed into two parts:
which is equivalent as replacing in Equation 4 with Equation 3. From Equation 6 we know that gradients of transducer loss w.r.t. the probability prediction of any node have following properties (closed-form gradients can be found in Equation 20):
However, this transducer loss aims to maximize log-probability of all possible alignments, regardless of their emission latency. In other words, as shown in Figure 2, emitting a vocabulary token and the blank token are treated equally, as long as the log-probability is maximized. It inevitably leads to emission delay because streaming ASR models learn to predict better by using more future context, causing significant emission delay.
By the decomposition in Equation 6, we propose a simple and effective transducer regularization method, FastEmit, which encourages predicting label instead of blank by additionally maximizing the probability of ‘predict label’ based on Equation 1, 5 and 6:
To interpret the gradients of FastEmit, intuitively it simply means that the gradients of emitting label tokens has a ‘higher learning rate’ back-propagating into the streaming ASR network, while emitting blank token remains the same. We also note that the proposed FastEmit regularization method is based on alignment probabilities instead of per-token or per-frame prediction of probability, thus we refer it as sequence-level emission regularization.
EXPERIMENTAL DETAILS
Our latency metrics of streaming ASR are motivated by real-world applications like Voice Search and Smart Home Assistants. In this work we mainly measure two types of latency metrics described below: (1) partial recognition latency on both LibriSpeech and MultiDomain datasets, and (2) endpointer latency on MultiDomain dataset. A visual example of two latency metrics is illustrated in Figure 3. For both metrics, we report both 50-th (medium) and 90-th percentile values of all utterances in the test set to better characterize latency by excluding outlier utterances.
Partial Recognition (PR) Latency is defined as the timestamps difference of two events as illustrated in Figure 3: (1) when the last token is emitted in the finalized recognition result, (2) the end of the speech when a user finishes speaking estimated by forced alignment. PR latency is especially descriptive of user experience in real-world streaming ASR applications like Voice Search and Assistants. Moreover, PR latency is the lower bound for applying other techniques like Prefetching , by which streaming application can send early server requests based on partial/incomplete recognition hypotheses to retrieve relevant information and necessary resources for future actions. Finally, unlike other latency metrics that may depend on hardware, environment or system optimization, PR latency is inherented to streaming ASR models and thus can better characterize the emission latency of streaming ASR. It is also noteworthy that models that capture stronger contexts can emit a hypothesis even before they are spoken, leading to a negative PR latency.
Endpointer (EP) Latency is different from PR latency and it measures the timestamps difference between: (1) when the streaming ASR system predicts the end of the query (EOQ), (2) the end of the speech when a user finishes speaking estimated by forced alignment. As illustrated in Figure 3, EOQ can be implied by jointly predicting the token with end-to-end Endpointing introduced in . The endpointer can be used to close the microphone as soon as the user finishes speaking, but it is also important to avoid cutting off users while they are still speaking. Thus, the prediction of the token has a higher latency compared with PR latency, as shown in Figure 3. Note that PR latency is also a lower bound of EP latency, thus reducing the PR latency is the main focus of this work.
2 Dataset and Training Details
We report our results on two datasets, a public dataset LibriSpeech and an internal large-scale dataset MultiDomain .
Our main results and ablation studies will be presented on a widely used public dataset LibriSpeech , which consists of about 1000 hours of English reading speech. For data processing, we extract 80-channel filterbanks feature computed from a 25ms window with a stride of 10ms, use SpecAugment for data augmentation, and train with the Adam optimizer. We use a single layer LSTM as the decoder. All of these training settings follow the previous work for fair comparison. We train our LibriSpeech models on 960 hours of LibriSpeech training set with labels tokenized using a 1,024 word-piece model (WPM), and report our test results on LibriSpeech TestClean and TestOther (noisy).
We also report our results a production dataset MultiDomain , which consists of 413,000 hours speech, 287 million utterances across multiple domains including Voice Search, YouTube, and Meetings. Multistyle training (MTR) is used for noise robustness. These training and testing utterances are anonymized and hand-transcribed, and are representatives of Google’s speech recognition traffic. All models are trained to predict labels tokenized using a 4,096 word-piece model (WPM). We report our results on a test set of 14K Voice Search utterances with duration less than 5.5 seconds long.
3 Model Architectures
FastEmit can be applied to any transducer model on any dataset without any extra effort. To demonstrate the effectiveness of our proposed method, we apply FastEmit on a wide range of transducer models including RNN-Transducer , Transformer-Transducer , ConvNet-Transducer and Conformer-Transducer . We refer the reader to the individual papers for more details of each model architecture. For each of our experiment, we keep the exact same training and testing settings including model size, model regularization (weight decay, variational noise, etc.), optimizer, learning rate schedule, input noise and augmentation, etc. All models are implemented, trained and benchmarked based on Lingvo toolkit .
All these model architectures are based on encoder-decoder transducers. The encoders are based on autoregressive models using uni-directional LSTMs, causal convolution and/or left-context attention layers (no future context is permitted). The decoders are based on prediction network and joint network similar to previous RNN-T models . For all experiments on LibriSpeech, we report results directly after training with the transducer objective. For all our experiments on MultiDomain, results are reported with minimum word error rate (MWER) finetuning for fair comparison.
RESULTS
In this section, we first report our results on LibriSpeech dataset and compare with other streaming ASR networks. We next study the hyper-parameter in FastEmit to balance transducer loss and regularization loss. Finally, we conduct large-scale experiments on the MultiDomain production dataset and compare FastEmit with other methods on a Voice Search test set.
We first present results of FastEmit on both Medium and Large size streaming ContextNet and Conformer in Table 1. We did a small hyper-parameter sweep of and set for ContextNet and for Conformer. FastEmit significantly reduces PR latency by . It is noteworthy that streaming ASR models that capture stronger contexts can emit the full hypothesis even before they are spoken, leading to a negative PR latency. We also find FastEmit even improves the recognition accuracy on LibriSpeech. By error analysis, the deletion errors have been significantly reduced. As LibriSpeech is long-form spoken-domain read speech, FastEmit encourages early emission of labels thus helps with vanishing gradients problem in long-form RNN-T , leading to less deletion errors.
2 Hyper-parameter λ𝜆\lambda in FastEmit
Next we study the hyper-parameter of FastEmit regularization by applying different values on M-size streaming ContextNet . As shown in Table 2, larger leads to lower PR latency of streaming models. But when the is larger than a certain threshold, the WER starts to degrade due to the regularization being too strong. Moreover, also offers flexibility of WER-latency trade-offs.
3 Large-scale Experiments on MultiDomain
Finally we show that FastEmit regularization method is also effective on the large scale production dataset MultiDomain. In Table 3, we apply FastEmit on RNN-Transducer , Transformer-Transducer and Conformer-Transducer . For RNN-T, we also compare FastEmit with other methods . All results are finetuned with minimum word error rate (MWER) training technique for fair comparison. In Table 3, CA denotes constrained alignment , MaskFrame denotes the idea of training RNN-T models with incomplete speech by masking trailing frames to encourage a stronger decoder thus can emit faster. We perform a small hyper-parameter search for both baselines CA and MaskFrame and report their WER, EP and PR latency on a Voice Search test set. FastEmit achieves latency reduction with significantly better accuracy over baseline methods in RNN-T , and generalizes further to Transformer-T and Conformer-T . By error analysis, as Voice Seach is short-query written-domain conversational speech, emitting faster leads to more errors. Nevertheless, among all techniques in Table 3, FastEmit achieves best WER-latency trade-off.