Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss

Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, Shankar Kumar

Introduction

In the past few years, models employing self-attention have achieved state-of-art results for many tasks, such as machine translation, language modeling, and language understanding . In particular, large Transformer-based language models have brought gains in speech recognition tasks when used for second-pass re-scoring and in first-pass shallow fusion . As typically used in sequence-to-sequence transduction tasks , Transformer-based models attend over encoder features using decoder features, implying that the decoding has to be done in a label-synchronous way, thereby posing a challenge for streaming speech recognition applications. An additional challenge for streaming speech recognition with these models is that the number of computations for self-attention increases quadratically with input sequence size. For streaming to be computationally practical, it is highly desirable that the time it takes to process each frame remains constant relative to the length of the input. Transformer-based alternatives to RNNs have recently been explored for use in ASR .

For streaming speech recognition models, recurrent neural networks (RNNs) have been the de facto choice since they can model the temporal dependencies in the audio features effectively while maintaining a constant computational requirement for each frame. Streamable end-to-end modeling architectures such as the Recurrent Neural Network Transducer (RNN-T) , Recurrent Neural Aligner (RNA) , and Neural Transducer utilize an encoder-decoder based framework where both encoder and decoder are layers of RNNs that generate features from audio and labels respectively. In particular, the RNN-T and RNA models are trained to learn alignments between the acoustic encoder features and the label encoder features, and so lend themselves naturally to frame-synchronous decoding.

Several optimization techniques have been evaluated to enable running RNN-T on device . In addition, extensive architecture and modeling unit exploration has been done for RNN-T . In this paper, we explore the possibility of replacing RNN-based audio and label encoders in the conventional RNN-T architecture with Transformer encoders. With a view to preserving model streamability, we show that Transformer-based models can be trained with self-attention on a fixed number of past input frames and previous labels. This results in a degradation of performance (compared to attending to all past input frames and labels), but then the model satisfies a constant computational requirement for processing each frame, making it suitable for streaming. Given the simple architecture and parallelizable nature of self-attention computations, we observe large improvements in training time and training resource utilization compared to RNN-T models that employ RNNs.

The RNN-T architecture), and Eq. (4), and ”RNN-T loss”, defined in Eq. (5), to refer to the loss used to train this architecture. (as depicted in Figure 1) is a neural network architecture that can be trained end-to-end with the RNN-T loss to map input sequences (e.g. audio feature vectors) to target sequences (e.g. phonemes, graphemes). Given an input sequence of real-valued vectors of length TT, x=(x1,x2,...,xT){\mathbf{x}}=(x_{1},x_{2},...,x_{T}), the RNN-T model tries to predict the target sequence of labels y=(y1,y2,...,yU){\mathbf{y}}=(y_{1},y_{2},...,y_{U}) of length UU.

Unlike a typical attention-based sequence-to-sequence model, which attends over the entire input for every prediction in the output sequence, the RNN-T model gives a probability distribution over the label space at every time step, and the output label space includes an additional null label to indicate the lack of output for that time step — similar to the Connectionist Temporal Classification (CTC) framework . But unlike CTC, this label distribution is also conditioned on the previous label history.

The RNN-T model defines a conditional distribution P(z∣x)P({\mathbf{z}}|{\mathbf{x}}) over all the possible alignments, where

is a sequence of (zi,ti)(z_{i},t_{i}) pairs of length U‾\overline{U}, and (zi,ti)(z_{i},t_{i}) represents an alignment between output label ziz_{i} and the encoded feature at time tit_{i}. The labels ziz_{i} can optionally be blank labels (null predictions). Removing the blank labels gives the actual output label sequence y{\mathbf{y}}, of length UU.

We can marginalize P(z∣x)P({\mathbf{z}}|{\mathbf{x}}) over all possible alignments z{\mathbf{z}} to obtain the probability of the target label sequence y{\mathbf{y}} given the input sequence x{\mathbf{x}},

where Z(y,T){\cal Z}({\mathbf{y}},T) is the set of valid alignments of length TT for the label sequence.

Transformer Transducer

In this paper, we present all experimental results with the RNN-T loss for consistency, which performs similarly to the monotonic RNN-T loss in our experiments.

The probability of an alignment P(z∣x)P({\mathbf{z}}|{\mathbf{x}}) can be factorized as

To compute Eq. (1) by summing all valid alignments naively is computationally intractable. Therefore, we define the forward variable α(t,u)\alpha(t,u) as the sum of probabilities for all paths ending at time-frame tt and label position uu. We then use the forward algorithm to compute the last alpha variable α(T,U)\alpha({T,U}), which corresponds to P(y∣x)P({\mathbf{y}}|{\mathbf{x}}) defined in Eq. (1). Efficient computation of P(y∣x)P({\mathbf{y}}|{\mathbf{x}}) using the forward algorithm is enabled by the fact that the local probability estimate (Eq. (4)) at any given label position and any given time-frame is not dependent on the alignment . The training loss for the model is then the sum of the negative log probabilities defined in Eq. (1) over all the training examples,

where TiT_{i} and UiU_{i} are the lengths of the input sequence and the output target label sequence of the ii-th training example, respectively.

2 Transformer

Experiments and Results

2 Transformer Transducer

3 Results

We first compared the performance of Transformer Transducer (T-T) models with full attention on audio to an RNN-T model using a bidirectional LSTM audio encoder. As shown in Table 2, the T-T model significantly outperforms the LSTM-based RNN-T baseline. We also observed that T-T models can achieve competitive recognition accuracy with existing wordpiece-based end-to-end models with similar model size. To compare with systems using shallow fusion with separately trained LMs, we also trained a Transformer-based LM with the same architecture as the label encoder used in T-T, using the full 810M word token dataset. This Transformer LM (6 layers; 57M parameters) had a perplexity of 2.492.49 on the dev-clean set; the use of dropout, and of larger models, did not improve either perplexity or WER. Shallow fusion was then performed using that LM and both the trained T-T system and the trained bidirectional LSTM-based RNN-T baseline, with scaling factors on the LM output and on the non-blank symbol sequence length tuned on the LibriSpeech dev sets. The results are shown in Table 2 in the “With LM” column. The shallow fusion result for the T-T system is competitive with corresponding results for top-performing existing systems.

Similarly, we explored the use of limited right context to allow the model to see some future audio frames, in the hope of bridging the gap between a streamable T-T model (left = 10, right = 0) and a full attention T-T model (left = 512, right = 512). Since we apply the same mask for every layer, the latency introduced by using right context is aggregated over all the layers. For example, in Figure 3, to produce y7y_{7} from a 3-layer Transformer with one frame of right context, it actually needs to wait for x10x_{10} to arrive, which is 90 ms latency in our case. To explore the right context impact for modeling, we did comparisons with fixed 512 frames left context per layer to compared with full attention T-T model. As we can see from Table 4, with right context of 6 frames per layer (around 3.2 secs of latency), the performance is around 16% worse than full attention model. Compared with streamable T-T model, 2 frames right context per layer (around 1 sec of latency) brings around 30% improvements.

Finally, Table 6 reports the results when using a limited left context of 10 frames, which reduces the time complexity for one-step inference to a constant, with look-ahead to future frames, as a way of bridging the gap between the performance of left-only attention and full attention models.

Conclusions

In this paper, we presented the Transformer Transducer model, embedding Transformer based self-attention for audio and label encoding within the RNN-T architecture, resulting in an end-to-end model that can be optimized using a loss function that efficiently marginalizes over all possible alignments and that is well-suited to time-synchronous decoding. This model achieves a new state-of-the-art accuracy on the LibriSpeech benchmark, and can easily be used for streaming speech recognition by limiting the audio and label context used in self-attention. Transformer Transducer models train significantly faster than LSTM based RNN-T models, and they allow us to trade recognition accuracy and latency in a flexible manner.

References