Streaming automatic speech recognition with the transformer model

Niko Moritz, Takaaki Hori, Jonathan Le Roux

Introduction

Hybrid hidden Markov model (HMM) based automatic speech recognition (ASR) systems have provided state-of-the-art results for the last few decades . End-to-end ASR systems, which approach the speech-to-text conversion problem using a single sequence-to-sequence model, have recently demonstrated competitive performance . The most popular and successful end-to-end ASR approaches are based on connectionist temporal classification (CTC) , recurrent neural network (RNN) transducer (RNN-T) , and attention-based encoder-decoder architectures . RNN-T based ASR systems achieve state-of-the-art ASR performance for streaming/online applications and are successfully deployed in production systems . Attention-based encoder-decoder architectures, however, are the best performing end-to-end ASR systems , but they cannot be easily applied in a streaming fashion, which prevents them from being used more widely in practice. To overcome this limitation, different methods for streaming ASR with attention-based systems haven been proposed such as the neural transducer (NT) , monotonic chunkwise attention (MoChA) , and triggered attention (TA) . The NT relies on traditional block processing with fixed window size and stride to produce incremental attention model outputs. The MoChA approach uses an extra layer to compute a selection probability that defines the length of the output label sequence and provides an alignment to chunk the encoder state sequence prior to soft attention. The TA system requires that the attention-based encoder-decoder model is trained jointly with a CTC objective function, which has also been shown to improve attention-based systems , and the CTC output is used to predict an alignment that triggers the attention decoding process . A frame-synchronous one-pass decoding algorithm for joint CTC-attention scoring was proposed in to further optimize and enhance ASR decoding using the TA concept.

Besides the end-to-end ASR modeling approach, the underlying neural network architecture is of paramount importance as well to achieve good ASR performance. RNN-based architectures, such as the long short-term memory (LSTM) neural network, are often applied for end-to-end ASR systems. Bidirectional LSTMs (BLSTMs) achieve state-of-the-art results among such RNN-based systems but are unsuitable for application in a streaming fashion, where unidirectional LSTMs or latency-controlled BLSTMs (LC-BLSTMs) must be applied instead . The parallel time-delayed LSTM (PTDLSTM) architecture has been proposed to further reduce the word error rate (WER) gap between unidirectional and bidirectional architectures and to improve the computational complexity compared to the LC-BLSTM . Recently, the transformer model, which is an encoder-decoder type of architecture based on self-attention originally proposed for machine translation , has been applied to ASR with promising results and improved WERs compared to RNN-based architectures .

In this work, we apply time-restricted self-attention to the encoder, and the TA concept to the encoder-decoder attention mechanism of the transformer model to enable the application of online/streaming ASR. The transformer model is jointly trained with a CTC objective to optimize training and decoding results as well as to enable the TA concept . For joint CTC-transformer decoding and scoring, we employ the frame-synchronous one-pass decoding algorithm proposed in .

Streaming Transformer

The streaming architecture of the proposed transformer-based ASR system is shown in Fig. 1. The transformer is an encoder-decoder type of architecture that uses two different attention layers: encoder-decoder attention and self-attention. The encoder-decoder attention can produce variable output lengths by using one or multiple query vectors, the decoder states, to control attention to a sequence of input values, the encoder state sequence. In self-attention (SA), the queries, values, and keys are derived from the same input sequence, which results in an output sequence of the same length. Both attention types of the transformer model are based on the scaled dot-product attention mechanism,

The encoder of our transformer architecture consists of a two-layer CNN module EncCNN and a stack of EE self-attention layers EncSA:

where X=(x1,…,xT)X=(\bm{x}_{1},\dots,\bm{x}_{T}) denotes a sequence of acoustic input features, which are 80-dimensional log-mel spectral energies plus 3 extra features for pitch information . Both CNN layers of EncCNN use a stride of size 22, a kernel size of 3×33\times 3, and a ReLU activation function. Thus, the striding reduces the frame rate of output sequence X0X_{0} by a factor of 4 compared to the feature frame rate of XX. The EncSA module of (5) consists of EE layers, where the ee-th layer, for e=1,…,Ee=1,\dots,E, is a composite of a multi-head self-attention layer

where x1:n+εenc0=X0[1:n+εenc]=(x10,…,xn+εenc0)\bm{x}_{1:n+\varepsilon^{\text{enc}}}^{0}=X_{0}[1\mathbin{:}n+\varepsilon^{\text{enc}}]=(\bm{x}_{1}^{0},\dots,\bm{x}_{n+\varepsilon^{\text{enc}}}^{0}), and εenc\varepsilon^{\text{enc}} denotes the number of look-ahead frames used by the time-restricted self-attention mechanism.

2 Decoder: Triggered attention

The encoder-decoder attention mechanism of the transformer model is using the TA concept to enable the decoder to operate in a streaming fashion. TA training requires an alignment between the encoder state sequence XEX_{E} and the label sequence Y=(y1,…,yL)Y=(y_{1},\dots,y_{L}) to condition the attention mechanism of the decoder only on past encoder frames plus a fixed number of look-ahead frames εdec\varepsilon^{\text{dec}}. This information is generated by forced alignment using an auxiliary CTC objective pctc(Y∣XE)p_{\text{ctc}}(Y|X_{E}) , which is jointly trained with the decoder model, where the encoder neural network is shared .

The triggered attention objective function is defined as

with νl=nl′+εdec\nu_{l}=n^{\prime}_{l}+\varepsilon^{\text{dec}}, where nl′n^{\prime}_{l} denotes the position of the first occurrence of label yly_{l} in the CTC forced alignment sequence , y1:l−1=(y1,…,yl−1)\bm{y}_{1:l-1}=(y_{1},\dots,y_{l-1}), and x1:νlE=(x1E,…,xνlE)\bm{x}_{1:\nu_{l}}^{E}=(\bm{x}_{1}^{E},\dots,\bm{x}_{\nu_{l}}^{E}), which corresponds to the truncated encoder sequence. The term p(yl∣y1:l−1,x1:νlE)p(y_{l}|\bm{y}_{1:l-1},\bm{x}_{1:\nu_{l}}^{E}) represents the transformer decoder model

for d=1,…,Dd=1,\dots,D, where DD denotes the number of decoder layers. Function Embed converts the input label sequence (⟨s⟩,y1,…,yl−1)(\langle\text{s}\rangle,y_{1},\dots,y_{l-1}) into a sequence of trainable embedding vectors z1:l0\bm{z}_{1:l}^{0}, where ⟨sos⟩\langle\text{sos}\rangle denotes the start of sentence symbol. Function DecTA finally predicts the posterior probability of label yly_{l} by applying a fully-connected projection layer to zlD\bm{z}_{l}^{D} and a softmax distribution over that output.

The CTC model and the triggered attention model of (10) are trained jointly using the multi-objective loss function

where hyperparameter γ\gamma controls the weighting between the two objective functions pctcp_{\text{ctc}} and ptap_{\text{ta}}.

3 Positional encoding

Sinusoidal positional positional encodings (PE) are added to the sequences X0X_{0} and Z0Z_{0}, which can be written as

4 Joint CTC-triggered attention decoding

The frame-by-frame processing of the CTC posterior probability sequence pctcp_{\text{ctc}} and the encoder state sequence XEX_{E} is shown from line 5 to 26, where pctc(n)p_{\text{ctc}}(n) denotes the CTC posterior probability distribution at frame nn. The function CTCPrefix follows the CTC prefix beam search algorithm described in , which extends the set of prefixes Ω\Omega using the CTC posterior probabilities pctcp_{\text{ctc}} of the current time step nn and returns the separate CTC prefix scores pbp_{\text{b}} and pnbp_{\text{nb}} as well as the newly proposed set of prefixes Ωctc\Omega_{\text{ctc}}. A local pruning threshold of 0.0001 is used by CTCPrefix to ignore labels of lower CTC probability. Note that no language model or any pruning technique is used by CTCPrefix, they will be incorporated in the following steps.

Experiments

The LibriSpeech data set, which is a speech corpus of read English audio books , is used to benchmark ASR systems presented in this work. LibriSpeech is based on the open-source project LibriVox and provides about 960 hours of training data, 10.7 hours of development data, and 10.5 hours of test data, whereby the development and test data sets are both split into approximately two halves named “clean” and “other”. The separation into clean and other is based on the quality of the recorded utterance, which was assessed using an ASR system .

2 Settings

The LM weight, CTC weight, and beam size of the full-sequence based joint CTC-attention decoding method are set to 0.7, 0.5, and 20 for the small transformer model and to 0.6, 0.4, and 30 for the large model setup. The parameter settings for CTC prefix beam search decoding are LM weight α0=0.7\alpha_{0}=0.7, pruning beam width θ1=16.0\theta_{1}=16.0, insertion bonus β=2.0\beta=2.0, and pruning size K=30K=30. Parameters for joint CTC-TA decoding are CTC weight λ=0.5\lambda=0.5, CTC LM weight α0=0.7\alpha_{0}=0.7, LM weight α=0.5\alpha=0.5, pruning beam width θ1=16.0\theta_{1}=16.0, pruning beam width θ2=6.0\theta_{2}=6.0, insertion bonus β=2.0\beta=2.0, pruning size K=300K=300, and pruning size P=30P=30. All decoding hyperparameter settings are determined using the development data sets of LibriSpeech.

3 Results

Table 1 presents ASR results of our transformer-based baseline systems, which are jointly trained with CTC to optimize training convergence and ASR accuracy . Results of different decoding methods are shown with and without using the RNN-LM, SpecAugment , and the large transformer model. Table 1 demonstrates that joint CTC-attention decoding provides significantly better ASR results compared to CTC or attention decoding alone, whereas CTC prefix beam search decoding attains lower WERs compared to attention beam search decoding, except for the dev-clean, dev-other, and test-other conditions when no LM is used. For attention beam search decoding, we normalize the log posterior probabilities of the transformer model and the RNN-LM scores when combining both using the hypothesis lengths . Still our attention results are worse compared to the CTC results, which is unexpected but demonstrates that joint decoding stabilizes the transformer results.

Conclusions

In this paper, a fully streaming end-to-end ASR system based on the transformer architecture is proposed. Time-restricted self-attention is applied to control the latency of the encoder and the triggered attention (TA) concept to control the output latency of the decoder. For streaming recognition and joint CTC-transformer model scoring, a frame-synchronous one-pass decoding algorithm is applied, which demonstrated similar LibriSpeech ASR results compared to full-sequence based CTC-attention as the number of look-ahead frames is increased. Combined with the time-restricted self-attention encoder, our proposed TA-based streaming ASR system achieved WERs of 2.8%2.8\% and 7.2%7.2\% for the test-clean and test-other data sets of LibriSpeech, which to our knowledge is the best published LibriSpeech result of a fully streaming end-to-end ASR system.

References