Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture

Haoran Miao, Gaofeng Cheng, Changfeng Gao, Pengyuan Zhang, Yonghong Yan

Introduction

In recent years, the end-to-end (E2E) automatic speech recognition (ASR) has gained popularity in ASR community . E2E ASR models simplify the hybrid DNN/HMM ASR models by replacing the acoustic, pronunciation and language models with one single deep neural network, and thus transcribe speech to text directly. To date, E2E ASR models have achieved significant improvement in ASR field . The hybrid Connectionist Temporal Classification (CTC) / attention E2E ASR architecture has attracted lots of attention because it combines the advantages of CTC models and attention models. During training, the CTC objective is attached to the attention-based encoder-decoder model as an auxiliary task. During decoding, the joint CTC/attention decoding approach is adopted in the beam search . However, it is difficult to deploy the online CTC/attention E2E ASR architecture because of global attention mechanisms and CTC prefix scores , which depend on the entire input speech. Our prior work has streamed this architecture from both the model structure and decoding algorithm aspects. On the model structure aspect, we proposed the stable monotonic chunk-wise attention (sMoChA) and monotonic truncated attention (MTA) to stream attention mechanisms, and applied the latency-controlled bidirectional long short-term memory (LC-BLSTM) as the low-latency encoder. On the decoding aspect, we proposed the online joint decoding approach, which includes truncated CTC (T-CTC) prefix scores and dynamic waiting joint decoding (DWJD) algorithm .

Recently, Transformer has gained success in ASR field . Transformer-based models are parallelizable and competitive to recurrent neural networks . However, the vanilla Transformer is inapplicable to online tasks for two reasons: First, the self-attention encoder (SAE) computes the attention weights on the whole input frames; Second, the self-attention decoder (SAD) computes the attention weights on the whole outputs of SAE.

In this paper, we stream the Transformer and integrate it into the CTC/attention E2E ASR architecture. On the SAE aspect, we propose the chunk-SAE which splits the input speech into isolated chunks of fixed length. Inspired by Transformer-XL , we further propose the state reuse chunk-SAE which reuses the stored states of the previous chunks to reduce the computational cost. On the SAD aspect, we propose the MTA based SAD, which performs attention on the truncated historical outputs of SAE. Finally, we propose the Transformer-based online CTC/attention E2E ASR architecture via the online joint decoding approach . Our experiments shows that the proposed online model with a 320 ms latency achieves 23.66% character error rate (CER) on HKUST, with only 0.19% absolute CER degradation compared with the offline baseline.

The rest of this paper is organized as follows. In Section 2, we describe the online CTC/attention E2E architecture proposed in our prior work . In Section 3, we introduce the Transformer architecture. In Section 4, we describe the online Transformer-based CTC/attention architecture. The experiments and conclusions are presented in Sections 5 and 6, respectively.

Online CTC/attention E2E Architecture

In our prior work , we proposed an online hybrid CTC/attention E2E ASR architecture, which consists of the LC-BLSTM encoder, sMoChA and LSTM decoder. During training, we introduce the CTC objective as an auxiliary task, and the loss function is defined by:

where α\alpha is a hyperparameter, Ldec\mathcal{L}_{\rm dec} and Lctc\mathcal{L}_{\rm ctc} are loss functions from the decoder and CTC. During decoding, we adopt the online joint decoding approach, which is defined by:

where Pdec(Y∣X)P_{\rm dec}(Y|X) and Pt-ctc(Y∣X)P_{\rm t\textit{-}ctc}(Y|X) are the probabilities of the hypothesis YY conditioned on input frames XX from the decoder and T-CTC , and Plm(Y)P_{\rm lm}(Y) is the language model probability. The hyperparameters λ\lambda and γ\gamma are tunable. For online decoding, we proposed DWJD algorithm to 1) coordinate the forward propagation in the encoder and the beam search in the decoder; 2) address the unsynchronized predictions of the sMoChA-based decoder and CTC outputs.

MTA , which performs attention on top of the truncated historical encoder outputs, outperforms the sMoChA by exploiting longer history. Formally, we denote qi{\rm\mathbf{q}}_{i} and hj{\rm\mathbf{h}}_{j} as the ii-th decoder state and the jj-th encoder output, respectively. Similar to monotonic chunk-wise attention , MTA defines the probability pi,jp_{i,j} of truncating encoder outputs at hj{\rm\mathbf{h}}_{j} as:

where the matrices W1{\rm\mathbf{W}}_{1}, W2{\rm\mathbf{W}}_{2}, vectors b{\rm\mathbf{b}}, v{\rm\mathbf{v}} and scalars gg, rr are trainable parameters. Then, the attention weight ai,ja_{i,j} is computed by:

where ai,ja_{i,j} indicates the probability of truncating encoder outputs at hj{\rm\mathbf{h}}_{j} and skipping the encoder outputs before hj{\rm\mathbf{h}}_{j}. During decoding, MTA determines a truncation end-point tit_{i} for the ii-th decoder step by:

where ri{\rm\mathbf{r}}_{i} is the letter-wise hidden vector for the ii-th decoder step.

During training, MTA performs attention on the whole encoder outputs:

where TT denotes the number of encoder outputs.

Transformer Architecture

Transformer follows the encoder-decoder architecture using stacked self-attention and position-wise feed-forward layers for both the encoder and decoder. We briefly introduce the Transformer architecture in this section.

Transformer adopts the scaled dot-product attention to map a query and a set of key-value pairs to an output as:

Instead of performing a single attention function, Transformer uses multi-head attention that jointly learns diverse relationships between queries and keys from different representation sub-spaces as follows:

Because Transformer lacks of modeling the sequence order, the work in suggested to use sine and cosine functions of different frequencies to perform the positional encoding.

2 Self-attention encoder (SAE)

The SAE consists of a stack of identical layers, each of which has two sub-layers, i.e. one self-attention layer and one position-wise feed-forward layer. The inputs of the SAE are acoustic frames in ASR tasks. The self-attention layer employs multi-head attention, in which the queries, keys and values are inputs of the previous layer. Besides, the SAE uses residual connections and layer normalization after each sub-layer.

3 Self-attention decoder (SAD)

The SAD also consists of a stack of identical layers, each of which has three sub-layers, i.e. one self-attention layer, one encoder-decoder attention layer and one position-wise feed-forward layer. The inputs of the SAD are embeddings of right-shifted output labels. To prevent the access to the future output labels in the self-attention, the subsequent positions are masked. In the encoder-decoder attention, the queries are current layer inputs while the keys and values are SAE outputs. Besides, the SAD also uses residual connections and layer normalization after each sub-layer.

Transformer-based Online CTC/attention E2E Architecture

In this section, we propose the Transformer-based online E2E model, which consists of the chunk-SAE with or without reusing stored states and MTA based SAD. The Transformer-based online CTC/attention E2E architecture is shown in Fig. 1.

To stream the SAE, we first propose the chunk-SAE, which splits a speech into non-overlapping isolated chunks of NcN_{c} central length. To acquire the contextual information, we splice NlN_{l} left frames before each chunk as historical context and NrN_{r} right frames after it as future context. The spliced frames only act as contexts and give no output. With the predefined parameters NcN_{c}, NlN_{l} and NrN_{r}, the receptive field of each chunk-SAE output is restricted to Nl+Nc+NrN_{l}+N_{c}+N_{r} and the latency of the chunk-SAE is limited to NrN_{r}.

2 State reuse chunk-SAE

In Eq. 12, the function SG(⋅){\rm SG}(\cdot) stands for stop-gradient. Therefore, the complexity of the state reuse chunk-SAE is reduced by a factor of Nl/(Nl+Nc+Nr)N_{l}/(N_{l}+N_{c}+N_{r}).

Moreover, the state reuse chunk-SAE captures long-term dependency beyond the chunks. Suppose the state reuse chunk-SAE consists of LL layers, the receptive field on the left side extends to as far as L⋅NlL\cdot N_{l} frames, which is much broader than that of chunk-SAE.

3 MTA based SAD

To stream the SAD, we propose the MTA based SAD to truncate the receptive field in a monotonic left-to-right way and perform attention on the truncated outputs of SAE. Specifically, we substitute MTA for the encoder-decoder attention in each SAD layer, as shown in Fig. 2. Suppose the representation dimension is dmd_{m}, MTA performs in parallel during training as follows:

MTA learns the appropriate offset for the pre-sigmoid activations in Eq. 14 via the trainable scalar rr. To prevent cumprod(1−P){\rm cumprod}(\mathbf{1}-\mathbf{P}) from vanishing to zeros, we initialize rr to a negative value, e.g. r=−4r=-4 in our experiments. To encourage the discreteness of the truncation probabilities, we simply add zero-mean, unit-variance Gaussian noise ε\varepsilon to the pre-sigmoid activations only during training.

During decoding, we have to compute the elements in Pl ⁣= ⁣{pi,jl}\mathbf{P}^{l}\!=\!\{p_{i,j}^{l}\} row by row, where Pl\mathbf{P}^{l} is the truncation probability matrix in the ll-th layer. we define tilt_{i}^{l} as the truncation end-point belonging to the ll-th layer when predicting the ii-th output label. Then, the end-point is determined by:

Experiments

We evaluated our models using HKUST Mandarin Chinese conversational telephone . The HKUST consists of about 200 hours train set for training and about 5 hours test set. We extracted 4000 utterances from the train set as our development set. To improve the recognition accuracy, we applied the speed perturbation on the rest train set by factors 0.9 and 1.1.

2 Model descriptions

We built all the online models using ESPnet toolkit . For the input, we used 83-dimensional features, including 80-dimensional filter banks, pitch, delta-pitch and Normalized Cross-Correlation Functions. The features were computed with a 25 ms window and shifted every 10 ms. For the output, we adopted a 3655-sized vocabulary set, including 3623 Chinese Mandarin characters, 26 English characters, as well as 6 non-language symbols denoting laughter, noise, vocalized noise, blank, unknown-character and sos/eos.

We used 2-layer convolutional neural networks (CNN) as the front-end. Each CNN layer had 256 filters, each of which has 3×33\times 3 kernel size with 2×22\times 2 stride, and thus the time reduction of the front-end was 1/41/4. The SAE and SAD had 12 and 6 layers, respectively. All sub-layers, as well as embedding layers, produced outputs of dimension 256. In the multi-head attention networks, the head number was 4. In the position-wise feed-forward networks, the inner dimension was 2048. Besides, we trained a 2-layer 1024-dimensional LSTM network on HKUST transcriptions as the external language model and adopted the above 3655-sized vocabulary set.

During training, we used the CTC/attention joint training (α=0.7\alpha=0.7) and the Adam optimizer with Noam learning rate schedule (25000 warm steps), and trained for 30 epochs. To prevent overfitting, we used dropout (dropout rate  ⁣= ⁣0.1\!=\!0.1) in each sub-layer, uniform label smoothing (penalty  ⁣= ⁣0.1\!=\!0.1) in the output layer and the model averaging approach that averages the parameters of models at the last 10 epochs. During decoding, we adopted online joint decoding approach, combining T-CTC prefix scores (λ ⁣= ⁣0.5\lambda\!=\!0.5) and language model scores (γ ⁣= ⁣0.3\gamma\!=\!0.3) to prune the hypotheses, and the beam size was 10.

3 Chunk-SAE with or without reusing states

In Table 1, we compared the speed and performance of the chunk-SAE with or without reusing states. The context configuration remained the same for online models during the comparison, i.e. Nl ⁣= ⁣Nc ⁣= ⁣Nr ⁣= ⁣64N_{l}\!=\!N_{c}\!=\!N_{r}\!=\!64. Firstly, we measured the speed of various encoders during decoding using a sever with Intel(R) Xeon(R) Silver 4114 CPU, 2.20GHz. For the clear comparison, we set the speed of chunk-SAE to 1.01.0 and give the speed ratio of other encoders. In lines 1 and 2 of Table 1, the chunk-SAE was slower than the SAE due to the redundant computation of the historical and future context. In lines 2 and 3 of Table 1, we observed that the state reuse chunk-SAE was 1.5x faster than the chunk-SAE, which is consistent with the theoretical analysis in Section 4.2. In addition to the faster speed, the state reuse chunk-SAE outperformed the chunk-SAE by 1.53%1.53\% and 0.38%0.38\% relative CERs reduction on HKUST development and test set, respectively. Because of the faster speed and better performance, we employed the state reuse chunk-SAE in our subsequent experiments.

4 Context investigation

in Table 2, we investigated our online model performance varying the historical, central and future context lengths. Firstly, comparing lines 2-4 in Table 2, we can see that the future context brought more improvement than the historical context, which indicates that the future context is more crucial to the performance of our online models. Secondly, comparing lines 5-7 in Table 2, we found that it was effective to increase the length of the historical context when we intended to reduce the latency of the state reuse chunk-SAE and maintain the recognition accuracy at the same time. Thirdly, comparing lines 7 and 8 in Table 2, we found that the CER reduced when we increased the length of central context.

Finally, our best online model achieved a 23.65%23.65\% CER, with a 640 ms latency and a 0.18%0.18\% absolute CER degradation compared with the offline baseline in line 1 of Table 2. In Table 3, we also compared our Transformer-based online CTC/attention model with other published ASR models. For a fair comparison, the latency of the online E2E models listed in Table 3 is 320 ms. These models were trained on HKUST with speed perturb except online Self-attention Aligner model.

Conclusion

In this paper, we propose the Transformer-based online E2E ASR model, which consists of the state reuse chunk-SAE and MTA based SAD, and integrate the proposed Transformer-based online E2E ASR model into the CTC/attention ASR architecture. Compared with the simple chunk-SAE, the state reuse chunk-SAE performs better and requires less computational cost, because it has broader historical context via storing the states in previous chunks. Compared with the SAD, the MTA based SAD truncates the SAE outputs in a monotonic left-to-right way and performs attention on the truncated SAE outputs, making it applicable to online recognition. We evaluate the proposed Transformer-based online CTC/attention E2E models on HKUST and achieves a 23.66%23.66\% CER with a 320 ms latency, which outperforms our prior LSTM-based online E2E models. In future, we plan to adopt teacher-student learning approach to further reduce the model latency.

References