Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict

Yosuke Higuchi, Shinji Watanabe, Nanxin Chen, Tetsuji Ogawa, Tetsunori Kobayashi

Introduction

Owing to the rapid development of neural sequence-to-sequence modeling , deep neural network (DNN)-based end-to-end automatic speech recognition (ASR) systems have become almost as effective as the traditional hidden Markov model-based systems . Various models and approaches have been proposed for improving the performance of the autoregressive (AR) end-to-end ASR model with the encoder-decoder architecture based on recurrent neural networks (RNNs) and Transformers .

Contrary to the autoregressive framework, non-autoregressive (NAR) sequence generation has attracted attention, including the revisitation of connectionist temporal classification (CTC) and the growing interest for non-autoregressive Transformer (NAT) . While the autoregressive model requires LL iterations to generate an LL-length target sequence, a non-autoregressive model costs a constant number of iterations K(≪L)K(\ll L), independent on the length of the target sequence. Despite the limitation in this decoding iteration, some recent studies in neural machine translation have successfully shown the effectiveness of the non-autoregressive models, performing comparable results to the autoregressive models. Different types of non-autoregressive models have been proposed based on the iterative refinement decoding , insert or edit-based sequence generation , masked language model objective , and generative flow .

Some attempts have also been made to realize the non-autoregressive model in speech recognition. CTC introduces a frame-wise latent alignment to represent the alignment between the input speech frames and the output tokens . While CTC makes use of dynamic programming to efficiently calculate the most probable alignment, the strong conditional independence assumption between output tokens results in poor performance compared to the autoregressive models . On the other hand, trains a Transformer encoder-decoder in a mask-predict manner : target tokens are randomly masked and predicted conditioning on the unmasked tokens and the input speech. To generate the output sequence in parallel during inference, the target sequence is initialized as all masked tokens and the output length is predicted by finding the position of the end-of-sequence token. However with this prediction of the output length, the model is known to be vulnerable to the output sequence with a long length. At the beginning of the decoding, the model is likely to make more mistakes in predicting long masked sequence, propagating the error to the later decoding steps. proposes Imputer, which performs the mask prediction in CTC’s latent alignments to get rid of the output length prediction. However, unlike the mask-predict, Imputer requires more calculations in each interaction, which is proportional to the square of the input length T(≫L)T(\gg L) in the self-attention layer, and the total computational cost can be very large.

Our work aims to obtain a non-autoregressive end-to-end ASR model, which generates the sequence in token-level with low computational cost. The proposed Mask CTC framework trains a Transformer encoder-decoder model with both CTC and mask-predict objectives. During inference, the target sequence is initialized with the greedy CTC outputs and low-confidence tokens are masked based on the CTC probabilities. The masked low-confidence tokens are predicted conditioning on the high-confidence tokens not only in the past but also in the future context. The advantages of Mask CTC are summarized as follows.

No requirement for output length prediction: Predicting the output token length from input speech is rather challenging because the length of the input utterances varies greatly depending on the speaking rate or the duration of silence. By initializing the target sequence with the CTC outputs, Mask CTC does not have to care about predicting the output length at the beginning of the decoding.

Accurate and fast decoding: We observed that the results of CTC outputs themselves are quite accurate. Mask CTC does not only retain the correct tokens in the CTC outputs but also recovers the output errors by considering the entire context. Token-level iterative decoding with a small number of masks makes the model well-suited for the usage in real scenarios.

Mask CTC framework

The following subsections first explain a conventional autoregressive framework based on attention-based encoder-decoder and CTC. Then a non-autoregressive model trained with mask prediction is explained and finally, the proposed Mask CTC decoding method is introduced.

Attention-based encoder-decoder models the joint probability of YY given XX by factorizing the probability based on the probabilistic left-to-right chain rule as follows:

The model estimates the output token yly_{l} at each time-step conditioning on previously generated tokens y<ly_{<l} in an autoregressive manner. In general, during training, the ground truth tokens are used for the history tokens y<ly_{<l} and during inference, the predicted tokens are used.

2 Connectionist temporal classification

CTC predicts a frame-level alignment between the input sequence XX and the output sequence YY by introducing a special token. The alignment A={at∈V∪{<blank>}∣t=1,...,T}A=\{a_{t}\in\mathcal{V}\cup\{\texttt{<blank>}\}|t=1,...,T\} is predicted with the conditional independence assumption between the output tokens as follows:

Considering the probability distribution over all possible alignments, CTC models the joint probability of YY given XX as follows:

where β−1(Y)\beta^{-1}(Y) returns all possible alignments compatible with YY. The summation of the probabilities for all of the alignments can be computed efficiently by using dynamic programming.

To achieve robust alignment training and fast convergence, an end-to-end ASR model based on an attention-based encoder-decoder framework is trained with CTC . The objective of the autoregressive joint CTC-attention model is defined as follows by combining Eq. (1) and Eq. (3):

3 Joint CTC-CMLM non-autoregressive ASR

Mask CTC adopts non-autoregressive speech recognition based on a conditional masked language model (CMLM) , where the model is trained to predict masked tokens in the target sequence Note that CMLM is used as an ASR decoder network conditioned on the encoder output as well, and it is different from an external language model often used in shallow fusion during decoding.. Taking advantages of Transformer’s parallel computation , CMLM can predict any arbitrary subset of masked tokens in the target sequence by attending to the entire sequence including tokens in the past and the future.

We observed that applying the original CMLM to non-autoregressive speech recognition shows poor performance, having the problem of skipping and repeating the output tokens. To deal with this, we found that jointly training with CTC similar to provides the model with absolute positional information (conditional independence) explicitly and improves the model performance reasonably well. With the CTC objective from Eq. (3) and Eq. (5), the objective of joint CTC-CMLM training for non-autoregressive ASR model is defined as follows:

4 Mask CTC decoding

Non-autoregressive models must know the length of the output sequence to predict the entire sequence in parallel. For example, in the beginning of the CMLM decoding, the output length must be given to initialize the target sequence with the masked tokens. To deal with this problem, in machine translation, the output length is predicted by training a fertility model or introducing a special token in the encoder . In speech recognition, however, due to the different characteristics between the input acoustic signals and the output linguistic symbols, it appeared that predicting the output length is rather challenging, e.g., the length of the input utterances of the same transcription varies greatly depending on the speaking rate or the duration of silence. simply makes the decoder to predict the position of token to deal with the output length. However, they analyzed that this prediction is vulnerable to the output sequence with a long length because the model is likely to make more mistakes in predicting a long masked sequence and the error is propagated to the later decoding, which degrades the recognition performance. To compensate this problem, they use beam search with CTC and a language model to obtain the reasonable performance, which leads to a slow down of the overall decoding speed, making the advantage of non-autoregressive framework less effective.

To tackle this problem regarding the initialization of the target sequence, we consider using the CTC outputs as the initial sequence for decoding. Figure 1 shows the decoding of CTC Mask based on the inference of CTC. CTC outputs are first obtained through a single calculation of the encoder and the decoder works as to refine the CTC outputs by attending to the whole sequence.

In this work, we use “greedy” result of CTC Y^={y^l∈β(A)∣l=1,...,L′}\hat{Y}=\{\hat{y}_{l}\in\beta(A)|l=1,...,L^{\prime}\}, which is obtained without using prefix search , to keep an inference algorithm non-autoregressive. The errors caused by the conditional independence assumption are expected to be corrected using the CMLM decoder. The posterior probability of y^l\hat{y}_{l} is approximately calculated by using the frame-level CTC probabilities as follows:

where Al={aj}jA_{l}=\{a_{j}\}_{j} is the consecutive same alignments that corresponds to the aggregated token y^l\hat{y}_{l}. Then, a part of Y^\hat{Y} is masked-out based on a confidence using the probability P^\hat{P} as follows:

Top CC masked tokens with the highest probabilities are predicted in each iteration. By defining C=[L/K]C=[L/K], the number of total decoding iterations can be controlled in a constant KK iterations.

With this proposed non-autoregressive training and decoding with Mask CTC, the model does not have to take care about predicting the output length. Moreover, decoding by refining CTC outputs with the mask prediction is expected to compensate the errors come from the conditional independence assumption.

Experiments

To evaluate the effectiveness Mask CTC, we conducted speech recognition experiments to compare different end-to-end ASR models using ESPnet . The performance of the models was evaluated based on character error rates (CERs) or word error rates (WERs) without relying on external language models.

The experiments were carried out using three tasks with different languages and amounts of training data: the 81 hours Wall Street Journal (WSJ) in English , the 581 hours Corpus of Spontaneous Japanese (CSJ) in Japanese and the 16 hours Voxforge in Italian . For the network inputs, we used 80 mel-scale filterbank coefficients with three-dimensional pitch features and applied SpecAugment during model training. For the tokenization of the target, we used characters: Latin alphabets for English and Italian, and Japanese syllable characters (Kana) and Chinese characters (Kanji) for Japanese.

2 Experimental setup

3 Evaluated models

CTC-attention: An autoregressive model trained with the joint CTC-attention objective as in Eq. (4). During inference, the joint CTC-attention decoding is applied with beam search .

CTC: A non-autoregressive model simply trained with the CTC objective.

4 Results

Table 1 shows the results for WSJ based on WERs and real time factors (RTFs) that were measured for decoding eval92 with Intel(R) Core(TM), i9-7980XE, 2.60GHz. By comparing the results for non-autoregressive models, we can see that the greedy CTC outputs of Mask CTC outperformed the simple CTC model by training with the mask-predict objective. By applying the refinement based on the proposed CTC masking, the model performance was steadily improved. The performance was further improved by increasing the number of decoding iterations and it resulted in the best performance with #mask iterations, which means one mask is predicted in each iteration. The results of Mask CTC are reasonable comparing to the results of prior work . Our models also approached the results of autoregressive models from the initial CTC result. In terms of the decoding speed measured in RTF, Mask CTC is, at most, 116 times faster than the autoregressive models. Since most of the CTC outputs are fairly accurate and the number of masks are quite small, there was not so much degradation in the speed as the number of the decoding iterations was increased.

Figure 2 shows an example decoding process of a sample in the WSJ evaluation set. Here, we can see that the CTC outputs include errors mainly coming from substitution errors due to the incomplete word spelling. By applying Mask CTC decoding, the spelling errors were successfully recovered by considering the conditional dependence between characters in word-level. However, as can be seen in the error for “sifood,” Mask CTC cannot recover errors derived from character-level insertion or deletion errors because the length allocated to each word is fixed by the CTC outputs.

Table 2 shows WERs for Voxforge. Mask CTC yielded better scores than the standard CTC model as the similar results to WSJ, demonstrating that our model can be adopted to other languages with a relatively small amount of training data.

Table 3 shows character error rates (CERs) and sentence error rates (SERs) for CSJ. While Mask CTC performed quite close or even better CERs than the autoregressive model, the results showed a little improvement from the simple CTC model, compared to the results of the aforementioned tasks. Since Japanese includes a large number of characters and the characters themselves often form a certain word, the simple CTC model seemed to be dealing with the short dependence between the characters reasonably well, performing almost the same scores without applying Mask CTC. However, when we look at the results in sentence-level, we observed some clear improvements for all of the evaluation sets, again showing that our model effectively recovers the CTC errors by considering the conditional dependence.

These experimental results on different tasks indicate that Mask CTC framework is especially effective on languages having tokens with a small unit (i.e., Latin alphabet and other phonemic scripts). It is our future work for investigating the effectiveness when we use byte pair encodings (BPEs) for the languages with such a small unit.

Conclusions

This paper proposed Mask CTC, a novel non-autoregressive end-to-end speech recognition framework, which generates a sequence by refining the CTC outputs based on mask prediction. During inference, the target sequence was initialized with the greedy CTC outputs and low-confidence masked tokens were iteratively refined conditioning on the other unmasked tokens and input speech features. The experimental comparisons demonstrated that Mask CTC outperformed the standard CTC model while maintaining the decoding speed fast. Mask CTC approached the results of autoregressive models; especially for CSJ, they were comparable or even better. Our future plan is to reduce the gap of masking strategies between training using random masking and inference using CTC outputs. Furthermore, we plan to explore the integration of external language models (e.g., BERT ) in Mask CTC framework.

References