ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context

Wei Han, Zhengdong Zhang, Yu Zhang, Jiahui Yu, Chung-Cheng Chiu, James Qin, Anmol Gulati, Ruoming Pang, Yonghui Wu

Introduction

Convolution Neural Network (CNN) based models for end-to-end (E2E) speech recognition is attracting an increasing amount of attention . Among them, the Jasper model recently achieves close to the state-of-the-art word error rate (WER) 2.95% on LibriSpeech test-clean with an external neural language model. The main feature of the Jasper model is a deep convolution based encoder with stacked layers of 1D convolutions and skip connections. Depthwise separable convolutions have been utilized to further increase the speed and accuracy of CNN models . The key advantage of a CNN based model is its parameter efficiency; however, the WER achieved by the best CNN model, QuartzNet , is still behind the RNN/transformer based models .

A major difference between the RNN/Transformer based models and a CNN model is the length of the context. In a bidirectional RNN model, a cell in theory has access to the information of the whole sequence; in a Transformer model, the attention mechanism explicitly allows the nodes at two distant time stamps to attend each other. However, a naive convolution with a limited kernel size only covers a small window in the time domain; hence the context is small and the global information is not incorporated. In this paper, we argue that the lack of global context is the main cause of the gap of WER between the CNN based ASR model and the RNN/Transformer based models.

To enhance the global context in the CNN model, we draw inspirations from the squeeze-and-excitation (SE) layer introduced in , and propose a novel CNN model for ASR, which we call ContextNet. An SE layer squeezes a sequence of local feature vectors into a single global context vector, broadcasts this context back to each local feature vector, and merges the two via multiplications. When we place an SE layer after a naive convolution layer, we grant the convolution output the access to global information. Empirically, we observe that adding squeeze-and-excitation layers to ContextNet introduces the most reduction in the WER on LibriSpeech test-other.

Previous works on hybrid ASR have successfully introduced the context to acoustic models by either stacking a large number of layers, or having a separately trained global vector to represent the speaker and the environment information . In , SE has been adopted to RNN for unsupervised adaptation. In this paper, we show that SE can also be effective for CNN encoders.

The architecture of ContextNet is also inspired by the design choices of QuartzNet , such as the usage of depthwise separable 1D convolution in the encoder. However, there are some key differences in the architectures in addition to the incorporation of the SE layer. For instance, we use a RNN-T decoder instead of the CTC decoder . Moreover, we use the Swish activation function , which contributes a slight but consistent reduction in WER. Overall, ContextNet achieves the WER of 1.9%/4.1% on LibriSpeech test-clean/test-other. This is a big improvement over previous CNN based architectures such as QuartzNet, and it outperforms transformer and LSTM based models .

This paper also studies how to reduce the computation cost of ContextNet for faster training and inference. First, we adopt a progressive downsampling scheme that is commonly used in vision models. Specifically, we progressively reduce the length of the encoded sequence eight times, significantly lower the computation while maintaining the encoder’s representation power and the overall model accuracy. As a benefit, this downsampling scheme allows us to reduce the kernel size of all the convolution layers to five without significantly reducing the effective receptive field of an encoder output node.

We can scale ContextNet by globally changing the number of channels in convolutional filters. Figure 1 illustrates the trade-off of ContextNet between model size and WER, as well as its comparison against other methods. Clearly, our scaled model achieves the best trade-offs among all.

In summary, the main contributions of this paper are: (1) an improved CNN architecture with global context for ASR, (2) a progressive downsampling and model scaling scheme to achieve superior accuracy and model size trade-off.

Model

This section introduces the architecture details of ContextNet. Section 2.1 discusses the high-level design of ContextNet. Then Section 2.2 introduces our convolutional encoder, and discusses how we progressively reduce the temporal length of the input utterance in the network to reduce the computation while maintaining the accuracy of the model.

Our network is based on the RNN-Transducer framework . The network contains three components: audio encoder on the input utterance, label encoder on the input label, and a joint network to combine the two and decode. We directly use the LSTM based label encoder and the joint network from , but propose a new CNN based audio encoder.

2 Encoder Design

where ∘\circ represents element-wise multiplication, W1,W2W_{1},W_{2} are weight matrics, and b1,b2b_{1},b_{2} are bias vectors.

2.2 Depthwise separable convolution

For simplicity, we use the same kernel size on all depthwise convolution layers in the network.

2.3 Swish activation function

where β=1\beta=1 for all our experiments. We’ve observed that the swish function works consistently better than ReLU.

2.4 Convolution block

2.5 Progressive downsampling

We use strided convolution for temporal downsampling. More downsampling layers reduces computation cost, but excessive downsampling in the encoder may negatively impact the decoder. Empirically, we find that a progressive 8×8\times downsampling scheme achieves a good trade-off between speed and accuracy. These trade-offs are discussed in Section 3.3.

2.6 Configuration details of ContextNet

ContextNet has 2323 convolution blocks C0,…,C22C_{0},\ldots,C_{22}. All convolution blocks have five layers of convolution, except C0C_{0} and C22C_{22}, which only have one layer of convolution each. Table 1 summarizes the architecture details. Note that a global parameter α\alpha controls the scaling of our model. Increasing α\alpha when α>1\alpha>1 increases the number of channels of the convolutions, giving the model more representation power with a larger model size.

Experiments

We conduct experiments on the Librispeech dataset which consists of 970 hours of labeled speech and an additional text only corpus for building language model. We extract 80 dimensional filterbanks features using a 25ms window with a stride of 10ms.

We use SpecAugment with mask parameter (F=27F=27), and ten time masks with maximum time-mask ratio (pS=0.05p_{S}=0.05), where the maximum size of the time mask is set to pSp_{S} times the length of the utterance. Time warping is not used. We use a 3-layer LSTM LM with width 4096 trained on the LibriSpeech langauge model corpus with the LibriSpeech960h transcripts added, tokenized with the 1k WPM built from LibriSpeech 960h. The LM has word-level perplexity 63.9 on the dev-set transcripts. The LM weight λ\lambda for shallow fusion is tuned on the dev-set via grid search. All models are implemented with Lingvo toolkit .

We evaluate three different configurations of ContextNet on LibriSpeech. The models are all based on Table 1, but differ in the network width, α\alpha; hence, they differ in model size. Specifically, we choose α\alpha in {0.5,1,2}\{0.5,1,2\} for the small, medium and large ContextNet. We also build our own LSTM baseline as a reference.

Table 2 summarizes the evaluation results as well as the comparisons with a few previously published systems. The results suggest improvements of ContextNet over previously published systems. Our medium model, ContextNet(M), only has 31M31M parameters and achieves similar WER compared with much larger systems . The large model, ContextNet(L), outperforms the previous SOTA by 13% relatively on test-clean and 18% relatively on test-other. Our scaled-down model, ContextNet(S), also shows an improvement to previous systems of similar size , with or without a language model.

2 Effect of Context Size

To validate the effectiveness of adding global context to the CNN model for ASR, we perform an ablation study on how the squeeze-and-excitation module affects the WER on LibriSpeech test-clean/test-other. ContextNet in Table 1 with all squeeze-and-excitation modules removed and α=1.25\alpha=1.25 serves as the baseline of zero context.

The vanilla squeeze-and-excitation module uses the whole utterance as context. To investigate the effect of different context sizes, we replace the global average pooling operator of the squeeze-and-excitation module by a stride-one pooling operator where the context can be controlled by the size of the pooling window. In this study, we compare the window size of 256256, 512512 and 10241024 on all convolutional blocks.

As illustrated in Table 3, the SE module provides major improvement over the baseline. In addition, the benefit becomes greater as the length of the context window increases. This is consistent with the observation in a similar study of SE on image classification models .

3 Depth, Width, Kernel Size and Downsampling

Depth: We perform a sweeping on the number of convolutional blocks and our best configuration is in Table 1. We find that with this configuration, we can train a model in a day with stable convergence.

Width: We globally scale the width of the network (i.e., the number of channels) on all encoder layers and study how it impacts the model performance. Specifically, we take the ContextNet model from Table 1, sweep α\alpha, and report the model size and the WER on LibriSpeech. Table 5 summarizes the result; it demonstrates the good trade-off between model size and WER of ContextNet.

Downsampling and kernel size: Table 4 summarizes the FLOPS and WER on LibriSpeech with various choices of downsampling and fileter size. We use the same model with only one downsampling layer added to C3C_{3} as the baseline; hence the baseline only does 2×2\times temporal reduction. We sweep the kernel size in {3,5,11,21}\{3,5,11,21\}, each kernel size is applied to all the depthwise convolution layers. The results suggest that progressive downsampling introduces significant saving in the number of FLOPS. Moreover, it actually benefits the accuracy of the model slightly. In addition, with progressive downsampling, increasing the kernel size decreases the WER of the model.

4 Large Scale Experiments

Finally, we show that the proposed architecture is also effective on large scale datasets. We use a experiment setup similar to , where the training set has public Youtube videos with semi-supervised transcripts generated by the approach in . We evaluate on 117 videos with a total duration of 24.12 hours. This test set has diverse and challenging acoustic environments Reproduced results. The train and eval set has been changed recently so the numbers in Table 6 are different from reported in .. Table 6 summarizes the result. We can see that ContextNet outperforms the previous best architecture from , which is a combination of convolution and bidirectional LSTM, by 12% relatively with fewer parameters and FLOPS.

Conclusion

In this work, we proposed and evaluated a CNN based architecture for end-to-end speech recognition. A couple of modeling choices are discussed and compared. This model achieves a better accuracy on the LibriSpeech benchmark with much fewer parameters compared to previously published CNN models. The proposed architecture can easily be used to search for small ASR models by limiting the width of the network. Initial study on a much larger and more challenging dataset also confirms our findings.

References