QuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions

Samuel Kriman, Stanislav Beliaev, Boris Ginsburg, Jocelyn Huang, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, Yang Zhang

Introduction

In the last few years, end-to-end (E2E) neural networks (NN) have achieved new state-of-the-art (SOTA) results on many automatic speech recognition (ASR) tasks. Such models replace the traditional multi-component ASR system with a single, end-to-end trained NN which directly predicts character sequences and therefore greatly simplify training, fine-tuning and inference. The latest E2E models also have very good accuracy, but this often comes at the cost of increasingly large models with high computational and memory requirements.

The motivation of this work is to build an ASR model that achieves SOTA-level accuracy, while utilizing significantly fewer parameters and less compute power. Smaller models offer multiple advantages: (1) they are faster to train, (2) they are more feasible to deploy on hardware with limited compute and memory, and (3) they have higher inference throughput.

We achieve this goal by building a very deep NN with 1D time-channel separable convolutions. This new network reaches near-SOTA word error rate (WER) on LibriSpeech (see Table 4) and WSJ (see Table 7) datasets with fewer than 20 million parameters, compared to previous end-to-end ASR designs which typically have over 100 million parameters. We have released the source code and pre-trained models in the NeMo toolkit . https://github.com/NVIDIA/NeMo

Related Work

There has been a lot of work done in exploring compact network architectures and on investigating the trade-off between accuracy and size of neural networks, such as SqueezeNet , ShuffleNet , and EfficientNet . Our approach is directly related to MobileNets and Xception , which uses depthwise separable convolutions . Each depthwise separable convolution module is made up of two parts: a depthwise convolutional layer and a pointwise convolutional layer. Depthwise convolutions apply a single filter per input channel (input depth). Pointwise convolutions are 1×11\times 1 convolutions, used to create a linear combination of the outputs of the depthwise layer. BatchNorm and ReLU are applied to the outputs of both layers.

Hannun et al applied a similar approach to ASR. They introduced an encoder-decoder model with time-depth separable (TDS) convolutions. The TDS model operates on data in time-frequency-channels (T×w×cT\times w\times c) format, where TT is the number of time-steps, ww is the input width and cc is the number of channels. The basic TDS block is composed of a 2D convolutional block with k×1k\times 1 convolutions over (T×w)(T\times w), and a fully-connected block, consisting of two 1×11\times 1 pointwise convolutions operating on (w⋅c)(w\cdot c) channels interleaved with layer-norm layers. In contrast, in our work we operate on data in time-channel format (T×cT\times c) and completely decouple the time and channel-wise parts of convolution. TDS block has k×c2+2×(w⋅c)2k\times c^{2}+2\times(w\cdot c)^{2} parameters, while QuartzNet model has k×c+c2k\times c+c^{2} parameters, which allows for a dramatic reduction in model size while still achieving good WER.

Another very small ASR model was introduced by Han et al , which uses multiple parallel streams of self-attention with dilated, factorized, although not separable, 1D convolutions. The parallel streams capture multiple resolutions of speech frames from the input by using different dilation rates per stream, and the results of the individual streams are concatenated into a final embedding. The best model has five streams with dilation rates 1-2-3-4-5.

Model architecture

QuartzNet’s design is based on the Jasper architecture, which is a convolutional model trained with Connectionist Temporal Classification (CTC) loss . The main novelty in QuartzNet’s architecture is that we replaced the 1D convolutions with 1D time-channel separable convolutions, an implementation of depthwise separable convolutions. 1D time-channel separable convolutions can be separated into a 1D depthwise convolutional layer with kernel length KK that operates on each channel individually but across K\boldsymbol{K} time frames and a pointwise convolutional layer that operates on each time frame independently but across all channels.

QuartzNet models have the following structure: they start with a 1D convolutional layer C1C_{1} followed by a sequence of blocks. Each block BiB_{i} is repeated SiS_{i} times and has residual connections between blocks. Each block BiB_{i} consists of the same base modules repeated RiR_{i} times and contains four layers: 1) KK-sized depthwise convolutional layer with coutc_{out} channels, 2) a pointwise convolution, 3) a normalization layer, and 4) ReLU. The last part of the model consists of three additional convolutional layers (C2,C3,C4C_{2},C_{3},C_{4}). The C1C_{1} layer has a stride of 2, and C4C_{4} layer has a dilation of 2.

Table 1 describes the QuartzNet-5x5, 10x5 and 15x5 models. There are five unique blocks across these models: B1B_{1} - B5B_{5}. The different models repeat the blocks a different number of times, represented by SiS_{i}. QuartzNet-5x5 (B1−B2−B3−B4−B5B_{1}-B_{2}-B_{3}-B_{4}-B_{5}) has each group of blocks repeated 1 time, QuartzNet-10x5 (B1−B1−B2−B2−...−B5−B5B_{1}-B_{1}-B_{2}-B_{2}-...-B_{5}-B_{5}) - repeated 2 times, and QuartzNet-15x5 (B1−B1−B1−...−B5−B5−B5B_{1}-B_{1}-B_{1}-...-B_{5}-B_{5}-B_{5}) - repeated 3 times.

A regular 1D convolutional layer with kernel size KK, cinc_{in} input channels, and coutc_{out} output channels has K×cin×coutK\times c_{in}\times c_{out} weights. The time-channel separable convolutions use K×cin+cin×coutK\times c_{in}+c_{in}\times c_{out} weights split into K×cinK\times c_{in} weights for the depthwise layer and cin×coutc_{in}\times c_{out} for the pointwise layer.

The depthwise convolution is applied independently for each channel, so it contributes a relatively small portion of the total number of weights. This allows us to use much wider kernels, roughly 3 times larger than kernels used in wav2letter or Jasper models. We experimented with four types of normalization: batch normalization , layer normalization , instance normalization , and group normalization , and found that models with batch normalization have most stable training and give the best WER.

2 Pointwise convolutions with groups

The total number of weights for a time-channel separable convolution block is K×cin+cin×coutK\times c_{in}+c_{in}\times c_{out} weights. Since KK is generally several times smaller than coutc_{out}, most weights are concentrated in the pointwise convolution part. In order to further reduce the number of parameters, we explore using group convolutions for this layer. We also added a group shuffle layer to increase cross-group interchange .

Using groups allows us to significantly reduce the number of weights at the cost of some accuracy. Table 3 shows the trade-off between accuracy and number of parameters for group sizes one, two, and four, evaluated on LibriSpeech.

Experiments

We evaluate QuartzNet’s performance on LibriSpeech and WSJ datasets. We additionally experiment with a transfer learning showcasing how a QuartzNet model trained with LibriSpeech and Common Voice can be fine-tuned on a smaller amount of audio data, the WSJ dataset, to achieve better performance than training from scratch.

Our best results on the LibriSpeech dataset are achieved with the QuartzNet-15x5 model, consisting of 15 blocks with 5 convolutional modules per block (see Table 1). By combining our network with independently trained language models (i.e., n-gram language models and Transformer-XL (T-XL) ) we got WER comparable to the current SOTA.

The model with time-channel separable convolutions is much smaller than a model with regular convolutions and is less prone to over-fitting, so we use only data augmentation and weight decay for regularization during training. We experimented with SpecAugment , SpecCutout, and speed perturbation . We achieved the best results with 10% speed perturbation combined with Cutout which randomly cuts small rectangles out of the spectrogram. The models are trained using NovoGrad optimizer with a cosine annealing learning rate policy. We also found that learning rate warmup helps stabilize early training.

The training of the 15x5 model for 400 epochs took ≈5\approx 5 days on one DGX1 server with 8 Tesla V100 GPUs with a batch size of 32 per GPU. In order to decrease the memory footprint as well as training time, we used mixed-precision training . We reduced the training time to just over four hours by scaling training to SuperPod with 32 DGX2 nodes with larger number of epochs and with an increased global batch of 16K (see Table 5).Training even longer (3000 epochs) improved greedy WER on test-clean to 3.87% and on test-other to 10.61%.

2 Wall Street Journal

We trained a smaller QuartzNet-5x3 model on the open vocabulary task of the Wall Street Journal dataset . We used train-si284 set for training, nov93-dev for validation, and nov92-eval for testing. The QuartzNet-5x3 model (see Table 6) was trained for 1200 epochs with batch size 32 per GPU, data augmentation (10% speed perturbation, SpecCutout) and dropout of 0.2 using NovoGrad optimizer (β1=0.95\beta_{1}=0.95, β2=0.5\beta_{2}=0.5) with 1000 steps of learning rate warmup, a learning rate of 0.05, and weight decay 0.001.

We used 2 external language models during inference: 4-gram (beam size=2048, alpha=3.5, beta=1.5) and Transformer-XL (T-XL). Both language models were constructed using only the official LM data of WSJ.

We used following end-to-end models trained on standard speech featuresNote, that wav2letter++ with trainable front-end and convLM has even better WER: 6.8%\% for nov93-test, and 3.5%\% for nov92-dev. Here, we consider only models with a standard mel-filterbanks front-end. for comparison: 1) RNN-CTC - CTC model with 5 bidirectional LSTM layers, 500 cells in each layer; 2) ResCNN-LAS : Listen-Attend-Spell model with deep residual convLSTM encoder and LSTM decoder + label smoothing; 3) Wav2Letter++ - CTC model with 1D convolutional layers and instance norm.

3 Transfer Learning

As our model is smaller than other models, we were interested in how well it could learn to generalize to data from various sources, especially if the amount of target speech is much smaller than the training data. Our setup consists of training QuartzNet 15x5 on a combination of LibriSpeech and Mozilla’s Common Voice We used the validated set of Common Voice, ver. en_1087h_2019-06-12. datasets, and then fine-tuning this trained model on the 80 hour WSJ dataset. Table 8 shows the WER achieved on LibriSpeech prior to fine-tuning and the result on WSJ after fine-tuning.

Conclusions and future directions

We introduced a new end-to-end speech recognition model, based on deep neural network with 1D time-channel separable convolutional layers. The model showed close to state-of-the art performance on Wall Street Journal and on LibriSpeech while being significantly smaller than all other end-to-end systems with similar accuracy. The small model footprint opens new possibility for speech recognition on mobile and embedded devices.

This work described a CTC-based model, but we are exploring models where the QuartzNet encoder is combined with attention-based decoders.

References