Letter-Based Speech Recognition with Gated ConvNets

Vitaliy Liptchinsky, Gabriel Synnaeve, Ronan Collobert

Introduction

State of the art speech recognition systems leverage pronunciation models as well as speaker adaptation techniques involving speaker-specific features. These systems rely on lexicon dictionaries, which decompose words into one or more sequences of phones. Phones themselves are decomposed into smaller sub-word units, called senones. Senones are carefully selected through a procedure involving a phonetic-context-based decision tree built from another GMM/HMM system. In the recent literature, “end-to-end” speech systems attempt to break away from these hardcoded a-priori, the underlying assumption being that with enough data pronunciations should be implicitly inferred by the model, and speaker robustness should be also achieved. A number of works have thus naturally proposed ways how to learn to map audio sequences directly to their corresponding letter sequences. Recurrent models, structured-output learning or combination of both are the main contenders.

In this paper, we show that simple convolutional neural networks (CNNs) coupled with structured-output learning can outperform existing letter-based solutions. Our CNNs employ Gated Linear Units (GLU). Gated ConvNets have been shown to reduce the vanishing gradient problem, as they provide a linear path for the gradients while retaining non-linear capabilities, leading to state of the art performance both in natural language modeling and machine translation tasks (Dauphin et al., 2017; Gehring et al., 2017). We train our system with a structured-output learning approach, either with CTC (Graves et al., 2006) or ASG (Collobert et al., 2016). Coupled with a custom-made simple beam-search decoder, we exhibit word error rate (WER) performance matching the best existing letter-based systems, both for the WSJ and LibriSpeech datasets (Panayotov et al., 2015). While phone-based systems still lead on WSJ (8181h of labeled data), our system is competitive with the existing state of the art systems on LibriSpeech (960960h).

The rest of the paper is structured as follows: the next section goes over the history of the work in the automatic speech recognition area. We then detail the convolutional networks used for acoustic modeling, along with the structured-output learning and decoding approaches. The last section shows experimental results on WSJ and LibriSpeech.

Background

The historic pipeline for speech recognition requires first training an HMM/GMM model to force align the units on which the final acoustic model operates (most often context-dependent phone or senone states) Woodland and Young (1993). The performance improvements brought by deep neural networks (DNNs) (Mohamed et al., 2012; Hinton et al., 2012) and convolutional neural networks (CNNs) (Sercu et al., 2016; Soltau et al., 2014) for acoustic modeling only extend this training pipeline. Current state of the art models on LibriSpeech also employ this approach (Panayotov et al., 2015; Peddinti et al., 2015b), with an additional step of speaker adaptation (Saon et al., 2013; Peddinti et al., 2015a). Departing from this historic pipeline, Senior et al. (2014) proposed GMM-free training, but the approach still requires to generate a forced alignment. Recently, maximum mutual information (MMI) estimation (Bahl et al., 1986) was used to train neural network acoustic models (Povey et al., 2016). The MMI criterion (Bahl et al., 1986) maximizes the mutual information between the acoustic sequence and word sequences or the Minimum Bayes Risk (MBR) criterion (Gibson and Hain, 2006), and belongs to segmental discriminative training criterions, although compatible with generative models.

Even though connectionist approaches (Lee et al., 1995; LeCun et al., 1995) long coexisted with HMM-based approaches, they had a recent resurgence. A modern work that directly cut ties with the HMM/GMM pipeline used a recurrent neural network (RNN) (Graves et al., 2013) for phoneme transcription with the connectionist temporal classification (CTC) sequence loss (Graves et al., 2006). This approach was then extended to character-based systems (Graves and Jaitly, 2014) and improved with attention mechanisms (Bahdanau et al., 2016; Chan et al., 2016). But the best such systems are often still behind state of the art phone-based (or senone-based) systems. Competitive end-to-end approaches leverage acoustic models (often ConvNet-based) topped with RNN layers as in (Hannun et al., 2014a; Miao et al., 2015; Saon et al., 2015; Amodei et al., 2016; Zhou et al., 2018; Zeyer et al., 2018) (e.g. a state of the art on WSJ (Chan and Lane, 2015a)), trained with a sequence criterion (the most popular ones being CTC (Graves et al., 2006) and MMI (Bahl et al., 1986)). A survey of segmental models can be found in (Tang et al., 2017). On conversational speech (that is not the topic of this paper), the state of the art is still held by complex ConvNets+RNNs acoustic models (which are also trained or refined with a sequence criterion), coupled with domain-adapted language models (Xiong et al., 2017; Saon et al., 2017).

Architecture

Our acoustic model (see an overview in Figure 1) is a Convolutional Neural Network (ConvNet) (LeCun et al., 1995), with Gated Linear Units (GLUs) (Dauphin et al., 2017) and dropout applied to activations of each layer except the output one. The model is fed with log-mel filterbank features, and is trained with either the Connectionist Temporal Classification (CTC) criterion (Graves et al., 2006), or with the ASG criterion: a variant of CTC that does not have blank labels but employs a simple duration model through letter transition scores (Collobert et al., 2016). At inference, the acoustic model is coupled with a decoder which performs a beam search, constrained with a count-based language model. We detail each of these components in the following.

The acoustic model architecture is a 1D Gated Convolutional Neural Network (Gated ConvNet), trained to map a sequence of audio features to its corresponding letter transcription. Given a dictionary of letters L{\cal L}, the ConvNet (which acts as a sliding-approach over the input sequence) outputs one score for each letter in the dictionary, for each input frame. In the transcription, words are separated by a special letter, denoted #.

Gated ConvNets have been shown to reduce the vanishing gradient problem, as they provide a linear path for the gradients while retaining non-linear capabilities, leading to state of the art performance both for natural language modeling and machine translation tasks (Dauphin et al., 2017; Gehring et al., 2017).

2 Acoustic Model Training

We considered two structured-output learning approaches to train our acoustic models: the Connectionist Temporal Classification (CTC), and a variant called AutoSeG (ASG).

CTC (Graves et al., 2006) efficiently enumerates all possible sequences of sub-word units (e.g. letters) which can lead to the correct transcription, and promotes the score of these sequences. CTC also allows a special “blank” state to be optionally inserted between each sub-word unit. The rationale behind the blank state is two-fold: (i) modeling “garbage” frames which might occur between each letter and (ii) identifying the separation between two identical consecutive sub-word units in a transcription. Figure 2a shows the CTC graph describing all the possible sequences of letters leading to the word “cat”, over 6 frames. We denote Gctc(θ,T){\cal G}_{ctc}(\theta,T) the CTC acceptance graph over TT frames for a given transcription θ\theta, and π=π1, …, πT∈Gctc(θ,T)\pi={\pi_{1},\,\dots,\,\pi_{T}}\in{\cal G}_{ctc}(\theta,T) a path in this graph representing a (valid) sequence of letters for this transcription. CTC assumes that the network outputs probability scores, normalized at the frame level. At each time step tt, each node of the graph is assigned with its corresponding log-probability letter ii (that we denote fit(X)f^{t}_{i}(\mathbf{X})) output by the acoustic model (given an acoustic sequence X\mathbf{X}). CTC minimizes the Forward score over the graph Gctc(θ,T){\cal G}_{ctc}(\theta,T):

where the “logadd” operation (also called “log-sum-exp”) is defined as logadd⁡(a,b)=log⁡(exp⁡(a)+exp⁡(b))\operatorname*{logadd}(a,b)=\log(\exp(a)+\exp(b)). This overall score can be efficiently computed with the Forward algorithm.

2.2 The ASG Criterion

Blank labels introduce code complexity when decoding letters into words. Indeed, with blank labels “ø”, a word gets many entries in the sub-word unit transcription dictionary (e.g. the word “cat” can be represented as “c a t”, “c ø a t”, “c ø a t”, “c ø a ø t”, etc… – instead of only “c a t”). We replace the blank label by special letters modeling repetitions of preceding letters. For example “caterpillar” can be written as “caterpil1ar”, where “1” is a label to represent one repetition of the previous letter.

The AutoSeG (ASG) criterion (Collobert et al., 2016) removes the blank labels from the CTC acceptance graph Gctc(θ,T){\cal G}_{ctc}(\theta,T) (shown in Figure 2a) leading to a simpler graph that we denote Gasg(θ,T){\cal G}_{asg}(\theta,T) (shown in Figure 2b). In contrast to CTC which assumes per-frame normalization for the acoustic model scores, ASG implements a sequence-level normalization to prevent the model from diverging (the corresponding graph enumerating all possible sequences of letters is denoted Gasg(θ,T){\cal G}_{asg}(\theta,T), as shown in Figure 2c). ASG also uses unnormalized transition scores gi,j(⋅)g_{i,j}(\cdot) on each edge of the graph, when moving from label ii to label jj, that are trained jointly with the acoustic model. This leads to the following criterion::

The left-hand part in Equation (\refeq−asg)(\ref{eq-asg}) promotes the score of letter sequences leading to the right transcription (as in Equation (2) for CTC), and the right-hand part demotes the score of all sequences of letters. As for CTC, these two parts can be efficiently computed with the Forward algorithm.

When removing transitions in Equation (3), the sequence-level normalization becomes equivalent to the frame-level normalization found in CTC, and the ASG criterion is mathematically equivalent to CTC with no blank labels. However, in practice, we observed that acoustic models trained with a transition-free ASG criterion had a hard time to converge.

2.3 Other Training Considerations

We apply dropout at the output to all layers of the acoustic model. Dropout retains each output with a probability pp, by applying a multiplication with a Bernoulli random variable taking value 11 with probability pp and otherwise (Srivastava et al., 2014).

Following the original implementation of Gated ConvNets (Dauphin et al., 2017), we found that using both weight normalization (Salimans and Kingma, 2016) and gradient clipping (Pascanu et al., 2013) were speeding up training convergence. The clipping we implemented performs:

where CC is either the CTC or ASG criterion, and ϵ\epsilon is some hyper-parameter which controls the maximum amplitude of the gradients.

3 Beam-Search Decoder

We wrote our own one-pass decoder, which performs a simple beam-search with beam thresholding, histogram pruning and language model smearing (Steinbiss et al., 1994). We kept the decoder as simple as possible (under 1000 lines of C code). We did not implement any sort of model adaptation before decoding, nor any word graph rescoring. Our decoder relies on KenLM (Heafield et al., 2013) for the language modeling part. It also accepts unnormalized acoustic scores (transitions and emissions from the acoustic model) as input. The decoder attempts to maximize the following:

where Glex(θ,T){\cal G}_{lex}(\theta,T) is a graph constrained by lexicon, Plm(θ)P_{lm}(\theta) is the probability of the language model given a transcription θ\theta, α\alpha, β\beta, and γ\gamma are three hyper-parameters which control the weight of the language model, and the silence (#) insertion penalty, respectively.

The beam of the decoder tracks paths with highest scores according to Equation (5), by bookkeeping pairs of (language model, lexicon) states, as it goes through time. The language model state corresponds to the (n−1)(n-1)-gram history of the nn-gram language model, while the lexicon state is the sub-word unit position in the current word hypothesis. To maintain diversity in the beam, paths with identical (language model, lexicon) states are merged. Note that traditional decoders combine the scores of the merged paths with a max⁡(⋅)\max(\cdot) operation (as in a Viterbi beam-search) – which would correspond to a max⁡(⋅)\max(\cdot) operation in Equation (5) instead of logadd⁡(⋅)\operatorname*{logadd}(\cdot). We consider instead the logadd⁡(⋅)\operatorname*{logadd}(\cdot) operation (as first suggested by Bottou (1991)), as it takes into account the contribution of all the paths leading to the same transcription, in the same way we do during training (see Equation (3)). In Section 4.1, we show that this leads to better accuracy in practice.

Experiments

We benchmarked our system on WSJ (about 8181h of labeled audio data) and LibriSpeech (Panayotov et al., 2015) (about 960960h). We kept the original 16 kHz sampling rate. For WSJ, we considered the classical setup si284 for training, dev93 for validation, and eval92 for evaluation. For LibriSpeech, we considered the two available setups clean and other. All the hyper-parameters of our system were tuned on validation sets. Test sets were used only for the final evaluations.

The letter vocabulary L{\cal L} contains 30 graphemes: the standard English alphabet plus the apostrophe, silence (#), and two special “repetition” graphemes which encode the duplication (once or twice) of the previous letter (see Section 3.2.2). Decoding is achieved with our own decoder (see Section 3.3). We used standard language models for both datasets, i.e. a 4-gram model (with about 165K165K words in the dictionary) trained on the provided data for WSJ, and a 4-gram modelhttp://www.openslr.org/11. (about 200K200K words) for LibriSpeech. In the following, we either report letter-error-rates (LERs) or word-error-rates (WERs).

Training was performed with stochastic gradient descent on WSJ, and mini-batches of 44 utterances on LibriSpeech. Clipping parameter (see Equation (4)) was set to ϵ=0.2\epsilon=0.2. We used a momentum of 0.90.9. Input features, log-mel filterbanks, were computed with 40 coefficients, a 25 ms sliding window and 10 ms stride.

We implemented everything using Torch7http://www.torch.ch.. The CTC and ASG criterions, as well as the decoder were implemented in C (and then interfaced into Torch).

We tuned our acoustic model architectures by grid search, validating on the dev sets. We consider here two architectures, with low and high amount of dropout (see the parameter pp in Section 3.2.3). Table 1 reports the details of our architectures. The amount of dropout, number of hidden units, as well as the convolution kernel width are increased linearly with the depth of the neural network. Note that as we use Gated Linear Units (see Section 3.1), each layer is duplicated as stated in Equation (1). Convolutions are followed by a fully connected layer, before the final layer which outputs 3030 scores (one for each letter in the dictionary). Concerning WSJ, the Low Dropout (p=0.2p=0.2) architecture has about 1717M trainable parameters. For LibriSpeech, architectures have about 130130M and 208208M of parameters for the Low Dropout (p=0.2p=0.2) and High Dropout (p=0.2→0.6p=0.2\rightarrow 0.6) architectures, respectively.

Figure 3 shows the LER and WER on the LibriSpeech development sets, for the first 4040 training epochs of our Low Dropout architecture. LER and WER appear surprisingly well correlated, both on the “clean” and “other” version of the dataset.

In Table LABEL:tbl-variants-wer, we report WERs on the LibriSpeech development sets, both for our Low Dropout and High Dropout architectures. Increasing dropout regularize the acoustic model in a way which impacts significantly generalization, the effect being stronger on noisy speech.

Table LABEL:tbl-variants-wer-wsj and Table LABEL:tbl-variants-wer also report the WER for the decoder ran with the max⁡(⋅)\max(\cdot) operation (instead of logadd⁡(⋅)\operatorname*{logadd}(\cdot) for other results) used to aggregate paths in the beam with identical (language model, lexicon) states. It appears advantageous (as there is no code complexity increase in the decoder – one only needs to replace max⁡(⋅)\max(\cdot) by logadd⁡(⋅)\operatorname*{logadd}(\cdot) in the code) to use the logadd⁡(⋅)\operatorname*{logadd}(\cdot) aggregation.

Figure 4 depicts alignments of the models with CTC and ASG criterions when forced aligned to a given target. Our analysis shows that the model with CTC criterion exhibits 500 ms delay compared to the model with ASG criterion. Similar observation was also previously noted in Sak et al. (2015).

2 Comparison with other systems

In Table 4, we compare our system with existing phone-based and letter-based approaches on WSJ and LibriSpeech. Phone-based acoustic state of the art models are reported as reference. These systems output in general senones; senones are carefully selected through a procedure involving a phonetic-context-based decision tree built from another GMM/HMM system. Phone-based systems also require an additional word lexicon which translates words into a sequence of phones. Most state of the art systems also perform speaker adaptation; iVectors compute a speaker embedding capturing both speaker and environment information (Xue et al., 2014), while fMMLR is a two-pass decoder technique which computes a speaker transform in the first pass (Gales and Woodland, 1996). Even though Table 3 associates speaker adaptation exclusively with phone-based systems, speaker adaptation can be also applied to letter-based systems.

State of the art performance for letter-based models on LibriSpeech is held by Deep Speech 2 (Amodei et al., 2016) and (Zeyer et al., 2018) on noisy and clean subsets respectively. On WSJ state of the art performance is held by Deep Speech 2. Deep Speech 2 uses an acoustic model composed of a ConvNet and a Recurrent Neural Network (RNN). Deep Speech 2 relies on a lot of extra speech data at training, combined with a very large 5-gram language model at inference time to make the letter-based approach competitive. Our system outperforms Deep Speech 2 on clean data, even though our system has been trained with an order of magnitude less data. Acoustic model in (Zeyer et al., 2018) is also based on RNNs and in addition employs attention mechanism. With LSTM language model their system shows lower WER than our, but with a simple 4-gram language model our system has slightly lower WER.

On WSJ the state of the art is a phone-based approach (Chan and Lane, 2015b) which leverages an acoustic model combining CNNs, bidirectional LSTMs, and deep fully connected neural networks. The system also performs speaker adaptation at inference. We also compare with existing letter-based approaches on WSJ, which are abundant in the literature. They rely on recurrent neural networks, often bi-directional, and in certain cases combined with ConvNet architectures. Our system matches the best reported letter-based WSJ performance. The Gated ConvNet appears to be very strong at modeling complete words as it achieves 6.7%6.7\% WER on LibriSpeech clean data even with no decoder, i.e. on the raw output of the neural network.

Concerning LibriSpeech, we summarize existing state of the art systems in Table 3. We highlighted the acoustic model architectures, as well as the type of underlying sub-word units.

Conclusion

We have introduced a simple end-to-end automatic speech recognition system, which combines a ConvNet acoustic model with Gated Linear Units, and a simple beam-search decoder. The acoustic model is trained to map audio sequences to sequences of characters using a structured-output learning approach based on a variant of CTC. Our system outperforms existing letter-based approaches (which do not use extra data at training time or powerful LSTM language models), both on WSJ and LibriSpeech. Overall phone-based approaches are still holding the state of the art, but our system’s performance is competitive on LibriSpeech, suggesting pronunciations is implicitly well modeled with enough training data. Further work should include leveraging speaker identity, training from the raw waveform, data augmentation, training with more data, and better language models.

References