Hybrid Autoregressive Transducer (hat)

Ehsan Variani, David Rybach, Cyril Allauzen, Michael Riley

Introduction

The automatic speech recognition (ASR) problem is probabilistically formulated as a maximum a posteriori decoding :

where XX is a sequence of acoustic features and WW is a candidate word sequence. In conventional modeling, the posterior probability is factorized as a product of an acoustic likelihood and a prior language model (LM) via Bayes’ rule. The acoustic likelihood is further factorized as a product of conditional distributions of acoustic features given a phonetic sequence, the acoustic model (AM), and conditional distributions of a phonetic sequence given a word sequence, the pronunciation model. The phonetic model structure, pronunciations and LM are commonly represented and combined as weighted finite-state transducers (WFSTs) .

The probabilistic factorization of ASR allows a modular design that provides practical advantages for training and inference. The modularity permits training acoustic and language models independently and on different data sets. A decent acoustic model can be trained with a relatively small amount of transcribed audio data whereas the LM can be trained on text-only data, which is available in vast amounts for many languages. Another advantage of the modular modeling approach is it allows dynamic modification of the component models in well proven, principled methods such as vocabulary augmentation , LM adaptation and contextual biasing .

From a modeling perspective, the modular factorization of the posterior probability has the drawback that the parameters are trained separately with different objective functions. In conventional ASR this has been addressed to some extent by introducing discriminative approaches for acoustic modeling and language modeling . These optimize the posterior probability by varying the parameters of one model while the other is frozen.

Recent developments in the neural network literature, particularly sequence-to-sequence (Seq2Seq) modeling , have allowed the full optimization of the posterior probability, learning a direct mapping between feature sequences and orthographic-based (so-called character-based) labels, like graphemes or wordpieces. One class of these models uses a recurrent neural network (RNN) architecture such as the long short-term memory (LSTM) model followed by a softmax layer to produce label posteriors. All the parameters are optimized at the sequence level using, e.g., the connectionist temporal classification (CTC) criterion. The letter models and word model are examples of this neural model class. Except for the modeling unit, these models are very similar to conventional acoustic models and perform well when combined with an external LM during decoding (beam search). .

Another category of direct models are the encoder-decoder networks. Instead of passing the encoder activations to a softmax layer for the label posteriors, these models use a decoder network to combine the encoder information with an embedding of the previously-decoded label history that provides the next label posterior. The encoder and decoder parameters are optimized together using a time-synchronous objective function as with the RNN transducer (RNNT) or a label-synchronous objective as with listen, attend, and spell (LAS) . These models are usually trained with character-based units and decoded with a basic beam search. There has been extensive efforts to develop decoding algorithms that can use external LMs, so-called fusion methods . However, these methods have shown relatively small gains on large-scale ASR tasks . In the current state of encoder-decoder ASR models, the general belief is that encoder-decoder models with character-based units outperform the corresponding phonetic-based models. This has led some to conclude a lexicon or external LM is unnecessary .

This paper proposes the hybrid autoregressive transducer (HAT), a time-synchronous encoder-decoder model which couples the powerful probabilistic capability of Seq2Seq models with an inference algorithm that preserves modularity and external lexicon and LM integration. The HAT model allows an internal LM quality measure, useful to decide if an external language model is beneficial. The finite-history version of the HAT model is also presented and used to show how much label-context history is needed to train a state-of-the-art ASR model on a large-scale training corpus.

Time-Synchronous Estimation of Posterior: Previous Work

For an acoustic feature sequence X=X1:TX=X_{1:T} corresponding to a word sequence WW, assume Y=Y1:UY=Y_{1:U} is a tokenization of WW where Yi∈VY_{i}\in V, is either a phonetic unit or a character-based unit from a finite-size alphabet VV. The character tokenization of transcript Hat and the corresponding acoustic feature sequence X1:6X_{1:6} are depicted in the vertical and the horizontal axis of Figure 1, respectively. Define an alignment path Y~\widetilde{Y} as a sequence of edges Y~1:∣Y~∣{\widetilde{Y}}_{1:|\widetilde{Y}|} where ∣Y~∣|\widetilde{Y}|, the path length, is the number of edges in the path. Each path passes through sequence of nodes (Xti,Yui)\left(X_{t_{i}},Y_{u_{i}}\right) where 0≤ti≤T0\leq t_{i}\leq T, 0≤ui≤U0\leq u_{i}\leq U and ti≥ti−1t_{i}\geq t_{i-1}, ui≥ui−1u_{i}\geq u_{i-1}. The dotted path and the bold path in Figure 1 are two left-to-right alignment paths between X1:6X_{1:6} and Y0:3Y_{0:3}. For any sequence pairs XX and YY, there are an exponential number of alignment paths. Each time-synchronous model typically defines a subset of these paths as the permitted path inventory, which is usually referred to as the path lattice. Denote by B:Y~→YB:\widetilde{Y}\rightarrow{Y} the function that maps permitted alignment paths to the corresponding label sequence. For most models, this is many-to-one. Besides the definition of the lattice, time-synchronous models usually differ on how they define the alignment path posterior probability P(Y~∣X)P(\widetilde{Y}|X). Once defined, the posterior probability for a time-synchronous model is calculated by marginalizing over these alignment path posteriors:

The model parameters are optimized by maximizing P(Y∣X)P\left(Y|X\right).

The alignment path posterior is derived by imposing the conditional independence assumption between label sequence YY given feature sequence XX:

For inference with an external LM, the search is defined as:

For the CTC and RNNT models, the inference with an external LM is usually formulated as the following search problem:

Hybrid Autoregressive Transducer

The Hybrid Autoregressive Transducer (HAT) model is a time-synchronous encoder-decoder model which distinguishes itself from other time-synchronous models by 1) formulating the local posterior probability differently, 2) providing a measure of its internal language model quality, and 3) offering a mathematically-justified inference algorithm for external LM integration.

Blank (Duration) distribution: The input feature sequence XX is fed to a stack of RNN layers to output TT encoded vectors f1(X)=f1:T1{\bf{f^{1}}\left(X\right)={f}^{1}_{1:T}}. The label sequence is also fed to a stack of RNN layers to output g1(Y)=g1:U1{\bf{g}^{1}\left(Y\right)={g}^{1}_{1:U}}. The conditional Bernoulli distribution bt,ub_{t,u} is then calculated as:

where σ(⋅)\sigma(\cdot) is the sigmoid function, w{\bf w} is a weight vector, ∙\bullet is the dot-product and b{\bf b} is a bias term.

Label distribution: The encoder function f2(X)=f1:T2{\bf{f^{2}}\left(X\right)={f}^{2}_{1:T}} encodes input features XX and the function g2(Y)=g1:U2{\bf{g}^{2}\left(Y\right)={g}^{2}_{1:U}} encodes label embeddings. At each time position (t,u)\left(t,u\right), the joint score is calculated over all y∈Vy\in V:

where J(⋅){\bf J(\cdot)} can be any function that maps ft2+gu2{\bf{f}^{2}_{t}+{g}^{2}_{u}} to a ∣V∣|V|-dim score vector. The label posterior distribution is derived by normalizing the score functions across all labels in VV:

The alignment path posterior of the HAT model is computed by chaining the local edge posteriors of Eq 4, P(Y~∣X)=∏k=1T+U−1P(Y~k∣X,Y~1:k−1)P(\widetilde{Y}|X)=\prod_{k=1}^{T+U-1}P({\widetilde{Y}}_{k}|X,{\widetilde{Y}}_{1:k-1}). The posterior of the bold path is:

Like any other time-synchronous model, the total posterior is derived by marginalizing the alignment path posteriors using Eq 2.

2 Internal Language Model Score

The separation of blank and label posteriors allows the HAT model to produce a local and sequence-level internal LM score. The local score at each label position uu is defined as: S(y∣y0:u)≜J(gu2)S(y|y_{0:u})\triangleq{\bf J}\left({\bf{g}^{2}_{u}}\right). In other words, this is exactly the posterior score of Eq 6 but eliminating the effect of the encoder activations, ft2{\bf{f}^{2}_{t}}. The intuition here is that a language-model quality measure at label uu should be only a function of the uu previous labels and not the time frame tt or the acoustic features. Furthermore, this score can be normalized to produce a prior distribution for the next label P(y∣y0:u)=softmax(S(y∣y0:u))P\left(y|y_{0:u}\right)=\text{softmax}(S(y|y_{0:u})). The sequence-level internal LM score is:

Note that the most accurate way of calculating the internal language model is by P(y∣y0:u)=∑xP(y∣X=x,y0:u)P(X=x∣y0:u)P\left(y|y_{0:u}\right)=\sum_{x}P(y|X=x,y_{0:u})P(X=x|y_{0:u}). However, the exact sum is not easy to derive, thus some approximation like Laplace’s methods for integrals is needed. Meanwhile, the above definition can be justified for the special cases when J(ft2+gu2)≈J(ft2)+J(gu2){\bf J}\left({\bf{f}^{2}_{t}+{g}^{2}_{u}}\right)\approx{\bf J}\left({\bf{f}^{2}_{t}}\right)+{\bf J}\left({\bf{g}^{2}_{u}}\right) (proof in Appendix A.).

3 Decoding

where λ1\lambda_{1} and λ2\lambda_{2} are scalar weights. Subtracting the internal LM score, prior, from the path posterior leads to a pseudo-likelihood (Bayes’ rule), which is justified for combining with an external LM score. We use a conventional FST decoder with a decoder graph encoding the phone context-dependency (if any), pronunciation dictionary and the (external) n-gram LM. The partial path hypotheses are augmented with the corresponding state of the model. That is, a hypothesis consists of the time frame tt, a state in the decoder graph FST, and a state of the model. Paths with equivalent history, i.e. an equal label sequence without blanks, are merged and the corresponding hypotheses are recombined.

Experiments

The training set, 4040 M utterances, development set, 88 k utterances, and test set, 2525 hours, are all anonymized, hand-transcribed, representative of Google traffic queries. The training examples are 256256-dim. log Mel features extracted from a 6464 ms window every 3030 ms . The training examples are noisified with 2525 different noise styles as detailed in . Each training example is forced-aligned to get the frame level phoneme alignment. All models (baselines and HAT) are trained to predict 4242 phonemes and are decoded with a lexicon and an n-gram language model that cover a 44 M words vocabulary. The LM was trained on anonymized audio transcriptions and web documents. A maximum entropy LM is applied using an additional lattice re-scoring pass. The decoding hyper-parameters are swept on separate development set.

Three time-synchronous baselines are presented. All models use 55 layers of LSTMs with 20482048 cells per layer for the encoder. For the models with a decoder network (RNNT and HAT), each label is embedded by a 128128-dim. vector and fed to a decoder network which has 22 layers of LSTMs with 256256 cells per layer. The encoder and decoder activations are projected to a 768768-dim. vector and their sum is passed to the joint network. The joint network is a tanh layer followed by a linear layer of size 4242 and a softmax as in . The encoder and decoder networks are shared for the blank and the label posterior in the HAT model such that it has exactly the same number of parameters as the RNNT baseline. The CE and CTC models have 112112M float parameters while RNNT and HAT models have 115115M parameters.

HAT model performance: For the inference algorithm of Eq 9, we empirically observed that λ1∈(2.0,3.0)\lambda_{1}\in(2.0,3.0) and λ2≈1.0\lambda_{2}\approx 1.0 lead to the best performance. Figure 3 plots the WER for different values of λ2\lambda_{2}, for baseline HAT model and HAT model trained with multi-task learning (MTL). The best WER at λ2=0.95\lambda_{2}=0.95 on the convex curve shows that significantly down-weighting the internal LM and relying more on the external LM yields the best decoding quality. The HAT model outperforms the baseline RNNT model by 1.4%1.4\% absolute WER (17.5%17.5\% relative gain), Table 1. Furthermore, using a 2nd-pass LM gives an extra 10%10\% gain which is expected for the low oracle WER of 1.5%1.5\%. Comparing the different types of errors, it seems that the HAT model error pattern is different from all baselines. In all baselines, the deletion and insertion error rates are in the same range, whereas the HAT model makes two times more insertion than deletion errors.We are examining this behavior and hoping to have some explanation for it in the final version of the paper.

Internal language model: The internal language model score proposed in subsection 3.2 can be analytically justified when J(ft2+gu2)≈J(ft2)+J(gu2){\bf J}\left({\bf{f}^{2}_{t}+{g}^{2}_{u}}\right)\approx{\bf J}\left({\bf{f}^{2}_{t}}\right)+{\bf J}\left({\bf{g}^{2}_{u}}\right). The joint function J{\bf J} is usually a tanh function followed by a linear layer . The approximation holds iff ft2+gu2{\bf{f}^{2}_{t}+\bf{g}^{2}_{u}} falls in the linear range of the tanh function. Figure 3 presents the statistics of this 768768-dim. vector, μd\mu_{d} ({\color[rgb]{0,0,1}\bullet}), μd+σd\mu_{d}+\sigma_{d} ({\color[rgb]{0,1,0}\star}), and μd−σd\mu_{d}-\sigma_{d} ({\color[rgb]{1,0,0}+}) for 1≤d≤7681\leq d\leq 768 which are accumulated on the test set. The linear range of tanh function is specified by the two dotted horizontal lines in Figure 3. The means and large portion of one standard deviation around the means fall within the linear range of the tanh, which suggests that the decomposition of the joint function into a sum of two joint functions is plausible.

Figure 4(a) shows the prior cost, −1∣D∣∑y∈Dlog⁡PILM(y)\frac{-1}{|\mathcal{D}|}\sum_{y\in\mathcal{D}}\log P_{ILM}(y), during the first 1616 epochs of training with D\mathcal{D} being either the train or test set. The curves are plotted for both the HAT and RNNT models. Note that in case of RNNT, since blank and label posteriors are mixed together, a softmax is applied to normalize the non-blank scores to get the label posteriors, thus it just serves as an approximation. At the early steps of training, the prior cost is going down, which suggests that the model is learning an internal language model. However, after the first few epochs, the prior cost starts increasing, which suggests that the decoder network deviates from being a language model. Note that both train and test curves behave like this. One explanation of this observation can be: in maximizing the posterior, the model prefers not to choose parameters that also maximize the prior. We evaluated the performance of the HAT model when it further constrained with a cross-entropy multi-task learning (MTL) criterion that minimizes the prior cost. Applying this criterion results in decreasing the prior cost, Figure 4(b). However, this did not impact the WER. After sweeping λ2\lambda_{2}, the WER for the HAT model with MTL loss is as good as the baseline HAT model, blue curve in Figure 3. Of course this might happen because the internal language model is still weaker than the external language model used for inference.

This observation might also be explained by Bayes’ rule: log⁡P(X∣Y)∝log⁡P(Y∣X)−log⁡P(Y)\log P(X|Y)\propto\log P(Y|X)-\log P(Y). The observation that the posterior cost −log⁡P(Y∣X)-\log P(Y|X) goes down while the prior cost −log⁡P(Y)-\log P(Y) goes up suggests that the model is implicitly maximizing the log likelihood term in the left-side of the above equation. This might be why subtracting the internal language model log-probability from the log-posterior and replacing it with a strong external language model during inference led to superior performance.

Limited vs Infinite Context: Assuming that the decoder network is not taking advantage of the full label history, a natural question is how much of the label history actually is needed to achieve peak performance. The HAT model performance for different context sizes is shown in Table 2. A context size of means feeding no label history to the decoder network, which is similar to making a conditional independence assumption for calculating the posterior at each time-label position. The performance of this model is on a par with the performance of the CE and CTC baseline models, c.f. Table 1. The HAT model with context size of 11 shows 12%12\% relative WER degradation compared to the infinite history HAT model. However, the HAT models with contexts 22 and 44 are matching the performance of the baseline HAT model.

While the posterior cost of the HAT model with context size of 22 is about 10%10\% worse than the baseline HAT model loss, the WER of both models is the same. One reason for this behavior is that the external LM has compensated for any shortcoming of the model. Another explanation can be the problem of exposure bias , which refers to the mismatch between training and inference for encoder-decoder models. The error propagation in an infinite context model can be much severe than a finite context model with context cc, simply because the finite context model has a chance to reset its decision every cc labels.

Since the finite context of 22 is sufficient to perform as well as an infinite context, one can simply replace all the expensive RNN kernels in the decoder network with a ∣V∣2|V|^{2} embedding vector corresponding to all the possible permutations of a finite context of size 22. In other words, trading computation with memory, which can significantly reduce total training and inference cost.

Conclusion

The HAT model is a step toward better acoustic modeling while preserving the modularity of a speech recognition system. One of the key advantages of the HAT model is the introduction of an internal language model quantity that can be measured to better understand encoder-decoder models and decide if equipping them with an external language model is beneficial. According to our analysis, the decoder network does not behave as a language model but more like a finite context model. We presented a few explanations for this observation. Further in-depth analysis is needed to confirm the exact source of this behavior and how to construct models that are really end-to-end, meaning the prior and posterior models behave as expected for a language model and acoustic model.

References

Appendix A Internal Language Model Score

The softmax function is invertible up to an additive constant. In other words, for any two real valued vectors v=v1:d,w=w1:d{\bf v=v_{1:d}},{\bf w=w_{1:d}}, and real constant cc,

which proves the if condition. For the other side of condition, the proof goes as follow:

this completes the proof since the right-hand-side (RHS) of above equation is constant for any ii. ∎

For any two random variables XX and YY and some real valued function S(y,x)S(y,x), if the posterior distribution of YY given XX be

The proof is straight-forward following Lemma 1. ∎

For the label distribution of Eq 7 with the score function of Eq 6, if J(ft2+gu2)≈J(ft2)+J(gu2){\bf J}\left({\bf{f}^{2}_{t}+{g}^{2}_{u}}\right)\approx{\bf J}\left({\bf{f}^{2}_{t}}\right)+{\bf J}\left({\bf{g}^{2}_{u}}\right) for any tt and uu, then:

for some real valued constant c1c_{1}. The first equality holds from Corollary 1, and the second equality holds by Bayes’ rule.

Applying exponential function and marginalizing over XX results in:

note that exp⁡(c1)\exp(c_{1}) is dropped which make the left-hand-side (LHS) be proportional to the RHS of first equation. The second equation holds since P(X=x∣y,y0:u)P\left(X=x|y,y_{0:u}\right) is a distribution over XX.

where the first equality comes from the definition and the second one is the assumption made in the statement of the proposition. Applying exponential function and marginalizing over XX:

since the second term will be a constant and will not be function of yy. Comparing Eq 14 and Eq 16:

Finally note that P(y∣y0:u)=P(y,y0:u)/P(y0:u)P\left(y|y_{0:u}\right)=P\left(y,y_{0:u}\right)/P\left(y_{0:u}\right), thus: