A Density Ratio Approach to Language Model Fusion in End-To-End Automatic Speech Recognition

Erik McDermott, Hasim Sak, Ehsan Variani

Introduction

End-to-end models such as Listen, Attend & Spell (LAS) or the Recurrent Neural Network Transducer (RNN-T) are sequence models that directly define P(W∣X)P(W|X), the posterior probability of the word or subword sequence WW given an audio frame sequence XX, with no chaining of sub-module probabilities. State-of-the-art, or near state-of-the-art results have been reported for these models on challenging tasks .

End-to-end ASR models in essence do not include independently trained symbols-only or acoustics-only sub-components. As such, they do not provide a clear role for language models P(W)P(W) trained only on text/transcript data. There are, however, many situations where we would like to use a separate LM to complement or modify a given ASR system. In particular, no matter how plentiful the paired {audio, transcript} training data, there are typically orders of magnitude more text-only data available. There are also many practical applications of ASR where we wish to adapt the language model, e.g., biasing the recognition grammar towards a list of specific words or phrases for a specific context.

The research community has been keenly aware of the importance of this issue, and has responded with a number of approaches, under the rubric of “Fusion”. The most popular of these is “Shallow Fusion” , which is simple log-linear interpolation between the scores from the end-to-end model and the separately-trained LM. More structured approaches, “Deep Fusion” , “Cold Fusion” and “Component Fusion” jointly train an end-to-end model with a pre-trained LM, with the goal of learning the optimal combination of the two, aided by gating mechanisms applied to the set of joint scores. These methods have not replaced the simple Shallow Fusion method as the go-to method in most of the ASR community. Part of the appeal of Shallow Fusion is that it does not require model retraining – it can be applied purely at decoding time. The Density Ratio approach proposed here can be seen as an extension of Shallow Fusion, sharing some of its simplicity and practicality, but offering a theoretical grounding in Bayes’ rule.

After describing the historical context, theory and practical implementation of the proposed Density Ratio method, this article describes experiments comparing the method to Shallow Fusion in a cross-domain scenario. An RNN-T model was trained on large-scale speech data with semi-supervised transcripts from YouTube videos, and then evaluated on data from a live Voice Search service, using an RNN-LM trained on Voice Search transcripts to try to boost performance. Then, exploring the transition between cross-domain and in-domain, limited amounts of Voice Search speech data were used to fine-tune the YouTube-trained RNN-T model, followed by LM fusion via both the Density Ratio method and Shallow Fusion. The ratio method was found to produce consistent gains over Shallow Fusion in all scenarios examined.

A Brief History of Language Model incorporation in ASR

Generative models and Bayes’ rule. The Noisy Channel Model underlying the origins of statistical ASR used Bayes’ rule to combine generative models of both the acoustics p(X∣W)p(X|W) and the symbol sequence P(W)P(W):

for an acoustic feature vector sequence X=\mboxx1,...,\mboxxTX={\mbox{\bf x}}_{1},...,{\mbox{\bf x}}_{T} and a word or sub-word sequence W=s1,...,sUW=s_{1},...,s_{U} with possible time alignments SW={...,s,...}S_{W}=\{...,{\bf s},...\}. ASR decoding then uses the posterior probability P(W∣X)P(W|X). A prior p(s∣W)p({\bf s}|W) on alignments can be implemented e.g. via a simple 1st-order state transition model. Though lacking in discriminative power, the paradigm provides a clear theoretical framework for decoupling the acoustic model (AM) p(X∣W)p(X|W) and LM P(W)P(W).

Hybrid model for DNNs/LSTMs within original ASR framework. The advent of highly discriminative Deep Neural Networks (DNNs) and Long Short Term Memory models (LSTMs) posed a challenge to the original Noisy Channel Model, as they produce phoneme- or state- level posteriors P(s(t)∣\mboxxt)P({\bf s}(t)|{\mbox{\bf x}}_{t}), not acoustic likelihoods p(\mboxxt∣s(t))p({\mbox{\bf x}}_{t}|{\bf s}(t)). The “hybrid” model proposed the use of scaled likelihoods, i.e. posteriors divided by separately estimated state priors P(w)P(w). For bidirectional LSTMs, the scaled-likelihood over a particular alignment s{\bf s} is taken to be

using k(X)k(X) to represent a p(X)p(X)-dependent term shared by all hypotheses WW, that does not affect decoding. This “pseudo-generative” score can then be plugged into the original model of Eq. (1) and used for ASR decoding with an arbitrary LM P(W)P(W). For much of the ASR community, this approach still constitutes the state-of-the-art .

Shallow Fusion. The most popular approach to LM incorporation for end-to-end ASR is a linear interpolation,

with no claim to direct interpretability according to probability theory, and often a reward for sequence length ∣W∣|W|, scaled by a factor β\beta .

Language Model incorporation into End-to-end ASR, using Bayes’ rule

The model makes the following assumptions:

The source domain ψ\psi has some true joint distribution Pψ(W,X)P_{\psi}(W,X) over text and audio;

The target domain τ\tau has some other true joint distribution Pτ(W,X)P_{\tau}(W,X);

A source domain end-to-end model (e.g. RNN-T) captures Pψ(W∣X)P_{\psi}(W|X) reasonably well;

Separately trained LMs (e.g. RNN-LMs) capture Pψ(W)P_{\psi}(W) and Pτ(W)P_{\tau}(W) reasonably well;

pψ(X∣W)p_{\psi}(X|W) is roughly equal to pτ(X∣W)p_{\tau}(X|W), i.e. the two domains are acoustically consistent; and

The target domain posterior, Pτ(W∣X)P_{\tau}(W|X), is unknown.

The starting point for the proposed Density Ratio Method is then to express a “hybrid” scaled acoustic likelihood for the source domain, in a manner paralleling the original hybrid model :

Given the stated assumptions, one can then estimate the target domain posterior as:

with k(X)=pψ(X)/pτ(X)k(X)=p_{\psi}(X)/p_{\tau}(X) shared by all hypotheses WW, and the ratio Pτ(W)/Pψ(W)P_{\tau}(W)/{P_{\psi}(W)} (really a probablity mass ratio) giving the proposed method its name.

In essence, this model is just an application of Bayes’ rule to end-to-end models and separate LMs. The approach can be viewed as the sequence-level version of the classic hybrid model . Similar use of Bayes’ rule to combine ASR scores with RNN-LMs has been described elsewhere, e.g. in work connecting grapheme-level outputs with word-level LMs . However, to our knowledge this approach has not been applied to end-to-end models in cross-domain settings, where one wishes to leverage a language model from the target domain. For a perspective on a “pure” (non-hybrid) deep generative approach to ASR, see .

2 Top-down fundamentals of RNN-T

The RNN Transducer (RNN-T) defines a sequence-level posterior P(W∣X)P(W|X) for a given acoustic feature vector sequence X=\mboxx1,...,\mboxxTX={\mbox{\bf x}}_{1},...,{\mbox{\bf x}}_{T} and a given word or sub-word sequence W=s1,...,sUW=s_{1},...,s_{U} in terms of possible alignments SW={...,(s,t),...}S_{W}=\{...,({\bf s},{\bf t}),...\} of WW to XX. The tuple (s,t)({\bf s},{\bf t}) denotes a specific alignment sequence, a symbol sequence and corresponding sequence of time indices, consistent with the sequence WW and utterance XX. The symbols in s{\bf s} are elements of an expanded symbol space that includes optional, repeatable blank symbols used to represent acoustics-only path extensions, where the time index is incremented, but no non-blank symbols are added. Conversely, non-blank symbols are only added to a partial path time-synchronously. (I.e., using ii to index elements of s{\bf s} and t{\bf t}, ti+1=ti+1t_{i+1}=t_{i}+1 if si+1s_{i+1} is blank, and ti+1=tit_{i+1}=t_{i} if si+1s_{i+1} is non-blank). P(W∣X)P(W|X) is defined by summing over alignment posteriors:

Finally, P(si+1∣X,ti,s1:i)P(s_{i+1}|X,t_{i},s_{1:i}) is defined using an LSTM-based acoustic encoder with input XX, an LSTM-based label encoder with non-blank inputs ss, and a feed-forward joint network combining outputs from the two encoders to produce predictions for all symbols ss, including the blank symbol.

The Forward-Backward algorithm can be used to calculate Eq. (7) efficiently during training, and Viterbi-based beam search (based on the argmax over possible alignments) can be used for decoding when WW is unknown .

3 Application of Shallow Fusion to RNN-T

Shallow Fusion (Eq. (3)) can be implemented in RNN-T for each time-synchronous non-blank symbol path extension. The LM score corresponding to the same symbol extension can be “fused” into the log-domain score used for decoding:

This is only done when the hypothesized path extension si+1s_{i+1} is a non-blank symbol; the decoding score for blank symbol path extensions is the unmodified log⁡P(si+1∣X,ti,s1:i)\log P(s_{i+1}|X,t_{i},s_{1:i}).

4 Application of the Density Ratio Method to RNN-T

Eq. (6) can be implemented via an estimated RNN-T “pseudo-posterior”, when si+1s_{i+1} is a non-blank symbol:

This estimate is not normalized over symbol outputs, but it plugs into Eq. (8) and Eq. (7) to implement the RNN-T version of Eq. (6). In practice, scaling factors λψ\lambda_{\psi} and λτ\lambda_{\tau} on the LM scores, and a non-blank reward β\beta, are used in the final decoding score:

5 Implementation

The ratio method is very simple to implement. The procedure is essentially to:

Train an end-to-end model such as RNN-T on a given source domain training set ψ\psi (paired audio/transcript data);

Train a neural LM such as RNN-LM on text transcripts from the same training set ψ\psi;

Train a second RNN-LM on the target domain τ\tau;

When decoding on the target domain, modify the RNN-T output by the ratio of target/training RNN-LMs, as defined in Eq. (11), and illustrated in Fig. 1.

The method is purely a decode-time method; no joint training is involved, but it does require tuning of the LM scaling factor(s) (as does Shallow Fusion). A held-out set can be used for that purpose.

Training, development and evaluation data

The following data sources were used to train the RNN-T and associated RNN-LMs in this study.

Source-domain baseline RNN-T: approximately 120M segmented utterances (190,000 hours of audio) from YouTube videos, with associated transcripts obtained from semi-supervised caption filtering .

Source-domain normalizing RNN-LM: transcripts from the same 120M utterance YouTube training set. This corresponds to about 3B tokens of the sub-word units used (see below, Section 5.1).

Target-domain RNN-LM: 21M text-only utterance-level transcripts from anonymized, manually transcribed audio data, representative of data from a Voice Search service. This corresponds to about 275M sub-word tokens.

Target-domain RNN-T fine-tuning data: 10K, 100K, 1M and 21M utterance-level {audio, transcript} pairs taken from anonymized, transcribed Voice Search data. These fine-tuning sets roughly correspond to 10 hours, 100 hours, 1000 hours and 21,000 hours of audio, respectively.

2 Dev and Eval Sets

The following data sources were used to choose scaling factors and/or evaluate the final model performance.

Source-domain Eval Set (YouTube). The in-domain performance of the YouTube-trained RNN-T baseline was measured on speech data taken from Preferred Channels on YouTube . The test set is taken from 296 videos from 13 categories, with each video averaging 5 minutes in length, corresponding to 25 hours of audio and 250,000 word tokens in total.

Target-domain Dev & Eval sets (Voice Search). The Voice Search dev and eval sets each consist of approximately 7,500 anonymized utterances (about 33,000 words and corresponding to about 8 hours of audio), distinct from the fine-tuning data described earlier, but representative of the same Voice Search service.

Cross-domain evaluation: YouTube-trained RNN-T →→\rightarrow Voice Search

The first set of experiments uses an RNN-T model trained on {audio, transcript} pairs taken from segmented YouTube videos, and evaluates the cross-domain generalization of this model to test utterances taken from a Voice Search dataset, with and without fusion to an external LM.

The overall structure of the models used here is as follows:

Acoustic features: 768-dimensional feature vectors obtained from 3 stacked 256-dimensional logmel feature vectors, extracted every 20 msec from 16 kHz waveforms, and sub-sampled with a stride of 3, for an effective final feature vector step size of 60 msec.

Acoustic encoder: 6 LSTM layers x (2048 units with 1024-dimensional projection); bidirectional.

Label encoder (aka “decoder” in end-to-end ASR jargon): 1 LSTM layer x (2048 units with 1024-dimensional projection).

RNN-T joint network hidden dimension size: 1024.

Output classes: 10,000 sub-word “morph” units , input via a 512-dimensional embedding.

Total number of parameters: approximately 340M

RNN-LMs for both source and target domains were set to match the RNN-T decoder structure and size:

1 layer x (2048 units with 1024-dimensional projection).

Output classes: 10,000 morphs (same as the RNN-T).

Total number of parameters: approximately 30M.

The RNN-T and the RNN-LMs were independently trained on 128-core tensor processing units (TPUs) using full unrolling and an effective batch size of 4096. All models were trained using the Adam optimization method for 100K-125K steps, corresponding to about 4 passes over the 120M utterance YouTube training set, and 20 passes over the 21M utterance Voice Search training set. The trained RNN-LM perplexities (shown in Table 1) show the benefit to Voice Search test perplexity of training on Voice Search transcripts.

2 Experiments and results

In the first set of experiments, the constraint λψ=λτ\lambda_{\psi}=\lambda_{\tau} was used to simplify the search for the LM scaling factor in Eq. 11. Fig. 2 and Fig. 3 illustrate the different relative sensitivities of WER to the LM scaling factor(s) for Shallow Fusion and the Density Ratio method, as well as the effect of the RNN-T sequence length scaling factor, measured on the dev set.

The LM scaling factor affects the relative value of the symbols-only LM score vs. that of the acoustics-aware RNN-T score. This typically alters the balance of insertion vs. deletion errors. In turn, this effect can be offset (or amplified) by the sequence length scaling factor β\beta in Eq. (3), in the case of RNN-T, implemented as a non-blank symbol emission reward. (The blank symbol only consumes acoustic frames, not LM symbols ). Given that both factors have related effects on overall WER, the LM scaling factor(s) and the sequence length scaling factor need to be tuned jointly.

Fig. 2 and Fig. 3 illustrate the different relative sensitivities of WER to these factors for Shallow Fusion and the Density Ratio method, measured on the dev set.

In the second set of experiments, β\beta was fixed at -0.1, but the constraint λψ=λτ\lambda_{\psi}=\lambda_{\tau} was lifted, and a range of combinations was evaluated on the dev set. The results are shown in Fig. 4. The shading in Figs. 2, 3 and 4 uses the same midpoint value of 15.0 to highlight the results.

The best combinations of scaling factors from the dev set evaluations (see Fig. 2, Fig. 3 and Fig. 4) were used to generate the final eval set results, WERs and associated deletion, insertion and substitution rates, shown in Table 2. These results are summarized in Table 3, this time showing the exact values of LM scaling factor(s) used.

Fine-tuning a YouTube-trained RNN-T using limited Voice Search audio data

The experiments in Section 5 showed that an LM trained on text from the target Voice Search domain can boost the cross-domain performance of an RNN-T. The next experiments examined fine-tuning the original YouTube-trained RNN-T on varied, limited amounts of Voice Search {audio, transcript} data. After fine-tuning, LM fusion was applied, again comparing Shallow Fusion and the Density Ratio method.

Fine-tuning simply uses the YouTube-trained RNN-T model to warm-start training on the limited Voice Search {audio, transcript} data. This is an effective way of leveraging the limited Voice Search audio data: within a few thousand steps, the fine-tuned model reaches a decent level of performance on the fine-tuning task – though beyond that, it over-trains. A held-out set can be used to gauge over-training and stop training for varying amounts of fine-tuning data.

The experiments here fine-tuned the YouTube-trained RNN-T baseline using 10 hours, 100 hours and 1000 hours of Voice Search data, as described in Section 4.1. (The source domain RNN-LM was not fine-tuned). For each fine-tuned model, Shallow Fusion and the Density Ratio method were used to evaluate incorporation of the Voice Search RNN-LM, described in Section 5, trained on text transcripts from the much larger set of 21M Voice Search utterances. As in Section 5, the dev set was used to tune the LM scaling factor(s) and the sequence length scaling factor β\beta. To ease parameter tuning, the constraint λψ=λτ\lambda_{\psi}=\lambda_{\tau} was used for the Density Ratio method. The best combinations of scaling factors from the dev set were then used to generate the final eval results, which are shown in Table 3

Discussion

The experiments described here examined the generalization of a YouTube-trained end-to-end RNN-T model to Voice Search speech data, using varying quantities (from zero to 100%) of Voice Search audio data, and 100% of the available Voice Search text data. The results show that in spite of the vast range of acoustic and linguistic patterns covered by the YouTube-trained model, it is still possible to improve performance on Voice Search utterances significantly via Voice Search specific fine-tuning and LM fusion. In particular, LM fusion significantly boosts performance when only a limited quantity of Voice Search fine-tuning data is used.

The Density Ratio method consistently outperformed Shallow Fusion for the cross-domain scenarios examined, with and without fine-tuning to audio data from the target domain. Furthermore, the gains in WER over the baseline are significantly larger for the Density Ratio method than for Shallow Fusion, with up to 28% relative reduction in WER (17.5% →\rightarrow 12.5%) compared to up to 17% relative reduction (17.5% →\rightarrow 14.5%) for Shallow Fusion, in the no fine-tuning scenario.

Notably, the “sweet spot” of effective combinations of LM scaling factor and sequence length scaling factor is significantly larger for the Density Ratio method than for Shallow Fusion (see Fig. 2 and Fig. 3). Compared to Shallow Fusion, larger absolute values of the scaling factor can be used.

A full sweep of the LM scaling factors (λψ\lambda_{\psi} and λτ\lambda_{\tau}) can improve over the constrained setting λψ=λτ\lambda_{\psi}=\lambda_{\tau}, though not by much. Fig. 4 shows that the optimal setting of the two factors follows a roughly linear pattern along an off-diagonal band.

Fine-tuning using transcribed Voice Search audio data leads to a large boost in performance over the YouTube-trained baseline. Nonetheless, both fusion methods give gains on top of fine-tuning, especially for the limited quantities of fine-tuning data. With 10 hours of fine-tuning, the Density Ratio method gives a 20% relative gain in WER, compared to 12% relative for Shallow Fusion. For 1000 hours of fine-tuning data, the Density Ratio method gives a 10.5% relative gave over the fine-tuned baseline, compared to 7% relative for Shallow Fusion. Even for 21,000 hours of fine-tuning data, i.e. the entire Voice Search training set, the Density Ratio method gives an added boost, from 7.8% to 7.4% WER, a 5% relative improvement.

A clear weakness of the proposed method is the apparent need for scaling factors on the LM outputs. In addition to the assumptions made (outlined in Section 3.1), it is possible that this is due to the implicit LM in the RNN-T being more limited than the RNN-LMs used.

Summary

This article proposed and evaluated experimentally an alternative to Shallow Fusion for incorporation of an external LM into an end-to-end RNN-T model applied to a target domain different from the source domain it was trained on. The Density Ratio method is simple conceptually, easy to implement, and grounded in Bayes’ rule, extending the classic hybrid ASR model to end-to-end models. In contrast, the most commonly reported approach to LM incorporation, Shallow Fusion, has no clear interpretation from probability theory. Evaluated on a YouTube →\rightarrow Voice Search cross-domain scenario, the method was found to be effective, with up to 28% relative gains in word error over the non-fused baseline, and consistently outperforming Shallow Fusion by a significant margin. The method continues to produce gains when fine-tuning to paired target domain data, though the gains diminish as more fine-tuning data is used. Evaluation using a variety of cross-domain evaluation scenarios is needed to establish the general effectiveness of the method.

The authors thank Matt Shannon and Khe Chai Sim for valuable feedback regarding this work.

References