Mixture Models for Diverse Machine Translation: Tricks of the Trade

Tianxiao Shen, Myle Ott, Michael Auli, Marc'Aurelio Ranzato

Introduction

Machine translation (MT) is a challenging task not only because of the large and structured output space, but also because it is inherently a one-to-many mapping. There are often many plausible and semantically equivalent translations due to information asymmetry between different languages, e.g., translating from a language without grammatical gender to a language that has grammatical gender leads to two valid translation options, as well as different translation styles such as formal/informal, literal/not literal, etc. This raises the question of how to model such multi-modal output distributions and how to evaluate these models.

Our first contribution is a better evaluation protocol that uses multiple references during evaluation to measure both the quality of translation and diversity of a generated hypothesis set. The second contribution of this paper is an in-depth empirical analysis of mixture models for machine translation, although we conjecture that the findings are general and might apply to other text generation tasks, such as dialogue, summarization, image captioning, etc.

Conditional mixture models, also known as mixture of experts (MoE) (Jacobs et al., 1991), are in principle well suited to generating diverse hypotheses which can be achieved through different mixture components. However, they have been largely overlooked in favor of models with richer latent structure (Zhang et al., 2016; Kaiser et al., 2018). There has been some previous work on mixture models for sequence to sequence learning (Shazeer et al., 2017; He et al., 2018), but these did not evaluate generations in terms of both quality and diversity, or they focus on a particular model variant. There is a lack of consensus whether mixture models are competitive with more complex models that rely on approximate Bayesian inference, whether they are plagued by the same “posterior collapse” degeneracy as variational models (Bowman et al., 2016), how model configurations affect performance and which one works best in practice.

This work considers all the major design choices involved in the construction of mixture models, including hard versus soft EM training, different parameterizations of mixture components, the choice of conditional prior, update frequency of responsibilities (also called membership weights), and how regularization noise is injected. We experiment on the large scale WMT English to German benchmark with a state-of-the-art model architecture and the results demonstrate intricate dependencies between these design choices. They also reveal that some ingredients are key to successful training of mixture models.

First, we show that mixture models are prone to degeneracies when trained with dropout noise, but that this can be mitigated by turning off dropout in the computation of responsibilities. The key to the specialization of experts is to make consistent use of them, and even a small amount of regularization noise can hamper that. Second, hard mixtures yield more diverse generations than soft mixtures, similar to how K-Means tends to find centroids that are farther apart from each other compared to the means found by a mixture of Gaussians (Kearns et al., 1998). Third, employing a uniform prior encourages all mixture components to produce good translations for any input source sentence, which is highly desirable. Finally, using independently parameterized mixture components provides greater diversifying capacity than shared parameters; but if responsibilities are refreshed online, independent parameterization is prone to a degeneracy where only a single component is trained because of the “rich gets richer” effect. Conversely, the combination of shared parameters and offline responsibility assignment may lead to another degeneracy, in which the mixture components fail to specialize and behave the same.

We extend our evaluation to three WMT benchmark datasets for which test sets with multiple human references are available. We demonstrate that mixture models, when successfully trained, consistently outperform variational NMT (Zhang et al., 2016) and diverse decoding algorithms such as diverse beam search (Li et al., 2017; Vijayakumar et al., 2018) and biased sampling (Graves, 2013; Fan et al., 2018). Our qualitative analysis shows that different mixture components can capture consistent translation styles across examples, enabling users to control generations in an interpretable and semantically meaningful way.

Related Work

Prior studies have investigated the prediction uncertainty in machine translation. Dreyer & Marcu (2012) and Galley et al. (2015) introduced new metrics to address uncertainty at evaluation time. Ott et al. (2018a) inspected the sources of uncertainty and proposed tools to check fitting between the model and the data distributions. They also observed that modern conditional auto-regressive NMT models can only capture uncertainty to a limited extent, and they tend to oversmooth probability mass over the hypothesis space.

Recent work has explored latent variable modeling for machine translation. Zhang et al. (2016) leveraged variational inference (Kingma & Welling, 2014; Bowman et al., 2016) to augment an NMT system with a single Gaussian latent variable. This work was extended by Schulz et al. (2018), who considered a sequence of latent Gaussian variables to represent each target word. Kaiser et al. (2018) proposed a similar model, but with groups of discrete multinomial latent variables. In their qualitative analysis, Kaiser et al. (2018) showed that the latent codes do affect the output predictions in interesting ways, but their focus was on speeding up regular decoding rather than producing a diverse set of hypotheses. None of these works analyzed and quantified diversity introduced by such latent variables.

The most relevant work is by He et al. (2018), who propose to use a soft mixture model with uniform prior for diverse machine translation. However, they did not evaluate on datasets with multiple references, nor did they analyze the full spectrum of design choices for building mixture models. Moreover, they used weaker base models and did not compare to variational NMT or diverse decoding baselines, which makes their empirical analysis less conclusive. We provide a comprehensive study and shed light on the different behaviors of mixture models in a variety of settings.

Besides machine translation, there is work on latent variables for dialogue generation (Serban et al., 2017; Cao & Clark, 2017; Wen et al., 2017) and image captioning (Wang et al., 2017; Dai et al., 2017). The proposed mixture model departs from these VAE or GAN-based approaches and importantly, is much simpler. It could also be applied to other text generation tasks as well.

Mixture Models for Diverse MT

A standard neural machine translation (NMT) model has an encoder-decoder structure. The encoder maps a source sentence xx to a sequence of hidden states, which are then fed to the decoder to generate an output sentence one word at a time. At each time step, the decoder additionally conditions its output on the previous outputs, resulting in an auto-regressive factorization p(y∣x;θ)=∏t=1Tp(yt∣y1:t−1,x;θ)p(y|x;\theta)=\prod_{t=1}^{T}p(y_{t}|y_{1:t-1},x;\theta), where (y1,⋯ ,yT)(y_{1},\cdots,y_{T}) are the words that compose a target sentence yy.

However, the machine translation task has inherent uncertainty, due to the existence of multiple valid translations yy for a given source sentence xx. With the auto-regressive factorization all uncertainty is represented in the decoder output distribution, making it difficult to search for multiple modes of p(y∣x;θ)p(y|x;\theta). Indeed, widely used decoding algorithms such as beam search typically produce hypotheses of low diversity with only minor differences in the suffix (Ott et al., 2018a).

Mixture models provide an alternative approach to modeling uncertainty and generating diverse translations. While these models have primarily been explored as a means of increasing model capacity (Jacobs et al., 1991; Shazeer et al., 2017), they are also a natural way of modeling different translation styles (He et al., 2018).

Formally, given a source sentence xx and reference translation yy, a mixture model introduces a multinomial latent variable z∈{1,⋯ ,K}z\in\{1,\cdots,K\}, and decomposes the marginal likelihood as:

where the prior p(z∣x;θ)p(z|x;\theta) and likelihood p(y∣z,x;θ)p(y|z,x;\theta) are learned functions parameterized by θ\theta. Each value of zz represents an expert, and the posterior probability:

can be viewed as the responsibility each expert takes for explaining an observation (x,y)(x,y).

Given a training set {(x(i),y(i))}i=1N\{(x^{(i)},y^{(i)})\}_{i=1}^{N}, we want to find θ\theta that maximizes the log likelihood. To this end, for each training example we compute the gradient:

and train the model with the EM algorithm (Dempster et al., 1977) by iteratively applying the following two steps:

estimate the responsibilities of each expert rz(i)←p(z∣x(i),y(i);θ)r_{z}^{(i)}\leftarrow p(z|x^{(i)},y^{(i)};\theta) using the current parameters θ\theta;

update θ\theta through each expert with gradients ∇θlog⁡p(y(i),z∣x(i);θ)\nabla_{\theta}\log p(y^{(i)},z|x^{(i)};\theta) weighted by their responsibilities rz(i)r_{z}^{(i)}.

There are several ways to apply the E and M steps across training examples (Neal & Hinton, 1998); §3.3 will review those considered in this work.

Decoding

All the decoding strategies for p(y∣x;θ)p(y|x;\theta) of a baseline model can be applied equally to p(y∣z,x;θ)p(y|z,x;\theta) in a mixture model. We adopt the most straightforward one: generating KK hypotheses by first enumerating zz and then greedily decoding y^t=arg max⁡yp(y∣y^1:t−1,z,x;θ)\hat{y}_{t}=\operatorname*{arg\,max}_{y}p(y|\hat{y}_{1:t-1},z,x;\theta). Notably, this decoding procedure is efficient and easily parallelizable.

Degeneracies

Unfortunately, naïve implementations of mixture models for text generation are prone to two major types of degeneracies:

Only one component gets trained (Eigen et al., 2014; Shazeer et al., 2017) because of the “rich gets richer” effect whereby, once a component is slightly better than others, it is always picked while the other components starve and are eventually never used.

The latent variable is ignored, similar to the collapse of variational auto-encoders where the posterior is always equal to the prior (Bowman et al., 2016).

In both cases, the model operates like a baseline model without any benefit from the latent variable. In practice, however, the chance of these degeneracies is heavily affected by a number of design decisions, which we describe in the following subsections.

1 Model Variants

Ideally, we would like different experts to specialize on different translation styles, so they can generate diverse hypotheses. Moreover, we want all of them to work well with any source sentence, so they will produce high quality translations. To this end, we explore several variants of the canonical mixture model, which vary by whether they use hard (h) or soft (s) assignments of responsibilities, and whether they use a learned prior (lp) or uniform prior (up).

The specialization of experts implies that the responsibility p(z∣x,y;θ)p(z|x,y;\theta) for explaining a particular translation yy should be large for only one zz, i.e., only one element in the sum ∑zp(y,z∣x;θ)\sum_{z}p(y,z|x;\theta) dominates. To encourage this, a hard mixture model directly optimizes for max⁡zp(y,z∣x;θ)\max_{z}p(y,z|x;\theta) by assigning full responsibility for each training example to the expert with the largest joint probability. Training proceeds via hard-EM, where the M-step remains unchanged and the E-step becomes:

estimate the responsibilities of each expert rz(i)←\mathds1[z=arg max⁡z′p(y(i),z′∣x(i);θ)]r_{z}^{(i)}\leftarrow\mathds{1}[z=\operatorname*{arg\,max}_{z^{\prime}}p(y^{(i)},z^{\prime}|x^{(i)};\theta)] using the current parameters θ\theta.

This can also be seen as maximizing the marginal likelihood p(y∣x;θ)p(y|x;\theta) while minimizing the gap between the sum and the max, thus finding a balance between the two terms (Kearns et al., 1998).

To encourage all experts to generate good hypotheses for any source sentence, we may set the prior p(z∣x;θ)p(z|x;\theta) to be uniform. This can prevent the model from collapsing into only one working expert with extreme p(z∣x;θ)p(z|x;\theta) value. This also aligns with our simple decoding strategy that generates a single hypothesis from each expert.

The choices of soft (s) versus hard (h) mixture model (M) and learned prior (lp) versus uniform prior (up) give us four model variants with different loss functions:

where the constant log⁡K\log K term is omitted for models with a uniform prior.

Among these variants, the hMup objective is also known as multiple choice learning (MCL) for an ensemble of learners where the oracle loss is minimized (Guzman-Rivera et al., 2012), and He et al. (2018) has considered the sMup objective to train a sequence-to-sequence mixture model.

2 Parameterization

Another important design decision with mixture models is the degree of parameter sharing between experts. Using independently parameterized experts provides them with additional capacity to differentiate from one another, but may exacerbate overfitting since the number of parameters increases linearly with the number of experts. On the other hand, sharing parameters among experts may help mitigate degeneracy D1, whereby low quality experts are neglected and eventually “die” during training, since by sharing parameters even unpopular experts receive some gradients.

We test different model variants using both independent and shared parameters. With independent parameterization, each expert has a different decoder network. With shared parameterization, experts use the same decoder network but the beginning-of-sentence token at the start of the target sequence is replaced with an embedded representation of the latent variable. This requires a negligible increase in parameters over the baseline model.

3 Training Schedule

We consider two schedules for alternating between the E-step and M-step during training: online and offline. In the online EM algorithm, we minimize the loss via stochastic gradient descent, effectively interleaving the E-step and M-step for each mini-batch (Lee et al., 2016). In contrast, the offline EM algorithm performs the E-step for all training examples, trains each expert to convergence with the resulting responsibilities, and repeats. In practice, for offline training we perform the M-step for only a single epoch, rather than to convergence, before re-estimating the responsibilities.

4 Regularization

Deep neural networks with a large amount of parameters are prone to overfitting, and regularization via dropout is usually adopted to achieving good generalization performance. However, the key to make experts diversify as training progresses is to make consistent use of them, and we find that even a small amount of regularization noise in the computation of responsibilities (E-step) can hamper that. Indeed, we show in §5.1 that naïve use of dropout causes mixture models to ignore the latent variable, but this degeneracy is mitigated by disabling dropout in the E-step. We explore this phenomenon in more detail in Appendix A.

Metrics

In this section, we describe the metrics we use to quantitatively assess the quality and diversity of a set of translation hypotheses. We use BLEU (Papineni et al., 2002) based on word n-gram matching to measure corpus similarity.

Suppose {y1,⋯ ,yM}\{y^{1},\cdots,y^{M}\} are MM reference translations of a source sentence xx, and {y^1,⋯ ,y^K}\{\hat{y}^{1},\cdots,\hat{y}^{K}\} are KK hypotheses. Let BLEU{([r1,⋯ ,rn],h)}x∈data\{([r_{1},\cdots,r_{n}],h)\}_{x\in\text{data}} denote the corpus-level BLEU for all pairs where hh is a hypothesis and [r1,⋯ ,rn][r_{1},\cdots,r_{n}] is its reference list. Let [n][n] denote the set of {1,⋯ ,n}\{1,\cdots,n\}, and [y−i][y^{-i}] denote [y1,⋯ ,yi−1,yi+1,⋯ ,yM][y^{1},\cdots,y^{i-1},y^{i+1},\cdots,y^{M}]. We compute the following two metricsSee Appendix B for another diversity metric in terms of reference coverage. (Ott et al., 2018a):

Pairwise-BLEU: To measure similarity among the hypotheses, we compare them with each other and compute BLEU{([y^j],y^k)}x∈data,j∈[K],k∈[K],j≠k\{([\hat{y}^{j}],\hat{y}^{k})\}_{x\in\text{data},j\in[K],k\in[K],j\neq k}.Several text GAN papers use Self-BLEU to evaluate diversity of unconditional generation (Yu et al., 2017), where each generated sentence is regarded as hypothesis against all other sentences as its reference list, i.e. BLEU{([y^−k],y^k)}x∈data,k∈[K]\{([\hat{y}^{-k}],\hat{y}^{k})\}_{x\in\text{data},k\in[K]}. For machine translation, Pairwise-BLEU offers a more precise evaluation of diversity compared to Self-BLEU. For example, if a source sentence has two valid translations T1≠T2T_{1}\neq T_{2}, system 1 provides 4 hypotheses H1=H2=H3=H4=T1H_{1}=H_{2}=H_{3}=H_{4}=T_{1}, and system 2 provides H1=H2=T1,H3=H4=T2H_{1}=H_{2}=T_{1},H_{3}=H_{4}=T_{2}, then their Self-BLEU are both 100; whereas Pairwise-BLEU for system 1 is 100 and for system 2 it is not, indicating that the latter is more diverse and more desirable (while they both have perfect quality). The more diverse the hypothesis set, the lower the Pairwise-BLEU. Ideally, we would like a model with Pairwise-BLEU matching human Pairwise-BLEU.

BLEU: We calculate human BLEU in a leave-one-out manner by computing BLEU{([y−m],ym)}x∈data\{([y^{-m}],y^{m})\}_{x\in\text{data}} for m∈[M]m\in[M] and then averaging the MM scores. We also use M−1M-1 references when computing system BLEU, to be comparable with human scores, i.e. average BLEU{([y−m],y^k)}x∈data,k∈[K]\{([y^{-m}],\hat{y}^{k})\}_{x\in\text{data},k\in[K]} for m∈[M]m\in[M]. This measures the overall quality of a hypothesis set. If this metric scores low, it implies that some generated hypotheses have poor quality.

Figure 1 illustrates how these metrics behave in different situations. The degeneracies outlined in §3 can be easily identified with these metrics: The first degeneracy is a situation where a single expert is responsible for all inputs (D1). This case can be identified when we measure both very low Pairwise-BLEU as well as very low BLEU, i.e., all but one latent value produce very bad generations. The second degeneracy happens when the latent variable is ignored and all experts behave almost identically (D2). This is evident when we observe good BLEU but extremely high Pairwise-BLEU (close to 100), i.e., when all latent values give good yet very similar outputs.

Experiments

We test mixture models and baselines on three benchmark datasets that uniquely provide multiple human references (Ott et al., 2018a; Hassan et al., 2018).

WMT’17 English-German (En-De): We train on all available bitext and filter sentence pairs that have source or target longer than 80 words, resulting in 4.5M sentence pairs. We use the Moses tokenizer (Koehn et al., 2007) and learn a joint source and target Byte-Pair-Encoding (Sennrich et al., 2016) with 32K types. We develop on newstest2013 and test on a 500 sentence subset of newstest2014 that has 10 reference translations (Ott et al., 2018a).

WMT’14 English-French (En-Fr): We borrow the setup of Gehring et al. (2017) with 36M training sentence pairs and 40K joint BPE vocabulary. We validate on newstest2012+2013, and test on a 500 sentence subset of newstest2014 with 10 reference translations (Ott et al., 2018a).

WMT’17 Chinese-English (Zh-En): We pre-process the training data following Hassan et al. (2018) which results in 20M sentence pairs, 48K and 32K source and target BPE vocabularies respectively. We develop on devtest2017 and report results on newstest2017 with 3 reference translations.

Model Architecture

All models use a very similar architecture, built using the Transformer (Vaswani et al., 2017) implementation in the Fairseq toolkit (Ott et al., 2019). The encoder and decoder have 6 blocks. The number of attention heads, embedding dimension and inner-layer dimension are 8, 512, 2048 for the “base” configuration and 16, 1024, 4096 for the “big” configuration, respectively. We use the “base” configuration to compare mixture model variants in §5.1, and the “big” configuration for extended experiments in §5.2.

Mixture models with “independent” experts use independent decoders, while models with “shared” experts use a single shared decoder with an extra set of weights to embed each latent variable state. All mixture models use a shared encoder across the experts. Models that learn a prior, namely sMlp and hMlp, have an additional module that averages the top encoder layer hidden states and predicts the conditional prior distribution p(z∣x;θ)p(z|x;\theta) via a one hidden layer neural network with a tanh activation in between; the number of hidden units matches the embedding dimension.

Baselines

The first baseline we consider use the same architecture, without a latent variable, and use beam search or sampling to generate KK hypotheses. We also consider the following modified versions: Diverse Beam Search (Vijayakumar et al., 2018) and Biased Sampling (Graves, 2013; Fan et al., 2018; Edunov et al., 2018). The former performs beam search sequentially and penalizes the selection of words used in previous generations. The latter improves upon straight sampling by sampling over the top-kk most likely words instead of all words at each step, resulting in less noisy generations. Finally, we also compare against Variational NMT (Zhang et al., 2016) which uses a Gaussian latent variable and variational inference for posterior approximation. We use a 512-dimensional latent space. At test time we first sample zz from the conditional prior distribution and then do greedy decoding. Any additional hyper-parameters for these baselines are tuned via grid search over the validation set.

Experimental Details

Models are optimized with the Adam algorithm (Kingma & Ba, 2015) using β1=0.9\beta_{1}=0.9, β2=0.98\beta_{2}=0.98, and ϵ=1e−8\epsilon=1e-8. We use the same learning rate schedule as Ott et al. (2018b). We run experiments on between 8 and 128 Nvidia V100 GPUs with mini-batches of approximately 25K and 400K tokens for the experiments of §5.1 and §5.2, respectively, following Ott et al. (2018b).

1 Analysis of Mixture Models: Tricks of the Trade

In this section we compare the four variants of mixture models (§3.1), and consider for each such variant whether to share parameters among mixture components (§3.2), whether to update responsibilities once each epoch (offline mode) or after every gradient step (online mode) (§3.3), and the effect of regularization on them (§3.4). We train on the WMT’17 En-De dataset using the “base” Transformer configuration and K=3K=3 mixture components for efficiency considerations.

Naïvely optimizing the loss functions in Equation 3.1 with dropout noise causes mixture models to ignore the latent variable (D2). This can be seen by the very high Pairwise-BLEU (low diversity) when dropout is used in both the E-step and M-step (Table 1, E&M). Unfortunately, entirely disabling dropout exacerbates overfitting and gives quite worse translation quality, i.e., lower BLEU (Table 1, no). We provide a solution to this issue by decomposing the log-likelihood gradient as in Equation 3 and applying dropout only to the gradient computation of the experts (M-step) but not their responsibilities (E-step). This allows the model to make consistent use of the experts and helps them to diversify (Table 1, M; see also Appendix A). We adopt this dropout scheme for subsequent experiments.

Mixture Model Exploration:

Figure 2 shows the results of different modeling choices we introduced in §3. We plot Pairwise-BLEU in descending order versus BLEU, so the right side indicates high diversity and the upper side indicates high quality. First, both Variational NMT (V-NMT) as well as soft mixture models with shared parameterization and offline responsibility updates lead to degeneracy D2 (upper left corner), where the latent variable is ignored and the generated hypotheses are almost the same, as indicated by Pairwise-BLEU close to 100. This is a well-known failure mode of VAEs for text (Bowman et al., 2016).We also tried annealing the KL term for Variational NMT, but even then the latent variable is hardly used. In the mixture model case, we surmise that training a shared mixture for a long time with the initial (random) responsibilities prevents the latent variable embeddings from specializing, and thus they provide no useful information to the decoder.

Second, mixture models trained with independently parameterized components and online responsibility assignment are prone to degeneracy D1 (lower right corner), whereby only a single expert gets trained due to the “rich gets richer” effect (Shazeer et al., 2017), and BLEU dramatically drops because of poor generations from other experts.We observe that the BLEU score is high for a single latent state but nearly zero for others, see Table 7 in Appendix C. Parameter sharing alleviates this, as updates to one expert affect also the others.

Third, there are two settings that work consistently well: online responsibility update with shared parameters, and offline update with independent parameters. The latter yields more diverse but lower quality translations, as expected since components have more degrees of freedom to deviate from each other but also less data to train. Fourth and not surprisingly, soft assignment yields lower diversity than hard assignment (Kearns et al., 1998). Fifth, setting the prior to be uniform encourages the model to make use of all the components for each input source sentence, and thus gives better BLEU.

Overall, there are a few variants that work robustly as indicated by the small variance from random initialization, and strike a good balance between generation quality and diversity as indicated by proximity to human performance: all online-shared models, offline-shared hMup, as well as offline-independent sMup and hMup. All of these models perform well, offering slightly different trade-offs between quality and diversity.

If we also account for computational and memory cost, methods using shared parameters and hard EM are preferable, since the extra parameters are negligible and it requires only a single backward pass for the selected expert.With K=10K=10 and the “base” Transformer architecture, hard mixture models are 2.4 times faster than the soft counterpart. If we further consider simplicity of implementation, uniform prior and online responsibility update are favorable because they do not require additional model components, nor storing the responsibilities as for offline responsibility update. Taking these considerations into account, we adopt hMup (online-shared) for subsequent experiments, keeping in mind that the model variants mentioned above are likely to work similarly well.

2 Large-Scale Evaluation

We now extend our evaluation to three large benchmark datasets using the “big” Transformer configuration and a larger numbers of mixture components (K=10,10,3K=10,10,3 for WMT En-De, En-Fr and Zh-En datasets respectively, to match the number of references we have on the test set).

Table 2 and Figure 3 show results of the hMup mixture model compared to other baseline approaches. The BLEU of beam search is fairly close to human, implying a remarkably high generation quality. However, it severely lacks diversity, as Pairwise-BLEU is close to 80. This suggests a scenario similar to Figure 1-B. In contrast, sampling hypotheses are very diverse, typically even more so than human references on En-De and En-Fr, but have a very poor BLEU, as illustrated in Figure 1-A. Diverse beam search increases the diversity of beam search to some extent, but pays a cost in translation quality. Overall, hMup achieves the best trade-off between quality and diversity.

We also explore the impact of varying KK on the WMT’17 En-De dataset (Figure 3, left). In general, when more hypotheses are generated, they become more diverse but of worse quality. Despite the improvement offered by hMup, there is still a large gap between the diversity of human references and system generations. The complete results of these experiments and the biased sampling baseline are provided in Appendix C.

3 Qualitative Analysis

In this section we perform a qualitative analysis with the WMT’17 Zh-En dataset. In Table 3 we show two source sentences, the corresponding reference translations, and generated hypotheses from different approaches. We see that beam search tends to produce generations that differ only in the last few words. Diverse beam search improves the diversity over beam search, but is not as diverse as hMup and may produce duplicate hypotheses (e.g., if the diversity penalty is not sufficiently high). hMup shows significant diversity in wording, word order, clause structure, etc.

To investigate whether the latent variable in hMup learns different translation styles, we examine the hypotheses generated from each latent state. For each value of zz, we compute word frequencies of the corresponding generations and look for words whose frequency is significantly different as we change the value of the latent variable. We first discover that for words like was, were and had, z = 1z\,{=}\,1’s frequency is more than three times higher than z = 3z\,{=}\,3’s; conversely, for has and says, z = 3z\,{=}\,3’s frequency is more than twice higher than z = 1z\,{=}\,1’s. Since Chinese does not have tense unless a time phrase is explicitly stated, we speculate that when translating into English, the first latent value tends to translate with past tense whereas the third latent value tends to translate with present tense. Indeed we find that this is a consistent behavior, as seen from the first four examples in Table 4. Similarly, we find that different latent values exhibit different preferences for using this or that (see the fourth and fifth examples in Table 4), % or per cent (see the last example in Table 4 and the first example in Table 3), and so on.

Conclusion

Using large scale benchmarks and state-of-the-art architectures, we investigated 3232 variants of mixture models, arising from the following design choices: use of hard versus soft EM, uniform versus learned prior, shared versus independent parameterization of mixture components, online versus offline responsibility update, and use of standard dropout regularization versus removal of this noise in responsibility computation. To the best of our knowledge this is the most extensive study of mixture models for conditional text generation to date, with machine translation as a use case. The simplicity of mixture models provides important advantages: the latent variable assignment can be explicitly enumerated and the posterior can be computed exactly. Despite their simplicity, however, mixture models exhibit complex behaviors depending on different combinations of design choices. In particular, when instantiated with sub-optimal choices, they are prone to two typical failure modes—only one component gets trained and other components “die”, or the latent variable is ignored and all components behave the same. Our study provides insights into training deep sequence models with discrete latent variables.

Our recommended configurations enable mixture models to offer much better trade-offs between quality and diversity than variational models as well as heuristic diverse decoding approaches. In the future, we hope to broaden the scope of this work by looking at other generation tasks such as dialogue and image captioning. We would also like to investigate models with richer and more structured latent representations, and narrow the gap between model and human performance.

We thank David Grangier for insightful discussions in the preliminary phase of this work. We also thank MIT NLP group for their helpful comments.

References

Appendix A Effect of Dropout

In Section 5.1 we observed that it is crucial to turn off dropout during the computation of responsibilities (E-step) to avoid model collapse, see Table 1. In this section, we further investigate the impact of dropout on the responsibility computation, using hMup as an example.

We speculate that dropout noise weakens the dependency on the latent variable, causing the hard E-step to select among the latent values at random. This prevents different latent states from specializing and ultimately causes the model to ignore them.

To test this hypothesis, we show in Figure 4 the effect of dropout noise on the E-step at the beginning of training (i.e., with a randomly initialized model). On the y-axis we plot how often the optimal value of zz changes after applying dropout with different rates. We see that as we increase the dropout probability, the optimal value of zz is quickly corrupted—even with a small dropout probability of 0.1 we observe a 42% chance that the optimal assignment of zz changes.

Appendix B Another Diversity Metric

In addition to Pairwise-BLEU, we also consider another metric to evaluate the diversity of a set of hypotheses, the reference coverage.

We pair each hypothesis to its best matching reference (breaking ties randomly), count how many distinct references are matched to at least one hypothesis, and average this count over all sentences in the test set. A low coverage number indicates that all hypotheses are close to a few references. Instead, we would like a diverse set that covers most of the references.

Compared to Pairwise-BLEU, this metric offers more intuitive numerical values, as it ranges between 1 and the total number of available references. In the next section, we will report both values for completeness.

Appendix C Detailed Results

Table 5 compares two well performing mixture model variants, namely (online, shared) hMup and sMup, to several baselines on the three banchmark datasets we have considered in this work, expanding on the results reported in Table 2 and Figure 3 of the main paper. Once again, mixture models offer a good trade-off between translation quality and diversity.

In Table 6 we compare different approaches for generating diverse translations on the WMT’17 En-De dataset. We additionally compare each approach as we vary the number of desired translations (KK) (see also Figure 3, left). We observe that sampling produces diverse but low quality outputs. We can improve translation quality by restricting sampling to the top-kk candidates at each output position (k = 2k\,{=}\,2 performed best), but translation quality is still worse than hMup. Beam search produces the highest quality outputs, but with low diversity. Diverse beam search provides a reasonable balance between diversity and translation quality, but hMup produces even more diverse and better quality translations. Finally, except for unrestricted sampling, hMup covers the largest number of references among all the approaches evaluated.

We conclude with Table 7 which shows the values used to generate Figure 2, together with other metrics, such as corpus level BLEU with hypotheses generated by a fixed latent variable state throughout the whole test set. This metric is useful to detect models affected by degeneracy D1, as there will be states that yield very low corpus level BLEU because they rarely generate good hypotheses.