FUDGE: Controlled Text Generation With Future Discriminators
Kevin Yang, Dan Klein
Introduction
Recent advances in large pretrained language models allow us to generate increasingly realistic text by modeling a distribution over natural language sequences . The distribution may be truly unconditional, as is common in language modeling, or it may model conditioned on some input , as in machine translation or summarization.
We are frequently interested in controlled text generation, the task of generating text conditioned on an additional desirable attribute which is not already built into . That is, we would like to model (or possibly ; henceforth we will drop from the notation for simplicity). For example, may be a pretrained translation model for Spanish inputs to English outputs , but we may wish to additionally constrain the outputs to possess a new attribute , e.g., formality, which we did not optimize for during training.
Unfortunately, once we have already obtained an unconditioned defined as the output distribution of some large generative model , it is nontrivial to add conditioning on a new attribute without either training a new model from scratch or fine-tuning with additional data. Although in principle we can trivially sample from via rejection sampling from , rejection sampling may be highly inefficient in practice. On the other hand, while generating according to attribute , should be left otherwise intact: in the previous translation formality example, it is pointless to generate formal English outputs if they do not preserve the original Spanish meaning.
In light of these concerns, we propose Future Discriminators for Generation (Fudge), a flexible and modular method for modeling which accesses only the output probabilities of the generative model which defines . Fudge learns a binary predictor for whether attribute will become true in the complete future, based on an incomplete sequence prefix (Sec. 3). Multiplying the output probabilities of this predictor with ’s original probabilities and then renormalizing yields a model for the desired via Bayes’ Rule.
We run experiments on three controlled text generation tasks — couplet completion in poetry, topic control in language generation, and formality change in machine translation — showing our method’s broad applicability. Additionally, we demonstrate the modularity of Fudge by composing multiple attribute constraints in both the couplet and topic control tasks. In our experiments, we find that Fudge is highly effective at attribute control, outperforming both a baseline which directly fine-tunes and also a strong gradient-based method (Pplm Dathathri et al. (2019)). Our code is available at https://github.com/yangkevin2/naacl-2021-fudge-controlled-generation.
Related Work
Ideally, a controlled text generation method should efficiently control for while preserving as much as possible. Recent work on controlled text generation has greatly advanced our ability to control for a required attribute flexibly and cheaply, with varying degrees of modification to the original model which defines .
One line of work fine-tunes a pretrained model for a desired attribute Ficler and Goldberg (2017); Yu et al. (2017); Ziegler et al. (2019). The result is a class-conditional language model (CCLM). However, it is difficult to isolate the desired attribute from the distribution shift between and the fine-tuning dataset Hu et al. (2017); John et al. (2018); Lazaridou et al. (2020), i.e., it is nontrivial to preserve the desirable qualities of the modeled by . One may also need to fine-tune separately for each attribute of interest. Ctrl Keskar et al. (2019) partially addresses these issues by providing 55 attribute control codes for a large language model trained from scratch, although this is expensive. Very recently, GeDi Krause et al. (2020) achieves strong performance by using CCLM generators as discriminators, though it relies on several heuristics. More broadly, text generation models for style transfer Hu et al. (2017); Lample et al. (2018b); Dai et al. (2019a), summarization See et al. (2017); Gehrmann et al. (2018); Zaheer et al. (2020), and machine translation Lample et al. (2018a); Ng et al. (2019); Lewis et al. (2019) can also be viewed as CCLM’s for different “attributes.”
A second type of approach instead conditions on a desired attribute by backpropagating gradients, either to directly modify model activations Dathathri et al. (2019); Liu et al. (2020) or to find a trigger string Wallace et al. (2019, 2020). Such methods often exhibit a high degree of attribute control, and can be used in adversarial attacks Wallace et al. (2020). In fact, Subramani et al. (2019) show that by carefully modifying the latent state, one can cause the base to produce arbitrary outputs.
A third class of methods, referred to as weighted decoding (WD), assumes access only to (i.e., ’s output logits), and operates directly on these logits Ghazvininejad et al. (2017); Holtzman et al. (2018); Cohn-Gordon et al. (2018); Shen et al. (2019). Compared to other approaches, WD methods are relatively interpretable in how they obtain from , but prior WD implementations have been observed to perform poorly in controlled text generation See et al. (2019); Dathathri et al. (2019). While Fudge shares a Bayesian motivation with other WD methods, Fudge follows the Bayesian factorization more closely in implementation (Sec. 3). The key distinguishing feature of Fudge is that it models whether attribute will be true in the future, rather than in the present. We find that Fudge substantially outperforms previous WD approaches in our experiments (Sec. 4.2).
Future Discriminators for Generation
We now explain the details of our proposed method, Future Discriminators for Generation (Fudge), and show that it corresponds to modeling the desired conditional distribution .
For a given language generation task, assume we have an autoregressive model (e.g., a large pretrained language model) which models for tokens . Letting denote a completed sequence, can sample from one token at a time by factoring :
To condition on attribute , we instead model . This requires a model for , modifying the previous factorization:
If we model directly, we obtain a class-conditional language model (CCLM). We can learn the CCLM by e.g., fine-tuning depending on the available data, possibly with some structural modification to to accommodate conditioning.
However, Fudge instead relies on the following Bayesian factorization, exchanging and conditioned on :
The second term is exactly the quantity modeled by the base . It then suffices to model the first term, , with a binary classifier for the attribute given a prefix . Intuitively, one can view as rescoring or reranking ’s original hypotheses.
We emphasize that although takes a prefix as input, it predicts whether attribute will in the future be satisfied for the completed generation . For instance, suppose we are given a dataset of examples with being the values of binary indicators for the desired (i.e., if is formality, then is 0 or 1 when is informal or formal respectively). For each training example , we train our classifier using all pairs ; that is, we construct a separate example from each prefix of . Our approach contrasts with previous methods such as Dathathri et al. (2019), which greedily optimize for on the immediate extension . One particular benefit is that Fudge naturally plans for the future: in the example for generating text on the “space” topic in Table 6, Fudge writes about a “mysterious ship” despite “ship” itself not being in the given “space”-topic bag of words, because “mysterious ship” easily leads into a mention of one of the targeted “space” words (“Earth”). Similarly, in the first couplet completion example in Table 3, Fudge needs to rhyme with “fear” after exactly ten syllables. After seven syllables, it could reasonably generate the word “clear,” but it first generates the adverb “pretty” in order to set up the generation of “clear” as the tenth syllable.
Fudge’s implementation is shown schematically in Figure 1, and is quite simple in practice. Fudge just needs to learn a (red in Figure 1) sharing tokenization with (dark blue). It then converts ’s output into probabilities (red table in Figure 1), and multiplies with the original output probabilities from (dark blue table), to obtain unnormalized probabilities (purple table). Finally, renormalizing over the output vocabulary yields the desired distribution . In practice, we operate in the log-probability space for numerical stability.
To improve computational efficiency, we typically choose to be lightweight relative to . We also consider only the top 200 possibilities for according to at each step, as a cheap approximation to the full distribution, and find that this works well in practice.See Appendix H for ablations on the top-200 pruning. In each task in Sec. 4, running Fudge on the test set takes no more than 15 minutes on a single Quadro RTX 6000 GPU.
Finally, as with other controlled generation approaches such as Dathathri et al. (2019), it is likely that augmenting Fudge with reranking approaches such as rejection sampling could improve output quality at the cost of compute time, although we do not comprehensively evaluate such extensions in this work.
We highlight several additional potential advantages of Fudge compared to directly modeling via e.g., a fine-tuned CCLM:
Fudge requires access only to (i.e., ’s output logits) rather than itself.
can be freely swapped out for any other model that shares the same tokenization when larger models become available.
Given multiple conditionally independent attributes with predictors for each, Fudge can easily condition on the combination of these attributes in a modular fashion by summing their output log-probabilities (Sec. 4.1, 4.2).
Unfortunately, like previous methods, Fudge cannot fully guarantee that all outputs possess the desired attribute . In Fudge’s case, this is due to the approximation inherent in modeling , as well as only considering the top 200 possible for computational efficiency.
Experiments
We run experiments on a range of controlled text generation tasks to evaluate the effectiveness of our proposed method: poetry couplet completion (Sec. 4.1), topic-controlled language generation (Sec. 4.2), and machine translation formality change (Sec. 4.3). For each task we discuss the evaluation setup, the specific details of our method and baselines, and finally experimental results.
We begin with English poetry generation, a task that emphasizes well-formedness, and which has been studied in different forms by many previous works Zhang and Lapata (2014); Wang et al. (2016); Ghazvininejad et al. (2016, 2017). Our task here is couplet completion. Given the first line of an iambic pentameter couplet (e.g., Table 1), the model must generate a second line which (1) satisfies iambic pentameter, (2) rhymes with the first line, and (3) ends a sentence. The desired attribute is defined as possessing all three properties, as evaluated by a rule-based checker (Appendix A). Our test set is a collection of prefix lines of couplets, collected from the ending couplet of each of Shakespeare’s 154 sonnets.
Success, the fraction of couplet completions with the desired attribute , as checked by . This is the main metric.
Grammaticality, the probability of grammaticality given by a Roberta-based CoLA grammaticality model Liu et al. (2019); Warstadt et al. (2019), averaged over all outputs.
Perplexity of the completion conditioned on the prefix. Following Dathathri et al. (2019), since our models use GPT2-Medium Radford et al. (2019) as , we evaluate perplexity using GPT Radford et al. (2018).See Appendix E for other perplexity measurements.
Distinctness of completions, measured as the number of unique unigrams, bigrams, and trigrams across all samples, divided by the total number of words Li et al. (2015).
At test time, we decode until the model generates ten syllables followed by an end-of-sentence punctuation mark, or after the eleventh syllable (an automatic failure, since iambic pentameter requires exactly ten syllables).
Overall, because we define using a rule-based which is accessible during training, our formulation of couplet completion is a relatively clean task for evaluating the effectiveness of Fudge.
Fudge Instantiation. The obvious approach is to learn a predictor for directly. However, the three components of — meter, rhyme, and sentence-ending — should be roughly independent. Thus we assume conditional independence, and demonstrate the modularity of Fudge by constructing three separate predictors to be combined at test time:
takes a text prefix , and predicts whether the completion of prefix will be in iambic meter. The model is an LSTM followed by a linear output layer.
takes prefix , the number of syllables between and for , and a rhyme sound .Two words have the same “rhyme sound” if they rhyme according to the CMU Pronouncing Dictionary Weide (1998). It predicts whether the completion has the rhyme sound at the end of token . The model is an LSTM with attention dependent on and , followed by a shallow feedforward network, and is trained via noise-contrastive estimation Gutmann and Hyvärinen (2010).The output logits from are unnormalized, but this does not affect Fudge after they are added to the output logits of and softmaxed for sampling.
takes prefix and the number of syllables between and for , and predicts whether ends a sentence. The model is an LSTM followed by a shallow feedforward network.
The predictors vary in architecture because and require inputs other than — in truth, they are families of related predictors. We find that performance is not overly sensitive to the particulars of the predictor architectures (Appendix D).
To train the discriminators, we sample a dataset of 10 million generations of varied length from GPT2-Medium. From these generations, we sample random subsequences of roughly 10 to 30 syllables and truncate ending syllables. These truncations become inputs to the predictors. For simplicity, we did not balance the class labels for e.g., the iambic predictor during training, although it is likely that doing so would improve performance.
At test time, we extract from the given first line of the couplet, and initialize , updating at each step. We then modify the output logits of by simply adding the log-probabilities from , , and , demonstrating the ease of composing constraints in Fudge.
Baselines. We compare to four baselines.A system like Hafez Ghazvininejad et al. (2016, 2017), which enforces meter and rhyme at each decoding step using a hard constraint, could achieve perfect success rate. However, this approach relies on the meter and rhyme attributes being “prefix-checkable” at the word level: one can guarantee success by simply never selecting a word which immediately violates the constraint. This is often the case for simple rule-based constraints, but not for many other interesting attributes, such as the topic and formality attributes in our subsequent experiments. To preserve generality, Fudge does not rely on this “prefix-checkable” property, and neither do our baselines.
Finetune, a CCLM which finetunes on similar inputs to those used for in Fudge. Since it is not obvious how to compose multiple CCLM’s for different attributes, we train a single CCLM for all desired properties together. We condition by prefixing the input with (1) whether the last 10 syllables of the original untruncated are iambic, (2) the rhyme sound at the end of , and (3) whether a sentence ends with . A special token is inserted 10 syllables from the end of .
Pplm Dathathri et al. (2019), which uses shallow predictors learned from ’s top-level hidden layer to modify ’s states toward increasing probability of the desired attribute via gradient ascent. We decompose the predictors into the same iambic, rhyme sound, and end-of-sentence predictors as for Fudge, inserting an additional hidden layer in the shallow predictor when needed to incorporate additional input (the desired rhyme sound and/or number of syllables until end-of-sentence).
Shakespeare’s original couplet completions.
All non-Shakespeare methods use top- sampling with .
1.2 Results
Even though our GPT2-Medium-generated training dataset is completely different from the test domain, and contains essentially zero examples of correct couplets, Fudge is able to learn the desired attribute. As shown in Table 2, Fudge greatly outperforms all automated baselines in success rate.
Surprisingly, the Pplm baseline achieves zero success. We find that its iambic and rhyme predictors are very poor, so we hypothesize that the relevant information is not easily extractable from the last hidden layer of . In contrast, Fudge’s predictors operate directly on the raw text.
Funnily enough, Fudge even matches Shakespeare according to , although this is largely due to the narrowness of and should not be taken seriously. We define using somewhat narrow criteria (Appendix A), which capture only a subset of what Shakespeare considered to be well-written couplets. The purpose of this task is to evaluate Fudge’s ability to satisfy a difficult well-formedness constraint compared to automated baselines, rather than to perfectly capture the human notion of an iambic pentameter couplet. Thus Shakespeare is marked wrong when he (1) uses archaic pronunciations, (2) uses loose rhymes, (3) elides syllables to fit meter, or (4) uses words missing from the CMU Pronouncing Dictionary. See Appendix A.1 for details. Of course, Shakespeare is only included as a whimsical point of reference; our generations obviously do not hold a candle to Shakespeare’s originals. Similarly, the grammaticality and perplexity metrics are designed for our automated baselines, and thus assign poor scores to Shakespeare’s antiquated and flowery style.
Fudge also maintains relatively fluent generation despite lower grammaticality and perplexity compared to . See Table 3 for two successful examples. Interestingly, Fudge also increases diversity compared to , perhaps due to the difficult constraint forcing Fudge to use lower-probability regions of the base distribution .
Finally, it is possible (and trivial) to adjust the conditioning strength in Fudge by multiplying the binary predictors’ output logits by a constant. However, this deviates from our Bayesian factorization of , and we do not do so.
2 Topic-Controlled Language Generation
Next, we explore topic control in English language generation. The desired attribute is to be on-topic for a given topic, such as science or politics. To facilitate comparison with prior work, we largely follow the setup of Pplm Dathathri et al. (2019): the model is provided an approximation to the topic at test time, in the form of a bag of on-topic words . The goal is to sample text according to the topic approximated by , starting from a generic prefix. There are 7 topics (space, politics, military, legal, science, religion, and computers) and 20 prefixes, and the model generates 3 80-tokenAll models and baselines use GPT2 tokenization. samples from each topic-prefix pair, for a total of 420 generations.
Metrics. Unfortunately, we cannot easily construct a rule-based for being “on-topic.” Additionally, use rate of words in is a poor metric, because a model can score highly by e.g., simply returning the words in , without generalizing to the full topic that approximates. Instead, we adopt a notion of success which requires the model to generalize the bag to the full topic. The remaining metrics are measures of quality and diversity.
Success, the average number of distinct words in a heldout bag which appear in the model output. Specifically, for each word in , we add to the closest GloVe Pennington et al. (2014) word by cosine similarity, such that the new word does not contain (and is not contained by) any word in . (This excludes e.g., most plurals.) Usage of distinct words in measures the model’s ability to generalize to other on-topic words, of which is a non-exhaustive set. This is our main metric.
Grammaticality, identical to the couplet task.
Perplexity, identical to the couplet task.
Distinctness, defined as in the couplet task. However, it is calculated separately within the 60 generations for each topic, and then averaged over the 7 topics.
Additionally, following the evaluation procedure of prior work such as Dathathri et al. (2019), we run human evaluations via Amazon Mechanical Turk for Fudge against each baseline, comparing topic control and fluency. For each pairwise comparison, we ask 3 workers to evaluate each of 420 paired outputs. Workers were asked to mark which generation is more on topic (first, second, both, or neither), and to rate each generation’s fluency on a Likert scale from 1 to 5. We report the average fraction of outputs marked as on-topic as well as the average fluency rating for each method.
Fudge Instantiation. Since we model topics as bags of words, Fudge uses a binary predictor which takes a prefix and word , and classifies whether appears in the future for . (Since it is desirable to stay on topic even after successfully getting on topic, we use rather than .) Training examples are sampled from the same dataset of 10 million GPT2-Medium generations used for the couplet task, and is trained using noise-contrastive estimation. is a lightweight LSTM-based classifier similar to from the couplet task.
At test time, we can compose individual-word constraints if we assume conditional independence between words (although this may be imperfect). Given a bag of words and prefix , we could condition on all words in the bag appearing in the future by adding all log-probabilities to ’s logits. However, topic control does not require every word to appear; perhaps some number of on-topic words is enough to be “on-topic.” Therefore, we model the topic constraint as selecting a random subset of words from the original bag, and requiring that only those words all appear. Since each of the words is selected with probability , the quantity we add to the base logits is in expectation. In our experiments we use , based on a fantasy-topic bag of words used for validation (Appendix C).
Finetune, which finetunes on the same inputs used for Fudge. The future word is given as a prefix for conditioning. At test time, we compute logits for each prefix in the given and use the average as the true logits, as an ad hoc way to condition on the full .
Wdec, a simple weighted decoding implementation which greedily considers only the immediate next token when optimizing for . Instead of using , Wdec just adds a fixed to the logit for each word in . Note Wdec requires to be well-defined at the token level, so it is not easily transferable to certain tasks (e.g., couplet completion).
Pplm Dathathri et al. (2019), which modifies the activations of to make the desired bag of words more likely at the immediate next position. We use their method without reranking for fair comparison.
All methods use top- sampling with , following Dathathri et al. (2019)’s setup.
2.2 Results
Fudge achieves the highest success by a substantial margin (Table 4), and outperforms all baselines on human evaluations in both topic relevance and fluency (Table 5). Fudge simultaneously preserves high quality and diversity according to automated metrics. Table 6 shows two examples.
Unsurprisingly, performs poorly on success. Wdec and Finetune also perform poorly, in success and especially in distinctness. Wdec frequently degenerates into repeating the given words in the bag , despite tuning (Appendix C). Finetune also suffers from repetition, which appears to be the result of distribution shift from fine-tuning. Our fine-tuning dataset was built by sampling directly from the original modeled by to mitigate distribution shift, but it is well-known that language model generations are more repetitive than natural language Holtzman et al. (2018, 2019). We hypothesize that Finetune, being fine-tuned on language model generations rather than natural language, amplifies this repetitiveness. This repetition is reflected in the poor grammaticality for both Finetune and especially Wdec. In contrast, Fudge does not touch the original , largely avoiding Finetune’s distribution shift problem on this task.
Finally, Fudge outperforms the strong gradient-based Pplm method, despite requiring access only to ’s output logits. Non-reliance on gradients means Fudge is also many times faster than Pplm, which takes a few hours compared to Fudge’s 15 minutes for the full set of 420 generations on our hardware. Sometimes we do not even have gradients: for example, gradients are unavailable in the API for GPT3 at time of writing.
3 Machine Translation Formality Change
Finally, we turn to a somewhat more challenging task, changing formality in machine translation — specifically, from informal to formal. Given a source sentence written in an informal and conversational style, the goal is to output a translation which is also more formal. We test on the Fisher and CALLHOME Spanish–English Speech Translation Corpus Post et al. (2013), a collection of transcribed Spanish conversations with English translations. Both the source Spanish and target English are highly informal and disfluent. Salesky et al. (2019) augment the Fisher dataset with additional parallel English translations, rewritten to be more fluent (and hence more formal); see Table 7 for an example. Our task is to translate the original informal Spanish to into more formal English. However, we assume that Salesky et al. (2019)’s fluent references are unavailable during training.
Metrics. The desired attribute is formality, but we cannot sacrifice the source sentence’s meaning. The latter requirement makes generation more constrained than in the couplet and topic tasks, so perplexity and distinctness are less relevant. Instead, we use the following:
BLEU Score Papineni et al. (2002), using two of Salesky et al. (2019)’s fluent references per test example. This is our main metric.
Formality, the average probability that the model’s outputs are formal, according to an evaluator trained on the Family/Relationships domain of the GYAFC formality dataset Rao and Tetreault (2018). The evaluator is an LSTM followed by a linear layer.
Fudge Instantiation. We assume that the attribute , formality, is conditionally independent from the original conditioning in , i.e., the meaning of the Spanish input. Fudge uses a binary predictor which classifies whether the text starting with prefix is written in a formal style. is an LSTM followed by a linear layer, trained on the Entertainment/Music domain of GYAFC.
At test time, Fudge directly augments ’s logits using log-probabilities from . is a pretrained Marian Junczys-Dowmunt et al. (2018) transformer model for Spanish-English. We evaluate both when is fine-tuned on the original Fisher training dataset (i.e., using the original targets, not Salesky et al. (2019)’s more fluent targets) as well as zero-shot with no fine-tuning, which is challenging due to the highly informal and disfluent text.
, the original machine translation model.
+ st, a pipeline consisting of followed by a style transfer model. Our style transfer model is T5 Raffel et al. (2020), fine-tuned on the same GYAFC Entertainment/Music domain that we used to train in Fudge.
Since we do not assume access to Salesky et al. (2019)’s more formal targets during training, it is difficult to apply Pplm to this task: Pplm’s predictor would operate on the pretrained translation model’s hidden states, thus requiring a Spanish-English translation dataset with both formal and informal English.We nevertheless ran Pplm in a somewhat convoluted setup, but found that it performed poorly (Appendix B). We omit Finetune for the same reason. In contrast, Fudge requires only the original English dataset with formality annotations.
3.2 Results
As shown in Table 8, Fudge increases the formality of outputs compared to , even though the test-time formality predictor is trained on a different domain (Family/Relationships, rather than Entertainment/Music). Note that formality unsurprisingly decreases after fine-tuning , simply due to the informality of the fine-tuning dataset. As in the couplet task, one could adjust the strength of the formality control in Fudge, although this is unprincipled from the view of modeling .
Moreover, while Fudge and achieve similar BLEU after fine-tuning , Fudge achieves higher BLEU compared to when is not fine-tuned on the Fisher training set. In the latter case, controlling for formality somewhat remedies the struggles of when not fine-tuned on such disfluent text.
In contrast, the + st baseline achieves near-perfect formality but less than half the BLEU of , due to the style transfer model overfitting to the GYAFC Entertainment/Music dataset. This is similar to the distribution shift issue that we observed in topic control for Finetune, an issue which Fudge largely avoids. Nevertheless, there remains substantial room for improvement on this difficult task.
Discussion
Fudge is a principled approach to controlled text generation which models by closely following a Bayesian factorization, thus preserving the base as much as possible. Fudge achieves strong performance on a wide range of different tasks: poetry couplet completion, topic control, and informal-to-formal machine translation. Additionally, Fudge can easily compose different attributes in a modular fashion: the meter, rhyme, and end-of-sentence constraints for couplet completion, and the individual words within each topic bag for topic control. In principle, Fudge is applicable to any controlled generation task where we can train discriminators for the desired attribute or attributes.
Ethics of Controlled Text Generation
We recognize that strong controlled generation methods have the potential to produce harmful outputs and/or misinformation when used adversarially Wallace et al. (2019, 2020). However, such methods can also be a powerful tool for mitigating harmful biases learned by large pretrained language models Radford et al. (2019); Brown et al. (2020), for example by detoxifying language Dathathri et al. (2019); Krause et al. (2020). Overall, we believe it is still beneficial to continue research into general controlled text generation methods such as Fudge.
Acknowledgements
We thank Daniel Fried, David Gaddy, Eric Wallace, Kevin Lin, Nicholas Tomlin, Ruiqi Zhong, and the three anonymous reviewers for their helpful comments and feedback, which aided us in greatly improving the paper. We also thank the authors of Dathathri et al. (2019) for clarifying our questions about their topic control setup. This work was supported by Berkeley AI Research, DARPA under agreement HR00112020054, and the NSF through a fellowship to the first author. The content does not necessarily reflect the position or the policy of the government, and no official endorsement should be inferred.
References
Appendix A Details of ℱℱ\mathcal{F} for Couplet Completion
We provide the full details of the function we use to check iambic pentameter, rhyme, and sentence-ending in our couplet completion task. Note that iambic pentameter consists of two components: iambic meter as well as containing exactly ten syllables.
Iambic meter: Given a phrase, we obtain the sequence of stresses (0 for unstressed, 1 for stressed, 2 for secondary stress) for each word, according to the CMU Pronouncing Dictionary Weide (1998). If any word does not exist in the dictionary (almost never for non-Shakespeare methods) we return False. We treat 2 as ambiguous stress, and additionally change 1 to 2 for any monosyllabic words, i.e. we allow monosyllabic stressed words to be unstressed but not vice versa. Finally, we check that all syllables at even indices (0-indexed) are unstressed or ambiguous, and all syllables at odd indices are stressed or ambiguous.
Number of syllables: We count the number of syllables in each word based on the number of stresses according to the CMU Pronouncing Dictionary. If a word does not exist in the dictionary, we estimate the number of syllables by rounding the number of letters divided by 3 to the nearest integer.
Rhyme: Two words rhyme if and only if they both exist in the CMU Pronouncing Dictionary and are a perfect rhyme according to the dictionary.
Sentence-ending: We check if the output ends with a period, question mark, or exclamation mark.
Of course, both Fudge and Finetune will fit to whatever output is given by . The purpose of the couplet task is to check Fudge’s ability to fit a difficult well-formedness constraint. We simply design an that corresponds to true iambic pentameter rhymes in most cases.
Shakespeare himself performs somewhat poorly according to , which is designed with the automated baselines in mind, not for Shakespeare. (The same is true for our grammaticality and perplexity metrics.)
One source of error is words which are out-of-vocabulary for the CMU Pronouncing Dictionary. Such words are almost never generated by either Fudge or our automated baselines, but appear in a fifth of Shakespeare’s lines, resulting in failures on the iambic meter and syllable checks.
Nevertheless, most of Shakespeare’s “errors” are the result of real — though slight — deviations from our very strict definitions of meter and rhyme. In particular, he frequently (1) elides syllables to fit meter, and (2) uses loose rhymes; both “error” types are likely exacerbated by differences between archaic and modern pronunciations. The example in Table 10 illustrates both types of “errors.” Although such deviations are often acceptable to a human, they are difficult to capture in an automatic metric, and we do not allow such deviations in . Again, Shakespeare is only included as a whimsical point of reference, and not as a serious baseline to be compared to.
Appendix B Pplm Baseline in Machine Translation
As discussed in the main text, it is difficult to apply Pplm in our machine translation setup, in which is learned from an English formality dataset without parallel Spanish. Since is a Spanish-English translation model, we must obtain hidden states for training Pplm’s by first “backtranslating” English into Spanish, accessing a second pretrained translator. For this purpose we use a second pretrained Marian transformer from HuggingFace (https://huggingface.co/Helsinki-NLP/opus-mt-en-es). Additionally, we needed to tune their suggested hyperparameters.
During evaluation, we observe that Pplm makes some reasonable modifications for formality compared to the base , like changing “hard” to “difficult,” but such improvements are also accompanied by occasional disfluencies and/or repetitions (although such problems plague all methods to some degree). Overall, while Pplm achieves similar BLEU to Fudge, it is substantially less formal (Table 11).
Appendix C Hyperparameter Choices
Fudge has essentially one hyperparameter in our topic control task, , which controls the strength of conditioning and corresponds to the number of words in the bag which should appear in the future.
To choose in topic control, we used a separate validation bag of words (on the topic of fantasy; Appendix K.4) to select a reasonable for our main paper experiments (). Unlike in the main paper where we use heldout bags to measure success, during validation we simply use the original bag. We use a set of 60 generations, considering values ranging from 1 to 6 (Table 12), although the result may be somewhat noisy. Of course, different choices of result in different tradeoffs (Appendix G).
We also optimized the conditioning strength for the Wdec baseline on the same fantasy bag of words, considering values ranging from 1 to 32. We selected the only value (4) which achieved reasonable success without a total collapse in diversity (Table 13), but diversity still collapsed when tested on our seven main test bags of words.
We do not optimize any model hyperparameters in the couplet completion and informal-to-formal translation tasks. LSTM’s and feedforward networks are 3 layers (including the output layer of dimension 1) and 300-dimensional unless otherwise specified. They are bidirectional (150-dimensional in each direction) for the couplet rhyme predictor and the topic control future words predictor, and otherwise unidirectional. Attention mechanisms use key-query-value attention. For the rhyme and future words predictors the output hidden state is multiplied element-wise by the embedding of the rhyme sound or future word, then concatenated to the embeddings, before the final feedforward network. Since a selling point of our method is the lightweight process of constructing and training predictors, noise-contrastive estimation is a natural choice for the rhyme and future word predictors: we avoid softmaxing over the output dimension during training. (This is primarily relevant for the future word predictor, as the number of distinct rhyme sounds is not too large, but we use noise-contrastive estimation for both for consistency’s sake.)
For the Pplm baseline, we used step size 0.01 for both couplet completion and MT after tuning, and kept their other hyperparameters fixed. For topic control we simply evaluated their provided generations instead of rerunning their model.
Appendix D Ablations on Predictor Architectures
Some variation in predictor architectures is necessary due to the diversity of our tasks (as evidenced by the difficulties in adapting PPLM). Specifically, while our core predictor architecture is word embeddings followed by LSTM and output layer, task-specific architectures vary because some “predictors” are actually families of related predictors. We model such families as a single predictor taking additional input (e.g., rhyme sound in poetry); this is needed in our poetry and topic tasks.
On these two tasks, we provide ablations with more homogenized predictors: additional inputs are simply embedded and concatenated to each input word embedding. The difference is relatively small in both cases (Tables 14 and 15). FudgeMod indicates the ablated version of Fudge.
Appendix E Alternative Perplexity Measurements
On the couplet completion task, we additionally measure perplexity using Transformer-XL Dai et al. (2019b) and using a GPT model fine-tuned on Shakespearean language as generated by Lau et al. (2018). We measure using Transformer-XL on the topic control task as well. Relative perplexities between most models remain largely similar when switching between GPT and Transformer-XL, with a few exceptions. Compared to the base GPT, Shakespeare’s perplexity naturally decreases while other models’ perplexities increase when measured with Shakespeare-finetuned GPT. The highly repetitive and disfluent Wdec baseline is rightly punished for this behavior when measured by Transformer-XL. Pplm also obtains slightly lower perplexity than Fudge on topic control when measured by Transformer-XL. Full results in Tables 16 and 15.
Appendix F Statistical Significance
In couplet completion, Fudge outperforms the strongest automated baseline (Finetune) on success rate with on a McNemar test, pairing the generations for each Shakespeare prefix.
In topic control, Fudge outperforms the strongest automated baseline Pplm with using a Wilcoxon matched pairs test, pairing the generations for topic-prefix combinations.
In translation formality, Fudge’s generations are more formal than those of the base with according to a paired t-test.
Appendix G Effect of Varying Topic Control Strength
Although we use for Fudge in our main paper experiments for topic control, we experiment here with varying the conditioning strength. Specifically, we experiment with and . The conditioning is unsurprisingly stronger as increases, as shown quantitatively in Table 19, although the perplexity increases as well.
We also provide some example generations for and in Tables 18 and 21, for the same prompts and topics as in Table 6 for in the main text. The generations remain mostly fluent and interesting, despite their worse grammaticality and perplexity.
Appendix H Effect of Varying Candidate Pruning
For computational efficiency, we only feed the top 200 candidates returned by into Fudge’s predictor when predicting each next token. Here, we ablate on this number in our topic control setting, testing 100 and 400 (Table 20).
Appendix I Additional Couplet Completion Examples
We provide some additional examples of Fudge and baselines on our couplet completion task in Table 22.
We also show some unsuccessful examples for Fudge in 23. Overall, we find that most errors are due to the rhyme and ten-syllable end of sentence constraints, or due to Shakespeare’s prefix ending in a word not in the CMU Pronouncing Dictionary (e.g., “prognosticate” in the table). Fudge also sometimes overgenerates punctuation at the end of a sentence.
Appendix J Additional Topic Control Examples
In Tables 24, 25, and 26 we show additional example generations by our method using the same hyperparameter setting as in the main paper, . Specifically, we provide the first generation by Fudge for 3 separate prefixes for each of the 7 topics. Virtually all examples are clearly on topic, while avoiding repetitiveness.
Additionally, we provide example generations from , Finetune and Wdec in Tables 27, 28, and 29 respectively. For Pplm we refer the reader to the examples in the main paper and appendices of Dathathri et al. (2019)’s original work.
Appendix K Topic Control Bags of Words and Prefixes
We use the exact same bags of words and prefixes as Dathathri et al. (2019) for their topic control task, with non-proper nouns lower-cased (in practice, this only changes the religion wordlist). Note our success metric in the paper matches without casing.
We additionally provide the heldout bags of words computed from the original bags (before lower-casing), which we use for the success metric. Although a few words deviate somewhat (“actress” as a synonym for “star” in the space category), overall the heldout bags do represent the desired topic.
Finally, we provide the fantasy bag of words used for selecting the and conditioning strengths for Fudge and Wdec respectively. It is also taken directly from Dathathri et al. (2019).
Space: planet, galaxy, space, universe, orbit, spacecraft, earth, moon, comet, star, astronaut, aerospace, asteroid, spaceship, starship, galactic, satellite, meteor
Politics: affirm, appropriation, aristocracy, authoritarian, authority, authorization, brief, capitalism, communism, constitution, conservatism, court, deficit, diplomacy, direct, democracy, equality, exports, fascism, federation, government, ideology, imports, initiative, legislature, legitimacy, liberalism, liberty, majority, order, political, culture, politics, power, primary, property, ratification, recall, referendum, republic, socialism, state, subsidy, tariff, imports, tax, totalitarian
Military: academy, advance, aircraft, ally, ammo, ammunition, armor, arms, army, arrow, arsenal, artillery, attack, attention, ballistic, barracks, base, battalion, battery, battle, battlefield, bomb, bombard, bombardment, brig, brigade, bullet, camouflage, camp, cannon, captain, capture, carrier, casualty, catapult, cavalry, colonel, combat, command, commander, commission, company, conflict, conquest, convoy, corps, covert, crew, decode, defeat, defend, defense, destroyer, division, draft, encode, enemy, engage, enlist, evacuate, explosive, fight, fire, fleet, force, formation, fort, front, garrison, general, grenade, grunt, guerrilla, gun, headquarters, helmet, honor, hospital, infantry, injury, intelligence, invade, invasion, jet, kill, leave, lieutenant, major, maneuver, marines, MIA, mid, military, mine, missile, mortar, navy, neutral, offense, officer, ordinance, parachute, peace, plane, platoon, private, radar, rank, recruit, regiment, rescue, reserves, retreat, ribbon, sabotage, sailor, salute, section, sergeant, service, shell, shoot, shot, siege, sniper, soldier, spear, specialist, squad, squadron, staff, submarine, surrender, tactical, tactics, tank, torpedo, troops, truce, uniform, unit, veteran, volley, war, warfare, warrior, weapon, win, wound
Legal: affidavit, allegation, appeal, appearance, argument, arrest, assault, attorney, bail, bankrupt, bankruptcy, bar, bench, warrant, bond, booking, capital, crime, case, chambers, claim, complainant, complaint, confess, confession, constitution, constitutional, contract, counsel, court, custody, damages, decree, defendant, defense, deposition, discovery, equity, estate, ethics, evidence, examination, family, law, felony, file, fraud, grievance, guardian, guilty, hearing, immunity, incarceration, incompetent, indictment, injunction, innocent, instructions, jail, judge, judiciary, jurisdiction, jury, justice, law, lawsuit, lawyer, legal, legislation, liable, litigation, manslaughter, mediation, minor, misdemeanor, moot, murder, negligence, oath, objection, opinion, order, ordinance, pardon, parole, party, perjury, petition, plaintiff, plea, precedent, prison, probation, prosecute, prosecutor, proxy, record, redress, resolution, reverse, revoke, robbery, rules, sentence, settlement, sheriff, sidebar, standing, state, statute, stay, subpoena, suit, suppress, sustain, testimony, theft, title, tort, transcript, trial, trust, trustee, venue, verdict, waiver, warrant, will, witness, writ, zoning
Science: astronomy, atom, biology, cell, chemical, chemistry, climate, control, data, electricity, element, energy, evolution, experiment, fact, flask, fossil, funnel, genetics, gravity, hypothesis, lab, laboratory, laws, mass, matter, measure, microscope, mineral, molecule, motion, observe, organism, particle, phase, physics, research, scale, science, scientist, telescope, temperature, theory, tissue, variable, volume, weather, weigh
Religion: absolute, affect, aid, angel, anthem, apostle, archangel, Archbishop, balance, ban, belief, benefit, Bible, bishop, bless, blessing, bliss, bond, bow, Buddhism, canon, Cantor, cathedral, celestial, chapel, charity, choice, Christianity, church, comfort, community, conflict, connection, conquest, conservative, control, conversion, convert, core, counsel, courage, Covenant, creative, Creator, creed, cross, Crusade, Darkness, decision, deity, destiny, Devil, disciple, discipline, discussion, divine, divinity, doctrine, duty, effect, elder, energy, essence, eternal, ethics, event, evidence, exile, Exodus, faith, family, fate, Father, favor, fundamental, gift, glory, God, gospel, grace, growth, guru, habit, hallow, halo, happiness, harmony, healing, Heaven, Hebrew, holy, honor, hope, host, humane, immortal, influence, insight, instruction, issue, Jesuit, Jesus, joy, Judaism, judgment, justice, karma, keen, Keystone, Kingdom, Latin, life, light, love, loving, marriage, meaning, mercy, Messiah, minister, miracle, mission, mortal, mosque, movement, music, mystery, nature, nun, official, oracle, order, organ, Orthodox, outlook, pacific, pagan, parish, participation, pastor, patriarch, peace, perception, personal, perspective, petition, pilgrim, politics, power, practice, prayer, prelude, presence, priest, principle, privacy, prophet, protection, purpose, query, quest, question, quiet, radiant, radical, rally, rebirth, redemption, refuge, relationship, relative, religion, religious, Revelation, ritual, role, Sacrament, sacred, sacrifice, sage, saint, salvation, sanctuary, savior, scripture, scriptures, sect, security, sense, serious, serve, service, Sharia, shepherd, shrine, silence, sin, society, soul, source, spirit, spiritual, split, statue, Sunday, support, Supreme, teaching, temple, tests, text, Torah, tradition, traditional, trust, unique, unity, unknown, value, vanity, virtue, vision, voice, voices, watch, weight, whole, wisdom, wonder, yang, yin, zeal
Computers: algorithm, analog, app, application, array, backup, bandwidth, binary, bit, bite, blog, blogger, bookmark, boot, broadband, browser, buffer, bug, bus, byte, cache, caps, captcha, CD, client, command, compile, compress, computer, configure, cookie, copy, CPU, dashboard, data, database, debug, delete, desktop, development, digital, disk, document, domain, dot, download, drag, dynamic, email, encrypt, encryption, enter, FAQ, file, firewall, firmware, flaming, flash, folder, font, format, frame, graphics, hack, hacker, hardware, home, host, html, icon, inbox, integer, interface, Internet, IP, iteration, Java, joystick, kernel, key, keyboard, keyword, laptop, link, Linux, logic, login, lurking, Macintosh, macro, malware, media, memory, mirror, modem, monitor, motherboard, mouse, multimedia, net, network, node, offline, online, OS, option, output, page, password, paste, path, piracy, pirate, platform, podcast, portal, print, printer, privacy, process, program, programmer, protocol, RAM, reboot, resolution, restore, ROM, root, router, runtime, save, scan, scanner, screen, screenshot, script, scroll, security, server, shell, shift, snapshot, software, spam, spreadsheet, storage, surf, syntax, table, tag, template, thread, toolbar, trash, undo, Unix, upload, URL, user, UI, username, utility, version, virtual, virus, web, website, widget, wiki, window, Windows, wireless, worm, XML, Zip
K.2 Prefixes
"An illustration of", "Emphasised are", "Foundational to this is", "Furthermore,", "In brief,", "In summary", "In this essay", "It has been shown", "More importantly,", "Prior to this", "The central theme", "The connection", "The issue focused on", "The key aspect", "The relationship", "This essay discusses", "To conclude,", "To review,", "To summarise", "Views on"
K.3 Heldout Bags of Words
Note that our heldout bag construction process yielded two stopwords, which we removed; they are omitted below.
Space: actress, aeronautics, broadband, cosmonaut, cosmos, fireball, flyby, galaxies, heavens, interstellar, lander, lunar, mothership, Romulan, room, worlds
Politics: appropriated, aristocrats, authorisation, autocratic, capitalist, communist, credibility, cultural, democratic, diplomatic, efforts, energy, excise, exporting, fascist, federal, federated, freedom, gender, ideologies, immediate, imported, income, judge, jurisdiction, legislative, lengthy, minority, Nazism, progressivism, properties, purchase, ratify, referenda, remember, secondary, shortfall, socialist, subsidies, uphold
Military: aboard, academies, adjutant, advancing, airmen, allies, argue, armies, armistice, armour, armoury, assets, ATL, aviation, barrage, batteries, bleeding, bottom, bricks, cadre, camera, capturing, cargo, casing, casualties, citadel, civilian, civilians, clandestine, committee, companies, concern, conquered, cursor, customer, dead, decoding, defensive, deputy, detonated, dormitories, encoding, enemies, engaging, escorting, evacuating, execute, expert, explosion, fatigues, flames, flying, forcing, forming, fought, freedom, frigate, gatling, glider, groan, guerilla, hand-to-hand, highest, hires, honour, howitzer, ICBM, injuries, inundate, invading, Iraq, khaki, knowledge, lace, late, launchers, leaving, lob, longtime, Maj., manoeuvre, medical, militia, naval, offensively, offices, operation, paragraph, personnel, persuade, pirate, pistol, policeman, propel, proposal, public, pump, rear, relinquish, rescuing, rifle, rifleman, riflemen, rifles, rocket, sabotaging, samurai, scouts, secluded, seige, Sgt., ship, shoulders, significant, skipper, skirmish, sloop, sonar, stationed, strategic, strategy, subsidiary, sunk, sword, taken, team, tensions, terms, threat, tribute, victory, visor, wear, won, zone, zoning
Legal: accusation, acquittal, admit, aggrieved, agreement, alleging, amendment, appearing, appellant, asserted, assertion, assualt, authority, burglary, championship, convicted, conviction, criminal, custodial, debatable, decision, defensive, democratic, deposited, deputies, disagree, discoveries, dispute, disputes, edict, embezzlement, enforcement, ethical, event, exams, families, federal, felonies, findings, folder, forgive, heard, homicide, immune, incarcerated, inept, injunctive, inmates, innocence, insolvency, insolvent, investment, judgment, judicial, jurors, knowing, land-use, leave, legislative, liability, litem, maintain, major, malpractice, mandamus, mediator, mutual, negligent, objecting, offender, pants, parties, passageways, pixels, police, property, prosecuting, prosecution, proxies, purchase, quash, regulations, repress, requesting, rescind, reservation, respondent, restaurant, revelation, reversing, rulings, second-degree, sentencing, sitting, solicitor, statutory, step-by-step, sued, sworn, testify, track, transcribed, treasurer, waived, whether, widget, wrongful, wrongs
Science: action, astronomical, bacterium, bone, clinical, component, compounds, electron, electrons, evolved, flow, fuels, genomics, gravitational, humidity, hypotheses, idea, increasing, jug, ligand, magnesium, mathematics, measuring, microscopy, molecular, nothing, observatory, observing, parameter, phone, physicist, physiology, pounds, rain, reason, renewable, scaling, scientific, siphon, statutes, stored, studies, system, tests, theories, transition, warming
Religion: Adventure, Almighty, Always, Answer, Appeals, Aramaic, Assistance, Association, Atlantic, Attorney, Balancing, Baptist, Basilica, Baskets, Best, Buddhist, Bunyan, Calculator, Calvary, Catholic, Catholicism, Charitable, Charities, Chen, Cognition, Communities, Compassion, Connery, Constantinople, Contemporary, Cosmic, Cost, Court, Creativity, Criminal, Crisis, Cure, Curriculum, Dangerous, Database, Date, Death, Deities, Demon, Determining, Dharma, Diocese, Double, Dreams, Echoes, Economic, Elegant, Emanuel, Empires, EOS, Episcopal, Epistle, Ethical, Eucharist, Everlasting, Excel, Existence, Factors, Fallen, Families, Fervor, Focus, Foods, Forums, Freedom, Glad, Glorious, Heart, Heavy, Hell, Help, Him, Honour, Hospital, Hypothesis, Impact, Implications, Influencing, Injunction, Intel, Invitations, Involvement, Jewish, Judas, Judgement, Kenichi, Kiss, Kombat, Lamp, Laughter, Learning, Leviticus, Liberal, Liberation, Lisa, Lives, Lord, Loss, Lust, Maker, Mandir, Marital, Married, Mary, Masjid, Meditation, Melody, Merrell, Metatron, Methodist, Militant, Mind, Mirror, Modernity, Morality, Mother, Motivation, Muhammad, Mutual, Mysteries, Mystical, Nanak, Natural, Network, ODST, Oneness, Outreach, PDF, Piano, Policy, Political, Pope, Practicing, Praise, Preview, Prime, Prostitute, Provider, Punishment, Purchase, Pure, Qi, Queries, Radiance, Rallies, Reiki, Reincarnation, Remote, Renewable, Resurrection, Rev., Rites, Safety, sanctuaries, Saturday, Saviour, Scrolls, Sculpture, Secular, Secure, Self, Sermon, Serving, Shadows, Shari’a, Shinto, Significance, Silent, Sonata, Songs, Spangled, Spanish, SPCA, SQL, St., Stevie, Suites, Supply, Sweet, SWF, Talmud, Templar, Terrier, Testament, Testing, Thank, Theology, Thyme, Tie, TransCanada, Truth, Uncharted, Understanding, United, Venue, Videos, VoIP, Volume, Vote, Wetlands, Wiccan, Worship
Computers: 512MB, allows, Android, article, attribute, autocomplete, automatically, back-up, barcode, beach, binaries, button, C++, caching, cake, camera, capabilities, casing, chairs, change, cheat, chew, choice, click, coder, compiling, components, computation, computing, confidentiality, configuring, connections, console, copies, counterfeiting, creating, crucial, customer, CyanogenMod, cyber, debian, decimal, decrypt, decryption, deflate, deleting, demo, detect, developing, dialog, dialup, direction, disc, display, DNS, DSL, DVD, e-mail, edit, educational, elements, encyclopedia, Excel, execute, extract, Firefox, fixes, flames, Frequently, functionality, gamepad, garbage, glass, gmail, GPU, guest, hats, house, identifier, infected, initialize, inkjet, input, interactive, interview, ISP, iterative, Jacket, journalists, jQuery, latency, layout, little, logon, lurks, Macs, mails, mainboard, memories, mice, must, notebook, off-line, on-line, operand, original, overflow, packet, pane, paper, parasite, parsing, pasting, PDF, phishing, php, pixels, point-in-time, popup, prev, profit, pull, query, reasoning, rectangular, redirect, rename, restart, restoring, run-time, sailor, saving, search, secure, shoe, sidebar, signal, sites, Solaris, spyware, step, stored, storing, subdirectory, taxi, telephoto, text, tools, topic, torrent, touchpad, typeface, Ubuntu, update, usb, utilities, visuals, VPN, wifi, workstation, writer, XP, XSLT
K.4 Validation Fantasy Bag of Words
beast, Cerberus, demon, dragon, fairy, Frankenstein, ghost, giant, Godzilla, horror, hydra, imp, monster, mummy, ogre, orc, savage, spirit, sprite, titan, troll, undead, unicorn, vampire, witch, zombie
Appendix L Additional Machine Translation Formality Examples
We provide some additional examples of Fudge against baselines on our machine translation formality task in Table 30.
Appendix M Software
All models are implemented in PyTorch Paszke et al. (2019), and pretrained models are obtained from HuggingFace Wolf et al. (2019). Specifically, the Marian translation model is https://huggingface.co/Helsinki-NLP/opus-mt-es-en.