Controlled Text Generation as Continuous Optimization with Multiple Constraints
Sachin Kumar, Eric Malmi, Aliaksei Severyn, Yulia Tsvetkov
Introduction
Recent advances in language models trained on large-scale web text corpora have led to great improvements in state-of-the-art on many natural language processing (NLP) tasks including the ability to generate increasingly coherent text . However, once such models are trained, they are prone to degeneration and biased, non-factual outputs as it is difficult to control the characteristics or attributes of the generated text without architectural modifications and fine-tuning the models on attribute-specific corpora . This can be even more challenging if multiple attributes are involved as labeled data for each combination of attributes can be difficult to obtain.
We focus on controlled text generation where the goal is to decode from a text generation model such that the outputs satisfy certain constraints, which the model was not necessarily trained on. For example, given a dialogue generation model, additionally constraining the generated responses to be polite, although the model was not optimized for politeness during training. Recent works address this problem with left-to-right autoregressive decoding, and modify the vocabulary distribution at every step directly using attribute classifier probabilities , or indirectly via backpropagating gradients through model activations . While exhibiting high level of attribute control, by design these methods can only work with categorical attributes (typically only one attribute) and condition only on the left context while decoding. Additionally, they often require several heuristics to work and are prone to adversarial outputs .
To address these concerns, we propose the following decoding algorithm. Given a pretrained language model, we posit decoding from it as an optimization problem. First, we relax this discrete optimization problem to a continuous one by representing each token as a simplex on the target vocabulary . This allows to use continuous optimization techniques like gradient-descent considering each token distribution as parameters, and keeping the language model’s parameters fixed (§2). Second, we represent each target attribute to control as a differentiable function. We formulate controllable decoding as a multi-objective optimization problem, with maximizing the log-probability of the language model as well as target attributes as objectives. To make this optimization feasible via gradient-descent, we repurpose it to a constraint optimization problem and solve the dual using the modified differential method of multipliers . We call the algorithm MuCoCO, for incorporating multiple constraints through continuous optimization.
We validate MuCoCO on three conditional text generation tasks with different types of sentence level constraints: (1) Adding formality and cross-lingual similarity in a machine translation model; (2) Ensuring transfer and content-preservation in a style-transfer model, and finally (3) Incorporating multiple styles and attributes (e.g., formality, sentiment magnitude, writer’s age group) in a paraphrasing model. With automatic as well as human evaluations we find that our proposed method outperforms strong baselines.
MuCoCO: Constrained Decoding as Multi-Objective Optimization
In this work, however, given and an input sequence , we are interested in finding an output sequence that not only maximizes the output probability but also optimizes multiple objectives defined over and . More formally, we seek to find a that minimizes all of the following objectives
Here each is a function defined over the output sequence , for example, the negative log-probability of an attribute (e.g., formality) classifier we want the output sequence to satisfy. And each is a function defined over both the input and output sequence, for example, semantic similarity between and . We assume all and are differentiable. This is a multi-objective optimization with several possible solutions.
To make optimization feasible, a multi-objective problem generally yields itself to the following formulation:
for some statically or dynamically computed weights and for each and , where . Although this weighted summation formulation is intuitively appealing, it typically requires an expensive grid-search over the various scalings or use of a heuristic . Furthermore, this formulation by definition assumes a trade-off between the different objectives by essentially assigning an importance weight to each of them. This problem is further exacerbated when different objectives have widely varying scalesFor example, classifier log-probabilities are in while sentence similarities usually lie in (0,1). with smaller scale objectives just getting ignored. More concretely, a multi-objective formulation as we define in (1) admits several possible “optimal” solutions also known as the Pareto set . The image of the Pareto set is called the Pareto front. Since we define all objectives using neural networks, the Pareto front in our case is non-convex, where linear combinations of objectives are shown to be unsuccessful in finding good solutions (see figure 1 for an example).
Ideally, our goal is a tunable optimization algorithm that finds solution on the Pareto front, i.e., every solution on the Pareto front should have a hyperparameter value for which the optimization algorithm finds that solution. In order to achieve this, we reframe our optimization problem as a Lagrangian optimization problem instead. We choose one of the losses as the primary objective and consider other losses as constraints. The goal is to minimize the primary loss subject to the secondary losses, each below a threshold value. More formally,
Here and are tunable hyperparameters whose values’ change can result in different solutions on the Pareto front. This formulation leads to an intuitive interpretation of the decoding process that the generated text from the model should satisfy the constraints while being as faithful to the primary objective as much as possible.For example, defining as the probability of a desired attribute in leads to a natural threshold of . For a well-calibrated , an even higher threshold could be used for inducing highly indicative features of in . Consequently, the Lagrangian we end up with looks similar to our original total loss linearly combined as in (2) given by
The fundamental issue in both linear combination of objectives and solving the dual is that fixed scalings and (manually pre-determined or obtained by solving the dual) do not work well with gradient descent to minimize for . Following prior work on differential method of multipliers , we propose to use a single gradient descent to optimize for both Lagrangian multipliers and simultaneously as follows:
We follow the gradient of downwards for the (descent) and upwards for the multipliers (ascent) while making sure that the multipliers remain positive (by setting the multipliers to whenever they become negative). Intuitively, this algorithm works by increasing the value of the multiplier with each gradient step as long as the constraint is violated. But when the constraint is suddenly satisfied and the multiplier is still large, it might take a number of gradient steps before the gradient descent pushes it to , thus causing the solution to be pushed further away from the constraint. As soon as the multipliers become 0 (or negative), the constraint is ignored and the process continues. However when the optimization hits the constraint again, this whole cycle repeats, resulting in “oscillations”. We introduce a dampening parameter to each of the multipliers to reduce these oscillations (again following Platt and Barr ) and update the Lagrangian as follows:
where , and is a hyperparameter. does not affect the final , just how quickly the algorithm converges to it (We use in all experiments). stop-gradient indicates that the argument is detached from the computational graph and does not contribute to the gradient computation. When a constraint is not satisfied (, hence ), the dampening parameter being negative incurs higher penalty on the violation than when not using any dampening, without actually increasing the value of too much. But when the constraint is satisfied, it helps quickly reduce the value of penalty being incurred on the constraint while the multiplier converges to .
2 Optimization: Exponentiated Gradient Descent
3 Preventing adversarial solutions: Annealing the thresholds
Finally, it is well known that most neural network based models are not robust to noise and in fact gradient-based methods have been used to generate adversarial examples for text classifiers . We find in our early experiments that using these models to define constraints can also lead to such cases where the constraints are rapidly satisfied but the generated sentences are disfluent. To prevent this issue, we introduce an annealing schedule during the gradient descent where we start with relaxed thresholds such that they are all satisfied and only the primary loss is active. As the optimization progresses, we gradually decrease the value of the thresholds causing the constraints to get violated resulting in the optimization gradually shifting to updating to satisfy them. The exact schedule we use is described in the next section.
The final decoding algorithm we use in all our experiments is described in the Appendix algorithm 1.
Experimental Setup
We evaluate MuCoCO on the following controlled generation tasks: reinforcing target style in text generated by a style transfer model §3.1 and adding formality to a machine translation model (§3.2). Additionally, we conduct a qualitative analysis of rewriting a product review to adhere to multiple expected attributes like formality, sentiment magnitude, and age group of the author (§4). These tasks include constraints corresponding to both expected attributes in the target sentence (like formality) as well as both source and target sentences (like semantic similarity) with up to 6 constraints per task.
We begin with a style-transfer task, a task aiming to faithfully and fluently rewrite a given sentence such that a desired writing style is reflected in the generation. This task has been widely studied [21, 54, 28, among others] and differs from related tasks like sentiment transfer where flipping the sentiment usually comes at the cost of changing meaning.
Style transfer is usually evaluated across three dimensions: (1) does the output sentence conform to the expected style; (2) does the output sentence preserve the input’s meaning; and (3) is the generated sentence fluent. Most prior work in style transfer focused on devising training objectives serving as proxy for the desired outcomes, for example, back-translation or paraphrasing for content preservation and language modeling for style and fluency. But depending on training algorithm and available data, there is often an observed trade-off between transfer and content-preservation . To that end, we add the desired attributes via explicit constraints when decoding from an existing style transfer model.
More specifically, we consider the task of informal to formal transfer with the state-of-the-art unsupervised model strap from Krishna et al. . This model is trained in an unsupervised fashion by (1) generating a pseudo-parallel corpus by paraphrasing each formal sentence in the training set (which results in a demotion of stylistic attributes), and (2) training an inverse-paraphrase model to translate paraphrases back to the original formal style. At test time, given an informal input sentence , the model first generates its paraphrase , then using an inverse-paraphrase model to generate the output . We train this model by fine-tuning GPT2 (345M) with the GYAFC Corpus (Entertainment/Music domain; around K formal sentences) and evaluate it on the provided test set containing informal sentences. Krishna et al. report best results with greedy decoding. In MuCoCO we modify the decoding algorithm by considering the negative log-probability of given according to the model as the primary objective, and incorporate the following constraints:
Formality: We train a binary classifier by fine-tuning GPT2 on the same GYAFC training corpus, following default hyperparameter choices provided in HuggingFace . This classifier outputs the formality probability of a sentence . We add this output as a constraint to the decoder as . In other words, the constraint is satisfied if the classifier assigns at least 0.5 probability of the output being formal. We initialize the threshold to which is later annealed to .
Baselines and Evaluation Metrics We compare MuCoCO with the following baselines:
No-Constraints: We decode directly from the model greedily without any constraints. This replicates the best result reported by Krishna et al. .
fudge: Introduced by Yang and Klein , this method decodes in an autoregressive manner. It modifies the output vocabulary distribution at every step by interpolating the language model probability with that of a formality classifier. This classifier is trained to predict the probability of entire sentence being formal given only a prefix (we train it similarly to by fine-tuning GPT2). This method only works with categorical features like formality and is not extensible to constraints like semantic similarity. We decode using the hyperparameters recommended in .
Following the baseline model Krishna et al. , we evaluate the generated sentences with the following metrics: (a) fluency or grammatical wellformedness measured by the accuracy of a RoBERTa-based classifier model trained on CoLA , averaged over all outputs, (b) transfer: measured by a RoBERTa-based classifier model trained on the GYAFC training corpus, and finally (c) wsim , a subword embedding based similarity model trained on a large-scale paraphrase corpus which performs well on STS benchmarks as well. We measure this metric both with respect to the input and the provided references.Each input sentence has 4 references, we choose the highest textscwsim value to compute the average. In addition, we also report usim.
Results The style transfer results are summarized in table 1. If we only incorporate a formality constraint, we observe that compared to fudge our method significantly improves transfer accuracy at the expense of content preservation. Adding semantic similarity constraints on the other hand improves both transfer as well as content preservation with the largest gains achieved when all the constraints are considered together. Qualitative analysis shows that MuCoCO’s outputs are typically more fluent and have stronger formality signals, but all of the models are prone to propagating errors from the paraphrasing model (see examples in the Appendix table 3).
2 Style-controlled Machine Translation
We now evaluate MuCoCO in the task of formality transfer in machine translation. Given a trained MT model, decoding is often done using beam search and the highest probability beam candidate is chosen as the final output. Prior work has explored adding rule-based or heuristic constraints such as length penalty or coverage to rerank beam candidates, and adding lexical constraints like penalizing n-gram repetitions . In this experiment, we target sentence-level constraints which are otherwise difficult to incorporate in a left-to-right decoding process. Given a trained MT model and the source text , we use negative log-probability of the translation under the MT model as our primary objective and incorporate the following constraints for decoding in different combinations:
Formality Unlike style transfer, where the goal is to rewrite text in the desired style, here we seek to generate translations in a desired style directly from an MT model which was not explicitly trained to conform to a specific style. We train a classifier similarly to one described in previous section by fine-tuning GPT2, but with a different input-embedding table to match the vocabulary of the decoder of the MT model. Again, we use as the constraint.
Baselines and Evaluation Metrics We compare MuCoCO with the following two baselines:
BeamSearch: We decode directly from the translation model with a beam search of size .
fudge : defined similarly as in the style transfer task but trained to match the decoder vocabulary. As mentioned before, fudge only works with categorical attributes like formality and is not easily extensible to constraints like cross-lingual similarity. We use the recommended hyperparameters by Yang and Klein for decoding.
In Yang and Klein , the authors also compare fudge with other baselines such as PPLM and BeamSearch followed by style transfer. They show that fudge vastly outperforms these baselines. Hence, we only show comparisons with fudge in this work. We evaluate along the following metrics: (a) BLEU : a standard metric for evaluating MT, (b) BERTScore : an embedding-based metric which is more robust to changes in surface forms of the words than BLEU. (b) transfer: the same RoBERTa-based formality classifier as in our style transfer experiments. We also report xsim, the constraint we use for decoding.
We experiment with French to English translation with a subset of the OpenSubtitles test set containing 1360 sentence pairs.We create this subset by filtering the original test set to contain only sentence pairs for which beam search translations are classified as informal. This test set contains informal spoken language for both source and target. For the primary objective, we use the Marian Transformer based French (fr) to English (en) model through Huggingface. We summarize the results of this experiment in table 2 with selected examples in the Appendix table 4.
Results By just using a cross-lingual similarity metric without modifying the model at all, we observe +0.6 improvement in BLEU score as well as BERTScore. Adding a formality constraint leads to considerable gain in formality of the outputs with a drop in BLEU; using both xsim and formal helps recover some of the drop. The drop in BLEU is unsurprising: since BLEU is a surface-level metric it naturally penalizes the translations that are rephrased to conform to formality constraints. Indeed, as shown in table 4, adding a formality constraint leads to changes in sentence structure and vocabulary. On the other hand, we see improvements in BERTScore which is an embedding-based metric, more robust to paraphrasing.
To further validate our results, we conduct a human evaluation of the generated translations. We randomly sample 100 source sentences and their translations generated by beam search and MuCoCO with both formal and xsim constraints. Two annotators (highly proficient in French and English) to rank the translations on faithfulness (is the source meaning reflected in the translation?) and formality. The options are randomized. On the translation pairs where both annotators agree ( out of ), the ones generated by our method were favored by annotators % percent of the time, while beam search translations were favored only % of the time, and % translations were equally favored.
Discussion
Simultaneously controlling several attributes One of the main advantages of our proposed approach is its flexibility to introduce any number of constraints (as long as they are differentiable) to the decoding objective. To illustrate this advantage we consider the following problem: given a sentence annotated with following attributes: age group of the author, formality, and sentiment magnitude, rewrite it such that any chosen combination of the attributes are modified while keeping the others fixed and the content preserved . For our primary objective, we use a inverse-paraphrasing model as defined in §3.1 which we train on a corpus of Yelp ReviewsThis corpus is sentence-tokenized and lowercased with 2.2M sentences not labeled for any attributes. . First, we paraphrase each sentence in the corpus as described in Krishna et al. creating a pseudo-parallel corpus (of reviews and their paraphrases) and train as an inverse-paraphrase model to translate the paraphrases back to the original reviews. We use usim and wmd for semantic similarity constraints and three classifiers for (a) age group of the author (binary; years or years); (b) formality of the review (binary: informal or formal); (c) sentiment magnitude (five-class classifier ratings of 1 to 5). Here we focus on sentiment amplification rather than transfer. That is, changing the 4-star rating of an input to 5 (or 2 to 1). Details of the classifiers and the data used are provided in Appendix C.2.Due to lack of an established benchmark for this task and due to many possible combinations of attributes, we do not report quantitative results. Table 5 shows examples of generated sentences with different combinations of attribute values.
Finding other solutions on the Pareto front As described in §2, the thresholds are tunable hyperparameters that allow us to find different solutions on the Pareto front. In our experiments so far, based on expected outcomes and how the constraints are defined, we showed results with only one threshold for each constraint. For example, ideally for a well-calibrated text classifier based constraint, this technique should be able to find solutions for any probability as threshold, but most neural-network based classifiers are not well-calibrated and predict the highest probability output as the label, hence a natural threshold for binary-classifiers is a label probability . In Appendix table 6, we show how the outputs change if we modify this threshold to different values. We observe that in most cases the optimization converges to generate words more commonly associated with formality. On the other hand, semantic similarity between two sentences is even harder to define, is less robust to noise, and varies with writing styles of the input sentences. As shown, increasing this threshold for semantic similarity can lead to repetitions and disfluency.
Speed and memory requirements The presented decoding algorithm treats each token in the output sequence as a parameter for gradient-descent which involves multiple forward and backward passes through the primary generative model as well as attribute models. Given an expected sequence length , it optimizes parameters which is both memory and time intensive compared to left-to-right decoding. For example, on a single GeForce RTX 2080 Ti (12GB) on which we run all presented experiments, with a batch size of 1, our approach takes approximately 90 minutes on average to decode around 1200 sentences compared to around 20 minutes for fudge with a single constraint. For reference, unconstrained beam-search takes 2-5 minutes. Given enough GPU capacity, however, this approach can easily be extended to larger-batches to improve decoding speed. We do not conduct this experiment due to limited available resources. Using 16-bit floating point operations, this can further be improved. Another way of improving memory efficiency would be to optimize not for tokens directly but instead optimize for token embeddings . This formulation also removes the requirement for all the models to share a vocabulary. We plan to investigate this in future work. Finally, given the capability of this approach to incorporate multiple constraints, it can also be used to generate pseudo-parallel data with different attribute combinations which then could be used to train supervised models for attributes for interest resulting in faster models at inference.
Ethical considerations Controlled language generation is a growing research area, and state-of-the-art techniques are still noisy and not powerful enough to enable fine-grained control over generated content. In the current form, they have the potential to generate harmful and biased language. For example, language generators are prone to generating non-factual content , especially when used maliciously . Moreover, when style transfer techniques are used in conjunction with users’ personal attributes such as gender, they are likely to generalize and amplify harmful social biases . We thus opted not to include gender transfer in our experiments. Our goal in this work is to enable finer-grained control over generated texts that could potentially alleviate these issues. These approaches can be used as a useful tool for mitigating many problematic biases already encoded in large language models , for anonymizing personal attributes , and even as aids for humans to avoid implicit biases in their writing .
Related Work
Recent work on controllable text generation can be divided into two categories. The first focuses on directly training attribute-conditional models either through fine-tuning pretrained models with attribute-specific corpora or via training conditional generative networks . More broadly, this includes methods for text style transfer. Unlike MuCoCO, these methods are not easily extensible and require training or fine-tuning a new model to incorporate new attributes. For example, ctrl train a large scale language model (B parameters) from scratch with control codes capable of generating high-quality text but is very expensive to train. The second line of work, in line with MuCoCO, aims to incorporate control in pre-trained models without retraining them. For example, GeDi trains smaller class-conditional LMs and uses them as discriminators for guided generation. More recently, fudge and dexpert propose changes to left-to-right decoding in language models by modifying the vocabulary distribution at every step using attribute classifiers and ensemble of language models trained on attribute-specific corpora. Although lightweight, these approaches are, by design, prone to a trade-off between preserving content and enforcing the attributes in the generated text. Our work is most closely related to Plug and Play Language Models which use gradients from the attribute models to update the prediction. They work by updating the model activations rather than token probabilities which limits their applicability to only unconditional language models. Furthermore, due to their autoregressive nature, these approaches do not guarantee sequence-level control as they only look at the prefix generated up to a certain step. These are also limited to categorical attributes and can not enforce real-valued controls like semantic similarity.
Gradient-descent based optimization to generate text has been explored in prior work for improving machine translation , paraphrasing and generating adversarial examples . These methods however rely on linear combinations of various objectives which as we discuss in §2 are not optimal for non-convex neural-network based models. This phenomenon has also been studied in multi-task learning where linear combination of multiple task losses is the most common approach and approaches for multi-objective gradient descent have been proposed. These approaches can also be explored for text generation in the future.
Conclusion
We present MuCoCO, a decoding algorithm for controlled generation from (conditional) language models that flexibly combines pretrained LMs with any differentiable constraints. With experiments on style transfer and controlled machine translation, and multiple combination of constraints, we show the effectiveness of this approach. In addition to its potential applications in factual rewriting and debiasing text, this work holds promise in making language generation models personalizable and adaptive to different dialects or even individual speakers, since MuCoCO re-uses pre-trained LMs without adaptation and can incorporate constraints (e.g., dialect or user properties) trained on very little data. Future work will explore more sophisticated optimization techniques to improve the computational efficiency of our approach, and gradient-descent based methods for sampling which will allow to sample from the language models with constraints.
References
Appendix A Overview of the Method
Figure 2 shows an overview of our proposed approach.
Appendix B MuCoCO Decoding Algorithm
Appendix C Details of Attribute Models
We explain the semantic similarity models we use in our experiments in more detail here:
First, we fine-tune GPT2 on the combination of SNLI and MNLI corpora which are both designed for training natural language inference model and intended to capture semantics. Each corpus contains pairs of sentencse with one of the three annotations: inference, contradiction or neutral. For each input sentence , the model is trained as with classification objective with the final logits computed as , where is a trainable parameter. In other words the three vectors as shown are concatenated and multiplied with a weight matrix. We train this for 1 epoch on the combined corpora.
For details of training can be found in where this model is shown to perform competitively on STS benchmarks . We use this model for adding constraints in style-transfer (§3.1) and multi-attribute transfer (§4).
That is, representations of the model and for the source sentence are trained to be close together as are the cross-lingual representations of source and target. We parameterize also with pretrained GPT2 (345M) model. But GPT2 and the Marian Transformer based MT model we use do not have matching vocabularies. Since the vocabulary of the primary objective and constraints should match for the decoding to work, we replace input word embedding layer of GPT2 with that of the decoder of the translation model before we train the distilled model. We use the TED2020 French-English parallel corpus containing around 400K sentence-pairs to train xsim and obtain comparable performance as Reimers and Gurevych on the cross-lingual STS benchmark .
wmd Given two bags of words, and , and an embedding table , we define word mover’s distance between and as
where we define . Given fixed inputs and , wmd can easily be computed using linear program solver We solve it using the python library POT: https://pythonot.github.io/. To backpropagate through this objective. We use the following steps following Kumar et al. :
During the forward pass, we obtain as indicated in algorithm 1 and compute word embeddings for both the input and the prediction . Using the linear program solver, we compute as well the proportions
We use the embedding table from usim model as for this constraint.
C.2 Models used in multi-attribute transfer
In §4, we present a paraphrasing model with 4 different constraints: usim as described previously and three classifier constraints. All the classifiers are trained by finetuning GPT2we use Huggingface with recommended hyperparameters for training all classifiers: https://huggingface.co/transformers/v2.0.0/examples.html on the following corpora:
We use the NUFA corpus consisting Yelp Restaurant Reviews with K sentences per age group (greater than 30 years, and less than 30 years) in the training set. The age was self-declared by the reviewers in their Yelp profiles. Our classifier achieves an accuracy of on a balanced test set of 10K sentences.
Formality
We use GYAFC corpus as described in for this constraint (with an accuracy of around 92%) on the provided test set.
Sentiment
We collect Yelp restaurant reviews using scripts provided by Subramanian et al. https://github.com/facebookresearch/MultipleAttributeTextRewriting/tree/master/data/Yelp with a rating from 1 to 5 star. We subsample from this corpus to train our 5-class classifier on 100K reviews per rating obtaining a classification accuracy of around on a held-out test set also sampled from the same corpus.
Appendix D More Details of Human Evaluation
We conduct A/B testing to rank translations generated by our method and beam search. We show the annotators the source sentence and two randomized translations (one from beam search and one from our method). We ask them to choose one of the four options: : the first translation is both faithful and formal while the second is not, : the second translation is both faithful and formal while the second is not, : both are faithful and formal, and : both are either unfaithful or informal or both. Results are summarized in §3.2.
Appendix E Examples
We show selected examples from our style-transfer models in Table 3. Since the final output is generated from the paraphrase , not the input sentence , some of the content is at times modified in the final output in decoding without constraints. MuCoCO with content based constraints is able to recover content in some examples and also improve formality of the outputs. But it can still be prone to errors since the content-similarity metrics are not perfect. See §3.1 for more details.
E.2 Style-controlled Machine Translation
Table 4 lists few selected examples for inducing cross-lingual similarity and formality constraints in a French to English MT model. We find that inducing formality modifies some of the constructs (like removing contractions: “gonna” to “going to”) in the output sentences which are not measured accurately by a surface-level metric like BLEU. See §3.2 for more details.
E.3 Multiple Solutions on the Pareto Front
Table 6 shows a few examples of changing constraint thresholds for semantic similarity as well as formality constraints. Since the classifiers are not well calibrated, we find that with tighter constraints, the outputs tend to overly represent formality indicating words while losing some of the content which the semantic similarity models are not always robust enough to detect. See §4 for more details.
E.4 Multi-attribute Transfer
Table 5 shows a few examples of transfering multiple combinations of attributes in a given input sentence. We focus on sentiment amplification rather than transfer as it is by definition prone to losing content. See more details in §4.