A Causal Lens for Controllable Text Generation
Zhiting Hu, Li Erran Li
Introduction
Controllable text generation aims at producing fluent language with control over various attributes, ranging from sentiment, topic, politeness, to gender, persona, and so forth . The problem lies at the heart of many NLP applications such as emotional chatbot, news article writing, language detoxification, etc. Of particular interest in this increasingly significant area are two settings for control, namely (1) attribute-conditional generation which generates sentences that entail a given attribute, and (2) text attribute transfer which rewrites a given sentence to possess a desired attribute while preserving all other original characteristics (Figure 1). The goal is to learn the control in each setting with (attribute, text) training pairsThus, for text attribute transfer (a.k.a., text style transfer), there is no direct supervision data, i.e., (original text, attribute, target text) triples..
The two settings have usually been considered as separate tasks and each led to various solutions, respectively. Let denote a sentence and an attribute. Previous attribute-conditional generation work typically concerns the conditional distribution . Despite the success of simulating observed real text, the conditional distribution is known to be susceptible to capture spurious correlations or biases in the data . For example, when generating biographical text given a gender attribute, the conditional model tends to generate text related to specific occupations such as nurse and yoga teacher for female, and architect and attorney for male (Figure 1). The learned biases could impair the model generalization to new domains, and make negative social impact in downstream applications. A few very recent attempts have been made to mitigate the biases in the model with various machine learning techniques. Yet those methods are often specific to a particular attribute (e.g., gender) , or rely on access to additional resources, such as fully observed confounding labels or a priori debiased classifiers , which can be costly to obtain in real applications. Furthermore, it is unclear how the diverse methods designed for attribute-conditional generation could also be applied to debias text attribute transfer that has been formulated with distinct training objectives.
This paper studies controllable text generation from a principled causal perspective, that offers a unifying formulation of the two central tasks, and enables mitigation of spurious correlations with well-established causality techniques. A growing number of recent work has used causality with machine learning for disentangled representation , model explanation , and robust prediction . Yet most approaches have focused on the vision domain, taking advantage of image spatial structures, and thus are not directly applicable to text with abstract attributes (such as sentiment). Though previous research on text modeling has also studied related concepts such as counterfactuals, it either handles only correlation instead of causation , or focuses on different applications such as data augmentation and classification . We discuss more related work in §4.
We develop the first unified causal framework for text generation under control. In particular, we devise a structural causal model (SCM) that describes the causal relationships between different variables, where the text is outcome and the attribute to control (e.g., sentiment) is treatment. The SCM further accounts for spurious correlations with confounders (e.g., category) with latent variables. The resulting SCM enables us to formulate the two control tasks as performing causal inference at different rungs of the causal ladder (Figure 1), respectively. Specifically, (1) for attribute-conditional generation, we go beyond the association-based conditional and propose to instead use , corresponding to the intervention rung. The -operation effectively eliminates the effect of confounders on the control, leading to unbiased text outputs; (2) for text attribute transfer, the task naturally maps to counterfactual prediction on the SCM, which answers the question “what the text would have been if the attribute had been different” through the standard causal inference procedure . The unifying perspective also allows us to draw from existing successful techniques and train the SCM for accurate control and confounder balancing .
Previous causal work typically assumes access to confounding labels or relevant proxy information for the entire observed data . In many real applications, however, it is prohibitively expensive or impossible to measure all the confounding factors for unbiased training. For example, it is often not affordable to annotate massively the confounding labels for the entire (attribute, text) corpus. We thus consider a more practical yet challenging scenario where we observe confounding information for only a small subset (e.g., ) of samples . We experiment on difficult datasets where the target attributes and confounding factors have strong correlations. Results show the causal approach substantially improves over conventional conditional models with enhanced control accuracy and reduced bias, on both attribute-conditional generation and attribute transfer.
Background
We first briefly review the causal concepts most relevant to the paper. A structural causal model (SCM) is defined by a directed graph consisting of nodes (variables) and edges (direct causal dependence between variables), e.g., Figure 2. Different inference questions on an SCM correspond to different levels of the causal ladder (Figure 1) and require different reasoning tools: (1) “Association” deals with correlations in observed data with joint/marginal/conditional distributions. (2) “Intervention” concerns what would happen were some actions been performed. A typical question is to estimate the distribution of an outcome variable given an intervention on a treatment variable : , where the -operation represents an action on by setting it to a given value. With randomized experimental data (i.e., collected by randomly assigning treatment), equals to the standard conditional . Yet in practice, we usually only have access to passively observed data, such as the (attribute, text) pairs from existing corpus, and have to adjust for confounders (i.e., variables that correlate with both treatment and outcome) in order to estimate from observational distributions. For example, we apply backdoor adjustment in §3.2 for attribute-conditional generation. Finally, (3) “Counterfactuals” involves queries about what would have happened, given the knowledge of what in fact happened. We next show how the controllable text generation tasks are bridged together as different levels of causal inference, operationalized by the proposed SCM.
The Causal Framework for Controllable Text Generation
We now describe the unified causal perspective. We first develop the structural causal model that characterizes the causal structure in the controlled generation process (§3.1). We then show that intervention on the SCM leads to attribute-conditional generation (§3.2), while counterfactual prediction makes attribute transfer (§3.3). At last, §3.4 describes the training of the SCM with new objectives to encourage confounder balancing and de-correlation.
Compared to previous causal modeling in other domains (e.g., images), modeling text as the outcome is challenging due to the complex unstructured information encoded in the text. We show here that the unifying perspective enables us to bring to bear rich tools and inspirations from causal inference, disentangled representation, and controllable generation, for effective text causal modeling.
Figure 2 shows the SCM graphs as detailed below. In the appendix we illustrate the model architecture used in our experimental studies.
Figure 2(a) shows the SCM that describes the controlled generation process of text. Here the attribute of interest serves as the treatment, and the text is the outcome. For simplicity, we assume is binary (e.g., positive or negative sentiment), though the framework can straightforwardly be applied to more general cases where has multiple classes and dimensions. Note that as the condition for generating can be instantiated in different forms depending on the concrete application. For example, it can be a scalar as an input to the generator, or a word sequence such as ‘‘[sentiment] positive’’, ‘‘[sentiment] negative’’ that acts as a prompt for the generator to produce text continuation .
In general, the confounder that induces spurious correlations between and is infeasible to be fully specified or observed. For example, to control the sentiment of a restaurant review, the confounder could involve popularity of the restaurant, personal preferences of the customer, and other factors, whose values cannot be directly measured. Thus, following the recent causal approaches in other domains , we model the unobserved confounder as a high dimensional latent variable , and infer from “indirect” confounding variables that are measurable in practice (such as food type). The “indirect” variables are also called proxy variables in causality , which we denote as . More background of confounder and proxy is provided in the appendix.
The goal of controllable text generation is thus to generate coherent text (given original text in the case of text attribute transfer) with accurate target attribute while unbiased in terms of the confounders.
Previous causal studies have usually assumed the confounder proxy is available for all data. Similarly, recent work of debiasing attribute-conditional generation (based on machine learning) has relied on access to those extensive proxy labels . However, the assumption is often impractical due to the time and financial cost for obtaining the massive additional information beyond the common (attribute, text) data. We thus consider a more practical setting where we only have access to the proxy information for a small subset (e.g., ) of examples. In Figure 2(a), we use dashed circle of to denote the new challenging setting.
The resulting SCM thus defines a joint distribution
where the component applies when the proxy is observed for the example; is a standard Gaussian prior following the common practice; and all components with free parameters are modeled as deep neural networks. We use the amortized variational inference as in variational auto-encoders (VAEs) to infer the latent confounder from observations. Specifically, we introduce a variational distribution with parameters . To infer for those examples whose proxy is not available, we could apply an auxiliary predictor that estimates from the observed . The auxiliary predictor can be trained on the subset of examples with available . In this work, we instead set the default to a dummy value when inferring for simplicity.
Note that previous work has also used VAEs for controllable text generation . However, they reside purely at the association level. In particular, despite the latent variables, they do not explicitly model the confounder and/or its proxy information, rendering the causal effects between components not identifiable . As a result, those models are vulnerable to biases, as shown in the experiments.
2 Inference (I): Intervention for Attribute-Conditional Generation
We now discuss how to perform causal inference given the SCM for attribute-conditional generation. As mentioned in §1, in contrast to the conventional association-level methods based on the conditional , here we formulate the task with the interventional conditional . The -operation sets to a given value independently of (§2), which eliminates the dependence between and , leading to the new intervened causal graph in Figure 2(b), where the arrow from to is removed. Thus captures the causal effect of attribute on text outcome without confounding bias. We can use the backdoor adjustment to estimate from the observed data:
That is, we adjust for confounder by making fair considerations of every possible values and averaging the results by the distribution discussed below. The difference from the previous methods becomes even clearer if we similarly decompose the conditional , which as we can see depends on and inherits the correlations between and in the data.
We generate text samples from approximately by first drawing and then decoding with . To sample from the marginal which does not have a clean analytic form, we use a similar approach as in by fitting a simple generative adversarial network (GAN) , , on the learned latent space (i.e., on all training data). We found it is sufficient to use a single-layer GAN which is fast to train.
3 Inference (II): Counterfactual for Text Attribute Transfer
Given an observed text , text attribute transfer seeks to produce new text that possesses the given new attribute and preserves as many characteristics of the original as possible. The task can naturally be mapped to counterfactual prediction on the SCM, i.e., imagining the alternative outcome for should its attribute have been . The resulting inference procedure looks similar to the previous VAE-based attribute transfer method . However, besides the key modeling difference of confounder/proxy as above, our counterfactual based interpretation offers a principled causal account for the attribute transfer task. Moreover, the causal perspective inspires new training techniques that substantially improve the performance and reduce generation bias, as presented in §3.4.
Figure 2(c) illustrates the inference process. Specifically, from the causal perspective, counterfactual prediction is mathematically formulated as a three-step procedure : (1) Abduction that infers the “context” compatible with the observation . In our problem, it is sufficient to infer as the context, as we would only intervene its descendant node in the SCM. Thus the step is done by computing ; (2) Action that performs intervention on variable by setting ; and (3) Prediction that computes the counterfactual outcome based on the SCM, i.e., , where we set to the mean (vector) of the above abduction distribution for simplicity.
4 Learning
With the causal model and the inferences on it, we now discuss model training, which integrates variational learning and counterfactual reasoning for confounder balancing and disentanglement.
The base objective for learning the causal model is built on the common VAE approach . Briefly, since the model’s marginal log-likelihood (that marginalizes out the latent ) is intractable, VAEs derive a lower bound with the variational distribution . Formally, given a training example with the optional proxy :
where the first term is the reconstruction that aims to recover the observations given the inferred from ; the second term is a Kullback–Leibler regularizer that enforces the variational distribution to stay close to the prior . We refer readers to for more details of VAEs. In the objective, , and are balancing hyperparameters. We set to when proxy is not available, and otherwise select from based on validation, same as . We use the cyclic schedule from to anneal from to to avoid excessive regularization of the KL term.
Counterfactual objectives
Training with the above base objective alone can lead to model collapse where the attribute variable is ignored in the generation process, i.e., text sampled from is not effectively controlled by . This is because the training text has already contained the attribute information, allowing both the inference of and the subsequent reconstruction of not to depend on the attribute value . This issue highlights a key difference of our model compared to previous latent-confounder causal models in other domains (e.g., medication effect prediction), where the outcome is typically a simple binary variable (e.g., cured or not) that does not “leak” the treatment information (e.g., medication) . Causal controllable text generation thus requires new solutions to encourage effective control.
Besides, a key ingredient for accurate causal inference is to achieve balance of confounders between treatment groups . That is, we want to match the confounder representation of the examples whose and those of the examples whose , in order to enhance the generalization performance for inferring counterfactual outcomes . The concept is closely related to disentangled representation in machine learning which seeks to keep most dimensions of a representation invariant to the change of a particular dimension .
The above two desiderata can be resolved with a suite of counterfactual objectives that are based on the counterfactual outcomes inferred in §3.3. We now describe those intuitive objectives, which are related to the attribute , confounder , and proxy , respectively. We also discuss how we are able to draw inspirations from previous literature of disentangled representation and text attribute transfer, thanks to their connections with causal inference as above.
The first objective concerns the attribute to correctly learn its influence on the outcome. Intuitively, given the counterfactual outcome given the counterfactual attribute , we want to make sure truly entails . This can be achieved by using a pretrained attribute classifier that estimates the likelihood of text possessing attribute . More specifically, we train the model such that its predicted possesses with a high likelihood measured by the classifier:
As in , we use Gumbel-softmax approximation to the discrete text to enable gradient backpropagation for optimizing . Similar objective has been used in previous conditional generation of text and image . A crucial caveat is that, here the classifier itself pretrained with the (attribute, text) data can also be biased due to confounding factors. Thus relying only on this objective as in the previous work is not sufficient for accurate unbiased attribute control, as shown in our experiments. To this end, we further devise the following counterfactual objectives.
The second objective focuses on balancing the confounder . Intuitively, by the definition of counterfactuals, must have the same confounder representation as the original . We thus minimize the distance between the respective and :
where, with slight abuse of notation, is the mean (vector) of , is the mean (vector) of on the counterfactual , and is a distance metric. Though the vectors and have continuous values, we draw inspiration from the recent disentangled representation work and use a binary cross-entropy loss to match to , i.e.,
where normalizes to $\sigma(\cdot)\bm{z}^{\prime}\textit{mean}(\cdot)\bm{z}L_{2}\|\bm{z}^{\prime}-\bm{z}\|^{2}$ as used in earlier work .
The third objective carries similar intuition as above, though uses the proxy when it is available. Specifically, we want to be able to reconstruct (as is ):
where is the mean (vector) of , same as in Eq.(5).
In sum, the overall objective for training the causal model is:
with balancing hyperparameters , and . In practice, we found the model is not sensitive to the choices of those hyperparameters. We set each of them to either or based on validation.
Related Work
There is an emerging interest in integrating causality with machine learning in various problems. Several latest works have studied causal inference combined with deep generative models for images, to learn causal structures between attributes , synthesize novel images , and augment unbiased classifier training . The spatial structure of images can make it easier to learn causal mechanisms, e.g., the work specified independent modules for image background and texture. In contrast, text with abstract concepts (e.g., sentiment, topics) exhibits less independent structure. Previous causal modeling for text usually focuses on language understanding . Recent work has also studied text as outcome in causal inference for data augmentation or generating text in specific domains (e.g., court view ). We make the first study of causal modeling for the general problem of text generation under control and demonstrate the effectiveness for bias mitigation.
Controllable text generation
Various approaches have been developed for attribute-conditional generation, by learning conditional language models (LMs) , guided inference , or prompts . Recent work has focused on reducing gender bias in machine translation and generation . Other work studied more general unbiased generation with ML assuming access to unbiased classifiers . We use causal techniques to address a different and challenging setting where only limited confounding labels are observed. Unsupervised text attribute transfer has gained increasing attention , with the primary focus on learning to disentangle target attribute with other factors. We study the new challenge of attribute transfer in the presence of strong bias in the data, and show greatly improved performance.
Experiments
We study the challenging generation tasks with strong spurious correlations in training data. The causal framework substantially reduces bias and improves control accuracy.
We describe detailed model configurations in appendix. Briefly, the main model components, including the decoder , inference network , and classifier (Eq.4) are all based on the GPT-2 (117M) architecture with pretrained weights, respectively. In and , we use the GPT-2 final-step output feature as the representation of input sentence . We implement other components ( and ) as simple MLPs. The model is trained with AdamW optimizer using an initial learning rate of 1e-6. All experiments were conducted on 8 Tesla V100 GPUs.
We first evaluate the interventional inference for attribute-conditional generation (§3.2). We use two datasets where the target attribute has a correlation strength of over 90% with the confounding factor, following the challenging settings of the latest work on visual bias . That is, the target attribute and the confounding factor of over 90% examples are both positive or negative, while those of the rest 10% examples are opposite. Differing from previous studies, we further assume the model can observe the confounding labels of a small subset of data, a more practical setting as in §3.1.
Our first dataset is derived from the Yelp challengehttps://www.yelp.com/dataset/challenge that contains customer reviews of different categories. Sentiment (1:positive vs. 0:negative) is the attribute we aim to control, and the category of review object (1:restaurant vs. 0:others) is the confounding factor. Specifically, we extract a subset of data where 90% restaurant reviews are of positive sentiment, while 90% reviews to other entities (e.g., shopping) are of negative sentiment (thus a 90% correlation strength). We keep the category labels for less than 2% of training data. The resulting data has 510K/6K training/validation examples, wherein 10K training examples have observable confounding category labelsDue to the strong correlation, only 10%K1K examples have opposite sentiment and category labels, posing a significant challenge for the model to de-correlate the two factors.. For evaluation, we further create a balanced test set of 13K examples with correlation strength 50% (i.e., no correlation). Following the previous controllable generation , we focus on generating short text, by truncating the output text in the data to 20 tokens at maximum.
The second dataset is from the Bios corpus that contains online biographies with gender and occupation labels. We use gender (female/male in the corpus) as the attribute to control. Thus the goal is to generate biographical text of a given gender. For occupation which is the confounding factor, we subsample and merge the occupations into two groups, i.e., {nurse, dietitian, paralegal, …} and {rapper, DJ, surgeon, …} (see appendix for more details). The correlation strength of the resulting dataset is 95%. For example, 95% female biographies are about the occupations in group one. We randomly split the dataset into 43K training and 2K validation examples, and keep the binary occupation labels for only 3K randomly selected training examples (among which only 5%K150 examples have opposite gender and occupation labels). As above, we further create a balanced test set of 2K examples for evaluation, and truncate the output text to no more than 20 tokens.
Baselines and setup
We compare with the conditional language models that people would commonly train for the task. The first model, Conditional LM, conditions only on the target attribute and generates text accordingly. The second model, Conditional LM (full), makes full use of the attribute and confounding labels in hope of better de-correlating the two. Since the confounding labels are available only on a small subset of examples, we first train a classifier on the subset with data-reweighting (see appendix for details), and use it to predict confounding labels for the remaining examples. The language model is then trained on the resulting complete data, conditioning on both the attribute and the (real or estimated) confounding label. We also compare with latest attribute-conditional generation approaches, such as GeDi where a language model conditioning on the confounding information is used to reshape the generation distribution of the above Conditional LM. We include comparison with more baseline methods in the appendix.
For our approach, the available confounding labels serve as the proxy . The attribute classifier used to train our model (Eq.4) is pretrained on the biased training data. On Yelp, the resulting (sentiment) classifier has a mediocre accuracy of 83% on the balanced test set; On Bios, the (gender) classifier has an accuracy of 91%.
Evaluation
We conduct both automatic and human evaluation. For the former, we follow the common practice and evaluate the generations in terms of various aspects as following: (1) Control accuracy for which we use an “evaluation attribute classifier” that takes as inputs the generated sentences and measures how accurate they entail the input attributes. The evaluation attribute classifier is trained on a large unbiased set of examples from the original corpus and is of high test accuracy (87% for Yelp and 95% on Bios) for evaluation purpose (note the difference from the above classifier trained with only biased training data); (2) Bias which is measured by another classifier for the confounding factor. Intuitively, the better the predicted confounding labels match the input attributes, the more correlated the two factors in the generation. A 50% match indicates no correlation. The classifiers are trained similarly as the evaluation attribute classifiers, and achieve accuracy 85% on Yelp and 90% for Bios; (3) Fluency which is measured by applying GPT-2 language models (LMs) on the generated text and computing the perplexity; The LMs obtain perplexity of 32.4 and 18.0 on the real text of Yelp and Bios, respectively. (4) Diversity with the common Distinct- metric that measures the ratio of unique -grams against total number of -gram in the generation set. We evaluate 10K generated samples by each model.
For human evaluation, we ask human raters to annotate for each generated text the attribute label and confounding factor label, based on which we compute the control accuracy and bias as above. We also annotate language fluency using a 5-point Likert scale. On each dataset, we compare Conditional LM (full) and our approach, with 100 sentences from each model annotated by 3 raters. The Pearson correlation coefficient of human scores is 0.67, showing strong inter-rater agreement.
Results
Table 1 shows the automatic evaluation results on both Yelp and Bios. Our causal approach significantly improves over the association-based conditional models. For example, on Yelp, our model achieves 16% absolute improvement in terms of control accuracy, and at the same time reduces the bias (spurious correlation with the confounder) by 19%. In contrast, the conditional LMs mostly inherit the bias from the training data. As an ablation study, we also evaluate a simplified variant of our full approach by omitting the counterfactual objectives w.r.t and (Eqs.5 and 8) (which reduces to a training strategy similar to the previous methods [e.g., 22]). The variant improves the control accuracy over the conditional LMs, but fails to effectively reduce the generation bias. The results show the crucial role of confounder balancing in bias reduction. On the Bios dataset, our approach also obtains consistent improvement on both accuracy and bias.
Table 2 shows the human evaluation results on both datasets, which largely confirm the above observations with automatic evaluation.
2 Text Attribute Transfer
We next study text attribute transfer (§3.3) as the second core task of controllable generation. The proposed causal approach also achieves substantial improvement in terms of accurate control and bias reduction. Besides, for a broader comparison, we also apply our approach to another unbiased dataset widely studied in previous text attribute transfer research, showing superior performance.
We use the above biased Yelp dataset (§5.1) to study the attribute transfer, where we aim to modify a sentence to possess the opposite sentiment (e.g., from negative to positive), and at the same time preserve all other characteristics. In particular, we want the new sentence to keep the category unchanged, which is difficult for previous association-based controllable models given the strong correlation in the data between sentiment and category. Besides, since most previous attribute transfer studies have focused only on unbiased setting, we additionally evaluate our approach on the popular unbiased Yelp data (reviews with sentiment for restaurants only) for comparison.
Evaluation
We follow the standard practice for evaluation. For the biased setting, we measure control accuracy, bias, and fluency as in §5.1. We also assess the common aspect preservation, which evaluates the BLEU score between the generated and original sentences (i.e., self-BLEU). A higher score indicates better preservation of sentence properties. For the unbiased setting, we omit the bias evaluation, and additionally compute another preservation metric, ref-BLEU, which is the BLEU score between the generation and human-written golden text on a subset of test examples . We also conduct human evaluation which shows the same conclusions as the automatic evaluation in terms of model performance. We put the results in appendix due to space limitation.
Results
Table 3 shows results on the biased Yelp data, a substantially more challenging setting than the popular unbiased one (Table 4). We compare two of the previous best-performing methods with public code. Our approach again manages to reduce the bias while achieving decent transfer accuracy. The previous methods struggle to edit the text on many instances (e.g., generating the same sentences as inputs), leading to low control accuracy. Ablation comparison with our simplified variant (our w/o cf-z/c) further validates the effect of counterfactual objectives for confounder balancing (§3.4), as shown by the improved accuracy and mitigated bias of the full approach.
Finally, Table 4 shows the results on the common unbiased Yelp sentiment data. The results show our approach generates fluent output with improved accuracy and preservation.
Conclusions and Future Work
We have presented a principled causal perspective for the two core tasks of controllable text generation. Based on the proposed structural causal model, attribute-conditional generation is modeled as interventional inference, and text attribute transfer performs counterfactual prediction. We connect rich techniques in causality, disentangled representation, and text generative modeling, and develop learning objectives for accurate control and confounder balancing. Focusing on the challenging setting with partially available confounding information, the experiments show our approach achieves accurate control and mitigates the strong correlations in the data.
The proposed causal framework opens up a range of new opportunities for further improving and enriching controllable text generation. For example, though this work has focused on single control attribute and confounding factor, it would be interesting to generalize the approach for structured control of a richer set of text attributes, by modeling the underlying causal graph between attributes (as explored similarly in image generation ). Besides, we are interested in importing more causality tools through the causal perspective to enable new applications. For instance, the inverse propensity reweighting technique in causality can potentially be used to debias pretrained language models , with the following known equation between the unbiased interventional conditional and the biased standard conditional :
where is known as the propensity score , i.e., the propensity (probability) of the being assigned to the particular treatment . Plugging in the together with the parameterized estimates of and as learned in §3, we would effectively convert the pretrained LM into the unbiased . Further, rich studies in the causality literature have proposed stabilized and enhanced variants of the above inverse propensity reweighting [e.g., see 80], all of which present interesting topics to explore in the controllable generation setting in the future.
We would like to note that automatic text generation could be used maliciously to generate fake, toxic, or offensive content . We hope the unbiased modeling study could offer techniques to alleviate potential issues.