Causal Effects of Linguistic Properties

Reid Pryzant, Dallas Card, Dan Jurafsky, Victor Veitch, Dhanya Sridhar

Introduction

Social scientists have long been interested in the causal effects of language, studying questions like:

How should political candidates describe their personal history to appeal to voters (Fong and Grimmer, 2016)?

How can business owners write product descriptions to increase sales on e-commerce platforms (Pryzant et al., 2017, 2018a)?

How can consumers word their complaints to receive faster responses (Egami et al., 2018)?

What conversational strategies can mental health counselors use to have more successful counseling sessions (Zhang et al., 2020)?

To study the causal effects of linguistic properties, we must reason about interventions: what would the response time for a complaint be if we could make that complaint polite while keeping all other properties (topic, sentiment, etc.) fixed? Although it is sometimes feasible to run such experiments where text is manipulated and outcomes are recorded Grimmer and Fong (2020), analysts typically have observational data consisting of texts and outcomes obtained without intervention. This paper formalizes the estimation of causal effects of linguistic properties in observational settings.

Estimating causal effects from observational data requires addressing two challenges. First, we need to formalize the causal effect of interest by specifying the hypothetical intervention to which it corresponds. The first contribution of this paper is articulating the causal effects of linguistic properties; we imagine intervening on the writer of a text document and telling them to use different linguistic properties.

The second challenge of causal inference is identification: we need to express causal quantities in terms of variables we can observe. Often, instead of the true linguistic property of interest we have access to a noisy measurement called the proxy label. Analysts typically infer these values from text with classifiers, lexicons, or topic models Grimmer and Stewart (2013); Lucas et al. (2015); Prabhakaran et al. (2016); Voigt et al. (2017); Luo et al. (2019); Lucy et al. (2020). The second contribution of this paper is establishing the assumptions we need to recover the true effects of a latent linguistic property from these noisy proxy labels. In particular, we propose an adjustment for the confounding information in a text document and prove that this bounds the bias of the resulting estimates.

The third contribution of this paper is practical: an algorithm for estimating the causal effects of linguistic properties. The algorithm uses distantly supervised label propagation to improve the proxy label Zhur and Ghahramani (2002); Mintz et al. (2009); Hamilton et al. (2016), then BERT to adjust for the bias due to text Devlin et al. (2018); Veitch et al. (2020). We demonstrate the method’s accuracy with partially-simulated Amazon reviews and sales data, perform a sensitivity analysis in situations where assumptions are violated, and show an application to consumer finance complaints. Data and a package for performing text-based causal inferences is available at https://github.com/rpryzant/causal-text.

Causal Inference Background

Causal inference from observational data is well-studied (Pearl, 2009; Rosenbaum and Rubin, 1983, 1984; Shalizi, 2013). In this setting, analysts are interested in the effect of a treatment TT (e.g., a drug) on an outcome YY (e.g., disease progression). For ease, we consider binary treatments. The average treatment effect (ATE) on the outcome YY is,

For example, if the confounding variable CC is discrete, we group the data into values of CC, calculate the average difference in outcomes between the treated and untreated samples of each group, and take the average over groups.

Causal Effects of Linguistic Properties

We are interested in the causal effects of linguistic properties. To formalize this as a treatment, we imagine intervening on the writer of a text, e.g., telling people to write with a property (or not). We show that to estimate the effect of using a linguistic property, we must consider how a reader of the text perceives the property. These dual perspectives of the reader and writer are well studied in linguistics and NLP;Literary theory argues that language is subject to two perspectives: the “artistic” pole – the text as intended by the author – and the “aesthetic” pole – the text as interpreted by the reader Iser (1974, 1979). The noisy channel model Yuret and Yatbaz (2010); Gibson et al. (2013) connects these poles by supposing that the reader perceives a noisy version of the author’s intent. This duality has also been modeled in linguistic pragmatics as the difference between speaker meaning and literal or utterance meaning Potts (2009); Levinson (1995, 2000). Gricean pragmatic models like RSA Goodman and Frank (2016) similarly formalize this as the reader using the literal meaning to help make inferences about the speaker’s intent. we adapt the idea for causal inference.

Figure 1 illustrates a causal model of the setting. Let WW be a text document and let TT (binary) be whether or not a writer uses a particular linguistic property of interest.We leave higher-dimensional extensions to future work. For example, in consumer complaints, the variable TT can indicate whether the writer intends to be polite or not. The outcome is a variable YY, e.g., how long it took for this complaint to be serviced. Let ZZ be other linguistic properties that the writer communicated (consciously or unconsciously) via the text WW, e.g. topic, brevity or sentiment. The linguistic properties TT and ZZ are typically correlated, and both variables affect the outcome YY.

We are interested in the average treatment effect,

where we imagine intervening on writers and telling them to use the linguistic property of interest (setting T=1T=1, “write politely”) or not (T=0T=0). This causal effect is appealing because the hypothetical intervention is well-defined – it corresponds to an intervention we could perform in theory. However, without further assumptions, ψwri.\psi^{\textrm{wri.}} is not identified from the observational data. The reason is that we would need to adjust for the unobserved linguistic properties ZZ, which create open backdoor paths because they are correlated with both the treatment TT and outcome YY (Figure 1).

(overlap) For some constant ϵ>0\epsilon>0,

Then the ATE ψrea.\psi^{\textrm{rea.}} is identified as,

Moreover, the ATE ψrea.\psi^{\textrm{rea.}} is equal to ψwri.\psi^{\textrm{wri.}}.

Substituting Proxy Labels

The following result shows that the estimand ψproxy\psi^{\textrm{proxy}} only attenuates the ATE that we want, ψrea.\psi^{\textrm{rea.}}. That is, the bias due to proxy treatments is benign; it can only decrease the magnitude of the effect but it does not change the sign.

TextCause , A Causal Estimation Procedure

The first stage of TextCause is motivated by Theorem 2, which said that a more accurate proxy can yield lower estimation bias. Accordingly, this stage uses distant supervision to improve the fidelity of lexicon-based proxy labels T^\hat{T}. In particular, we exploit an inductive bias of frequently used lexicon-based proxy treatments: the words in a lexicon correctly capture the linguistic property of interest (i.e., high precision, Tausczik and Pennebaker, 2010), but can omit words and discourse-level elements that also map to the desired property (i.e., low recall, Kim and Hovy, 2006; Rao and Ravichandran, 2009).

Train a classifier to predict Pθ(T^ ∣ W)P_{\theta}(\hat{T}\,|\,W), e.g., logistic regression trained with bag-of-words features and T^\hat{T} labels.

Use T^∗\hat{T}^{*} as the new proxy treatment variable.

2 Adjusting for Text

Letting θ\theta be all parameters of the model, our training objective is to minimize,

where L(⋅)L(\cdot) is the cross-entropy loss and R(⋅)R(\cdot) is the original BERT masked language modeling objective, which we include following Veitch et al. (2020). The hyperparameter α\alpha is a penalty for the masked language modeling objective. The parameters Mt\mathbf{M}_{t} are updated on examples where T^i∗=t\hat{T}^{*}_{i}=t.

Once Q^(⋅)\hat{Q}(\cdot) is fitted, an estimator ψ^proxy\hat{\psi}^{\textrm{proxy}} for the effect ψproxy\psi^{\textrm{proxy}} (Eq. 7) is,

Experiments

We evaluate the proposed algorithm’s ability to recover causal effects of linguistic properties. Since ground-truth causal effects are unavailable without randomized controlled trials, we produce a semi-synthetic dataset based on Amazon reviews where only the outcomes are simulated. We also conduct an applied study using real-world complaints and bureaucratic response times. Our key findings are

More accurate proxies combined with text adjustment leads to more accurate ATE estimates.

Naive proxy-based procedures significantly underestimate true causal effects.

ATE estimates can lose fidelity when the proxy is less than 80% accurate.

Dataset. Here we use real world and publicly available Amazon review data to answer the question, “how much does a positive product review affect sales?” We create a scenario where positive reviews increase sales, but this effect is confounded by the type of product. Specifically:

The text WW is a publicly available corpus of Amazon reviews for digital music products Ni et al. (2019). For simplicity, we only include reviews for mp3, CD, or Vinyl. We also exclude reviews for products worth more than $100 or fewer than 5 words.

The observed covariate CC is a binary indicator for whether the associated review is a CD or not, and we use this to simulate a confounded outcome.

The proxy treatment T^\hat{T} is computed via two strategies: (1) a randomly noised version of TT fixed to 93% accuracy (to resemble a reasonable classifier’s output, later called “proxy-noised”), and (2) a binary indicator for whether any words in WW overlap with a positive sentiment lexicon Liu et al. (2010).

The final data set consists of 17,000 examples.

Protocol. All nonlinear models were implemented using PyTorch Paszke et al. (2019). We use the transformershttps://huggingface.co/transformers implementation of DistillBERT and the distilbert-base-uncased model, which has 66M parameters. To this we added 3,080 parameters for text adjustment (the Mtb\mathbf{M^{b}_{t}} and Mtc\mathbf{M^{c}_{t}} vectors). Models were trained in a cross-validated fashion, with the data being split into 12,000, 2,000, and 4,000-example train, validation, and test sets.See Egami et al. (2018) for an investigation into train/test splits for text-based causal inference. BERT was optimized for 3 epochs on each fold using Adam Kingma and Ba (2014), a learning rate of 2e−52e^{-5}, and a batch size of 32. The weighting on the potential outcome and masked language modeling heads was 0.1 and 1.0, respectively. Linear models were implemented with sklearn. For T-boosting, we used a vocab size of 2,000 and L2 regularization with a strength of c=1e−4c=1e^{-4}. Each experiment was replicated using 100 different random seeds for robustness. Each trial took an average of 32 minutes with three 1.2 GHz CPU cores and one TITAN X GPU.

Note that for clarity, we henceforth refer to the treatment-boosting and text-adjusting stages of TextCause as T-boost and W-Adjust .

1.2 Results

Our primary results are summarized in Table 1. Individually, T-boost and W-Adjust perform well, generating estimates which are closer to the oracle than the naive “unadjusted” and “proxy-lex’ baselines. However, these components fail to outperform the highly accurate “proxy-noised” baseline unless they are combined (i.e., the TextCause algorithm). Only the full \scTextCause{\sc TextCause} algorithm consistently outperformed (i.e. produced higher quality ATE estimates) than the baselines. This result is robust to varying levels of noise and treatment/confound strength. Indeed TextCause ’s estimates were on average within 2% of the semi-oracle. Furthermore, these results support Theorem 2: methods which adjusted for the text always attenuated the true ATE.

Our results suggest that adjusting for the confounding parts of text can be crucial: estimators that adjust for the covariates CC but not the text perform poorly, sometimes even worse than the unadjusted estimator ψ^naive\hat{\psi}^{\textrm{naive}}.

Does it always help to adjust for the text? We consider the case where confounding information in the text causes a naive estimator which does not adjust for this information (ψnaive\psi^{\textrm{naive}}) to have the opposite sign of the true effect ψ\psi. Does our proposed text adjustment help in this situation? Theorem 2 says it should, because ψproxy\psi^{\textrm{proxy}} estimates are bounded in [0, ψ\psi]. This ensures that the most important of bits, the bit of directional information, is preserved.

Table 2 shows results from such a scenario. We see that the true ATE of TT, ψ\psi, has a strong negative effect, while the naive estimator ψnaive+C\psi^{\textrm{naive+C}} produces a positive effect. Adding an adjustment for the confounding parts of the text with TextCause successfully brings the proxy-based estimate to 0, which is indicative of the bounded behavior that Theorem 2 suggests.

Sensitivity analysis. In Figure 3 we synthetically vary the accuracy of a proxy T^\hat{T} by dropping random subsets of the data. This is to evaluate the robustness of various estimation procedures. We would expect (1) methods that do not adjust for the text to behave unpredictably, and (2) methods that do adjust for the text to be more robust.

These results support our first hypothesis: boosting treatment labels without text adjustment can behave unpredictably, as proxy-lex and T-boost both overestimate the true ATE. In other words, the predictions of both estimators grow further from the oracle as T^\hat{T}’s accuracy increases.

The results are mixed with respect to our second hypothesis. Both methods which adjust for the text (W-Adjust and TextCause ) consistently attenuate the true ATE, which is in line with Theorem 2. However, we find that TextCause , which makes use of T-boost and W-Adjust , may not always provide the highest quality ATE estimates in finite data regimes. Notably, when T^\hat{T} is less than 90% accurate, both proxy-lex and T-boost can produce higher-quality estimates than the proposed TextCause algorithm.

Note that all estimates quickly lose fidelity as the proxy T^\hat{T} becomes noisier. It rapidly becomes difficult for any method to recover the true ATE when the proxy T^\hat{T} is less than 80% accurate.

2 Application: Complaints to the Financial Protection Bureau

We proceed to offer an applied pilot study which seeks to answer, “how does the perceived politeness of a complaint affect the time it takes for that complaint to be addressed?” We consider complaints filed with the Consumer Financial Protection Bureau (CFPB).https://www.consumer-action.org/downloads/english/cfpb_full_dbase_report.pdf/ This is a government agency which solicits and handles complaints about financial products. When they receive a complaint it is forwarded to the relevant company. The time it takes for that company to process the complaint is recorded. Some submissions are handled quickly (<< 15 days) while others languish. This 15-day threshold is our outcome YY. We additionally adjust for an observed covariate CC that captures what product and company the complaint is about (mortgage or bank account). To reduce other potentially confounding effects, we pair each Y=1Y=1 complaint with the most similar Y=0Y=0 complaint according to cosine similarity of TF-IDF vectors Mozer et al. (2020). From this we select the 4,000 most similar pairs for a total of 8,000 complaints.

For our treatment (politeness), we use a state-of-the-art politeness detection package geared towards social scientists Yeomans et al. (2018). This package reports a score from a trained classifier using expert features of politeness and a hand-labeled dataset. We take examples in the top and bottom 25% of the scoring distribution to be our T^=1\hat{T}=1 and T^=0\hat{T}=0 examples and throw out all others. The final dataset consists of 4,000 complaints, topics, and outcomes.

We use the same training procedure and hyperparameters as Section 6.1, except now W-Adjust is trained for 9 epochs and each cross validation fold is of size 2,000.

Results are given in Figure 3 and suggest that perceived politeness may have an effect on reducing response time. We find that the effect size increases as we adjust for increasing amounts of information. The “unadjusted” approach which does not perform any adjustment produces the smallest ATE. “proxy-lex”, which only adjusts for covariates, indicated the second-smallest ATE. The W-Adjust and TextCause methods, which adjust for covariates and text, produced the largest ATE estimates. This suggests that there is a significant amount of confounding in real world studies, and the choice of estimator can yield highly varying conclusions.

Related Work

Our focus fits into a body of work on text-based causal inference that includes text as treatments (Egami et al., 2018; Fong and Grimmer, 2016; Grimmer and Fong, 2020; Wood-Doughty et al., 2018), text as outcomes Egami et al. (2018), and text as confounders (Roberts et al. (2020); Veitch et al. (2020); see Keith et al. (2020) for a review of that space). We build on Veitch et al. (2020), which proposed a BERT-based text adjustment method similar to our W-Adjust algorithm. This paper is related to work by Grimmer and Fong (2020), which discusses assumptions needed to estimate causal effects of text-based treatments in randomized controlled trials. There is also work on discovering causal structure in text, as topics with latent variable models Fong and Grimmer (2016) and as words and n-grams with adversarial learning Pryzant et al. (2018b) and residualization Pryzant et al. (2018a). There is also a growing body of applications in the social sciences Hall (2017); Olteanu et al. (2017); Saha et al. (2019); Mozer et al. (2020); Karell and Freedman (2019); Sobolev (2019); Zhang et al. (2020).

This paper also fits into a long-standing body of work on measurement error and causal inference (Pearl, 2012; Kuroki and Pearl, 2014; Buonaccorsi, 2010; Carroll et al., 2006; Shu and Yi, 2019; Oktay et al., 2019; Wood-Doughty et al., 2018). Most of this work deals with proxies for confounding variables. The present paper is most closely related to Wood-Doughty et al. (2018), which also deals with proxy treatments, but instead proposes an adjustment using the measurement model.

Conclusion

This paper addressed a setting of interest to NLP and social science researchers: estimating the causal effects of latent linguistic properties from observational data. We clarified critical ambiguities in the problem, showed how causal effects can be interpreted, presented a method, and demonstrated how it offers practical and theoretical advantages over the existing practice. We also release a package for performing text-based causal inferences.https://github.com/rpryzant/causal-text This work opens new avenues for further conceptual, methodological, and theoretical refinement. This includes improving non-lexicon based treatments, heterogeneous effects, overlap violations, counterfactual inference, ethical considerations, extensions to higher-dimensional outcomes and covariates, and benchmark datasets based on paired randomized controlled trials and observational studies.

Acknowledgements

This project recieved partial funding from the Stanford Data Science Institute, NSF Award IIS-1514268 and a Google Faculty Research Award. We thank Justin Grimmer, Stefan Wager, Percy Liang, Tatsunori Hashimoto, Zach Wood-Doughty, Katherine Keith, the Stanford NLP Group, and our anonymous reviewers for their thoughtful comments and suggestions.

References

Appendix A Proof of Theorem 1

The proof is complete because the estimand ψrea.\psi^{\textrm{rea.}} is simply μ(1)−μ(0)\mu(1)-\mu(0).

Appendix B C𝐶C-restriction

Appendix C Ablating T-boost

Appendix D Lemma 1 (used by Theorems 2 and 3)

Apply the law of total probability and definition of conditional independence to the causal graph given in Figure 3. ∎

Appendix E Proof of Theorem 2

Now we write the inner part using Lemma D, collect terms, and use the law of total probability to write everything in terms of misclassification probabilities:

Appendix F Theorem about the bias due to noisy proxies

Here we show that the naive estimand which does not adjust for the text,

can be arbitrarily biased away from the effect of interest, ψrea.\psi^{\textrm{rea.}}.

The α\alpha and β\beta terms are related to the error of the proxy label. This theorem says that correlations between the outcome and errors in the proxy can induce bias. Intuitively, this is similar to bias from confounding, though it is mathematically distinct. This means that even a highly accurate proxy label can result in highly misleading estimates. Proof:

Where the first equality is by the tower property, the second by inverse probability weighting, and the third Bayes’ rule. We continue by invoking Lemma 1:

The result follows immediately. □\square

Appendix G Deriving semi-oracle, a Causal Estimator for when P​(T|T^)𝑃conditional𝑇^𝑇P(T|\hat{T}) is known

This is an ATE estimator which assumes access to P(T^∣T)P(\hat{T}|T) instead of TT but is still unbiased. We derive this estimator using the “matrix adjustment” technique of Wood-Doughty et al. (2018); Pearl(2012). We start by decomposing the joint distribution

We can write this as a product between a matrix Mc,y(T^,T)=P(T^∣Y,C,T)\mathbf{M}_{c,y}(\hat{T},T)=P(\hat{T}|Y,C,T) and vector Vc,y(T)=P(Y,C,T)\mathbf{V}_{c,y}(T)=P(Y,C,T):

For our binary setting Mc,y\mathbf{M}_{c,y} is:

Under fairly broad conditions, M\mathbf{M} has an inverse, which allows us to reconstruct the joint distribution:

Note also that this expression is similar to τME\tau_{ME} in Wood-Doughty et al. (2018) except their error terms are of the form P(T ∣ T^)P(T\,|\,\hat{T}).