From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models

Shangbin Feng, Chan Young Park, Yuhan Liu, Yulia Tsvetkov

Introduction

Digital and social media have become a major source of political news dissemination (Hermida et al., 2012; Kümpel et al., 2015; Hermida, 2016) with unprecedentedly high user engagement rates (Mustafaraj and Metaxas, 2011; Velasquez, 2012; Garimella et al., 2018). The volume of online discourse surrounding polarizing issues—climate change, gun control, abortion, wage gaps, death penalty, taxes, same-sex marriage, and more—has been drastically growing in the past decade (Valenzuela et al., 2012; Rainie et al., 2012; Enikolopov et al., 2019). While online political engagement promotes democratic values and diversity of perspectives, these discussions also reflect and reinforce societal biases—stereotypical generalizations about people or social groups (Devine, 1989; Bargh, 1999; Blair, 2002). Such language constitutes a major portion of large language models’ (LMs) pretraining data, propagating biases into downstream models.

Hundreds of studies have highlighted ethical issues in NLP models Blodgett et al. (2020a); Field et al. (2021); Kumar et al. (2022) and designed synthetic datasets Nangia et al. (2020); Nadeem et al. (2021) or controlled experiments to measure how biases in language are encoded in learned representations Sun et al. (2019), and how annotator errors in training data are liable to increase unfairness of NLP models Sap et al. (2019). However, the language of polarizing political issues is particularly complex Demszky et al. (2019), and social biases hidden in language can rarely be reduced to pre-specified stereotypical associations Joseph and Morgan (2020). To the best of our knowledge, no prior work has shown how to analyze the effects of naturally occurring media biases in pretraining data on language models, and subsequently on downstream tasks, and how it affects the fairness towards diverse social groups. Our study aims to fill this gap.

As a case study, we focus on the effects of media biases in pretraining data on the fairness of hate speech detection with respect to diverse social attributes, such as gender, race, ethnicity, religion, and sexual orientation, and of misinformation detection with respect to partisan leanings. We investigate how media biases in the pretraining data propagate into LMs and ultimately affect downstream tasks, because discussions about polarizing social and economic issues are abundant in pretraining data sourced from news, forums, books, and online encyclopedias, and this language inevitably perpetuates social stereotypes. We choose hate speech and misinformation classification because these are social-oriented tasks in which unfair predictions can be especially harmful (Duggan, 2017; League, 2019, 2021).

To this end, grounded in political spectrum theories (Eysenck, 1957; Rokeach, 1973; Gindler, 2021) and the political compass test,https://www.politicalcompass.org/test we propose to empirically quantify the political leaning of pretrained LMs (§2). We then further pretrain language models on different partisan corpora to investigate whether LMs pick up political biases from training data. Finally, we train classifiers on top of LMs with varying political leanings and evaluate their performance on hate speech instances targeting different identity groups Yoder et al. (2022), and on misinformation detection with different agendas Wang (2017). In this way, we investigate the propagation of political bias through the entire pipeline from pretraining data to language models to downstream tasks.

Our experiments across several data domains, partisan news datasets, and LM architectures (§3) demonstrate that different pretrained LMs do have different underlying political leanings, reinforcing the political polarization present in pretraining corpora (§4.1). Further, while the overall performance of hate speech and misinformation detectors remains consistent across such politically-biased LMs, these models exhibit significantly different behaviors against different identity groups and partisan media sources. (§4.2).

The main contributions of this paper are novel methods to quantify political biases in LMs, and findings that shed new light on how ideological polarization in pretraining corpora propagates biases into language models, and subsequently into social-oriented downstream tasks. In §5, we discuss implications of our findings for NLP research, that no language model can be entirely free from social biases, and propose future directions to mitigate unfairness.

Methodology

We propose a two-step methodology to establish the effect of political biases in pretraining corpora on the fairness of downstream tasks: (1) we develop a framework, grounded in political science literature, to measure the inherent political leanings of pretrained language models, and (2) then investigate how the political leanings of LMs affect their performance in downstream social-oriented tasks.

While prior works provided analyses of political leanings in LMs (Jiang et al., 2022a; Argyle et al., 2022), they primarily focused on political individuals, rather than the timeless ideological issues grounded in political science literature. In contrast, our method is grounded in political spectrum theories (Eysenck, 1957; Rokeach, 1973; Gindler, 2021) that provide more nuanced perspective than the commonly used left vs. right distinction (Bobbio, 1996; Mair, 2007; Corballis and Beale, 2020) by assessing political positions on two axes: social values (ranging from liberal to conservative) and economic values (ranging from left to right).

The widely adopted political compass test,2 which is based on these theories, measures individuals’ leaning on a two-dimensional space by analyzing their responses to 62 political statements.The 62 political statements are presented in Table 13. We also evaluated on other political ideology questionnaires, such as the 8 values test, and the findings are similar. Participants indicate their level of agreement or disagreement with each statement, and their responses are used to calculate their social and economic scores through weighted summation. Formally, the political compass test maps a set of answers indicating agreement level {\textscstrongdisagree\{\textsc{strong disagree}, disagree, agree, \textscstrongagree}62\textsc{strong agree}\}^{62} to two-dimensional point (ssoc,seco)(s_{\textit{soc}},s_{\textit{eco}}), where the social score ssocs_{\textit{soc}} and economic score secos_{\textit{eco}} range from $$. We employ this test as a tool to measure the political leanings of pretrained language models.

We probe a diverse set of LMs to measure their alignment with specific political statements, including encoder and language generation models (decoder and autoregressive). For encoder-only LMs, we use mask filling with prompts derived from the political statements. We construct the following prompt: “Please respond to the following statement: [statement] I with this statement.” Then, pretrained LMs fill the mask and return 10 highest probability tokens. By comparing the aggregated probability of pre-defined positive (agree, support, endorse, etc.) and negative lexicons (disagree, refute, oppose, etc.) assigned by LMs, we map their answers to {\textscstrongdisagree\{\textsc{strong disagree}, disagree, agree, \textscstrongagree}\textsc{strong agree}\}. Specifically, if the aggregated probability of positive lexicon scores is larger than the negative aggregate by 0.3,The threshold was set empirically. Complete lists of positive and negative lexicons as well as the specific hyperparameters used for response mapping are listed in Appendix A.1. we deem the response as strong agree, and define strong disagree analogously.

We probe language generation models by conducting text generation based on the following prompt: “Please respond to the following statement: [statement] \\backslashn Your response:”. We then use an off-the-shelf stance detector (Lewis et al., 2019) to determine whether the generated response agrees or disagrees with the given statement. We use 10 random seeds for prompted generation, filter low-confidence responses using the stance detector, and average the stance detection scores for a more reliable evaluation.We established empirically that using multiple prompts results in more stable and consistent responses.

Using this framework, we aim to systematically evaluate the effect of polarization in pretraining data on the political bias of LMs. We thus train multiple partisan LMs through continued pretraining of existing LMs on data from various political viewpoints, and then evaluate how model’s ideological coordinates shift. In these experiments, we only use established media sources, because our ultimate goal is to understand whether “clean” pretraining data (not overtly hateful or toxic) leads to undesirable biases in downstream tasks.

2 Measuring the Effect of LM’s Political Bias on Downstream Task Performance

Armed with the LM political leaning evaluation framework, we investigate the impact of these biases on downstream tasks with social implications such as hate speech detection and misinformation identification. We fine-tune different partisan versions of the same LM architecture on these tasks and datasets and analyze the results from two perspectives. This is a controlled experiment setting, i.e. only the partisan pretraining corpora is different, while the starting LM checkpoint, task-specific fine-tuning data, and all hyperparameters are the same. First, we look at overall performance differences across LMs with different leanings. Second, we examine per-category performance, breaking down the datasets into different socially informed groups (identity groups for hate speech and media sources for misinformation), to determine if the inherent political bias in LMs could lead to unfairness in downstream applications.

Experiment Settings

We evaluate political biases of 14 language models: BERT (Devlin et al., 2019), RoBERTa (Liu et al., 2019), distilBERT (Sanh et al., 2019), distilRoBERTa, ALBERT (Lan et al., 2019), BART (Lewis et al., 2020), GPT-2 (Radford et al., 2019), GPT-3 (Brown et al., 2020), GPT-J (Wang and Komatsuzaki, 2021), LLaMA (Touvron et al., 2023), Alpaca (Taori et al., 2023), Codex (Chen et al., 2021), ChatGPT, GPT-4 (OpenAI, 2023) and their variants, representing a diverse range of model sizes and architectures. The specific versions and checkpoint names of each model are provided in Appendix C. For the stance detection model used for evaluating decoder-based language model responses, we use a BART-based model (Lewis et al., 2019) trained on MultiNLI (Williams et al., 2018).

To ensure the reliability of the off-the-shelf stance detector, we conduct a human evaluation on 110 randomly sampled responses and compare the results to those generated by the detector. The stance detector has an accuracy of 0.97 for LM responses with clear stances and high inter-annotator agreement among 3 annotators (0.85 Fleiss’ Kappa). Details on the stance detector, the response-to-agreement mapping process, and the human evaluation are in Appendix A.2.

Partisan Corpora for Pretraining

We collected partisan corpora for LM pretraining that focus on two dimensions: domain (news and social media) and political leaning (left, center, right). We used the POLITICS dataset (Liu et al., 2022a) for news articles, divided into left-leaning, right-leaning, and center categories based on Allsides.https://www.allsides.com For social media, we use the left-leaning and right-leaning subreddit lists by Shen and Rose (2021) and the PushShift API (Baumgartner et al., 2020). We also include subreddits that are not about politics as the center corpus for social media. Additionally, to address ethical concerns of creating hateful LMs, we used a hate speech classifier based on RoBERTa (Liu et al., 2019) and fine-tuned on the TweetEval benchmark (Barbieri et al., 2020) to remove potentially hateful content from the pretraining data. As a result, we obtained six pretraining corpora of comparable sizes: {\textscleft,\textsccenter,\textscright}×{\textscreddit,\textscnews}\{\textsc{left},\textsc{center},\textsc{right}\}\times\{\textsc{reddit},\textsc{news}\}. Details about pretraining corpora are in Appendix C. These partisan pretraining corpora are approximately the same size. We further pretrain RoBERTa and GPT-2 on these corpora to evaluate their changes in ideological coordinates and to examine the relationship between the political bias in the pretraining data and the model’s political leaning.

Downstream Task Datasets

We investigate the connection between models’ political biases and their downstream task behavior on two tasks: hate speech and misinformation detection. For hate speech detection, we adopt the dataset presented in Yoder et al. (2022) which includes examples divided into the identity groups that were targeted. We leverage the two official dataset splits in this work: Hate-Identity and Hate-Demographic. For misinformation detection, the standard PolitiFact dataset (Wang, 2017) is adopted, which includes the source of news articles. We evaluate RoBERTa (Liu et al., 2019) and four variations of RoBERTa further pretrained on reddit-left, reddit-right, news-left, and news-right corpora. While other tasks and datasets (Emelin et al., 2021; Mathew et al., 2021) are also possible choices, we leave them for future work. We calculate the overall performance as well as the performance per category of different LM checkpoints. Statistics of the adopted downstream task datasets are presented in Table 1.

Results and Analysis

In this section, we first evaluate the inherent political leanings of language models and their connection to political polarization in pretraining corpora. We then evaluate pretrained language models with different political leanings on hate speech and misinformation detection, aiming to understand the link between political bias in pretraining corpora and fairness issues in LM-based task solutions.

Figure 1 illustrates the political leaning results for a variety of vanilla pretrained LM checkpoints. Specifically, each original LM is mapped to a social score and an economic score with our proposed framework in Section 2.1. From the results, we find that:

Language models do exhibit different ideological leanings, occupying all four quadrants on the political compass.

Generally, BERT variants of LMs are more socially conservative (authoritarian) compared to GPT model variants. This collective difference may be attributed to the composition of pretraining corpora: while the BookCorpus (Zhu et al., 2015) played a significant role in early LM pretraining, Web texts such as CommonCrawlhttps://commoncrawl.org/the-data/ and WebText (Radford et al., 2019) have become dominant pretraining corpora in more recent models. Since modern Web texts tend to be more liberal (libertarian) than older book texts (Bell, 2014), it is possible that LMs absorbed this liberal shift in pretraining data. Such differences could also be in part attributed to the reinforcement learning with human feedback data adopted in GPT-3 models and beyond. We additionally observe that different sizes of the same model family (e.g. ALBERT and BART) could have non-negligible differences in political leanings. We hypothesize that the change is due to a better generalization in large LMs, including overfitting biases in more subtle contexts, resulting in a shift of political leaning. We leave further investigation to future work.

Pretrained LMs exhibit stronger bias towards social issues (yy axis) compared to economic ones (xx axis). The average magnitude for social and economic issues is 2.972.97 and 0.870.87, respectively, with standard deviations of 1.291.29 and 0.840.84. This suggests that pretrained LMs show greater disagreement in their values concerning social issues. A possible reason is that the volume of social issue discussions on social media is higher than economic issues (Flores-Saviaga et al., 2022; Raymond et al., 2022), since the bar for discussing economic issues is higher (Crawford et al., 2017; Johnston and Wronski, 2015), requiring background knowledge and a deeper understanding of economics.

We conducted a qualitative analysis to compare the responses of different LMs. Table 2 presents the responses of three pretrained LMs to political statements. While GPT-2 expresses support for “tax the rich”, GPT-3 Ada and Davinci are clearly against it. Similar disagreements are observed regarding the role of women in the workforce, democratic governments, and the social responsibility of corporations.

The Effect of Pretraining with Partisan Corpora

Figure 3 shows the re-evaluated political leaning of RoBERTa and GPT-2 after being further pretrained with 6 partisan pretraining corpora (§3):

LMs do acquire political bias from pretraining corpora. Left-leaning corpora generally resulted in a left/liberal shift on the political compass, while right-leaning corpora led to a right/conservative shift from the checkpoint. This is particularly noticeable for RoBERTa further pretrained on Reddit-left, which resulted in a substantial liberal shift in terms of social values (2.972.97 to −3.03-3.03). However, most of the ideological shifts are relatively small, suggesting that it is hard to alter the inherent bias present in initial pretrained LMs. We hypothesize that this may be due to differences in the size and training time of the pretraining corpus, which we further explore when we examine hyperpartisan LMs.

For RoBERTa, the social media corpus led to an average change of 1.60 in social values, while the news media corpus resulted in a change of 0.64. For economic values, the changes were 0.90 and 0.61 for news and social media, respectively. User-generated texts on social media have a greater influence on the social values of LMs, while news media has a greater influence on economic values. We speculate that this can be attributed to the difference in coverage (Cacciatore et al., 2012; Guggenheim et al., 2015): while news media often reports on economic issues (Ballon, 2014), political discussions on social media tend to focus more on controversial “culture wars” and social issues (Amedie, 2015).

Pre-Trump vs. Post-Trump

News and social media are timely reflections of the current sentiment of society, and there is evidence (Abramowitz and McCoy, 2019; Galvin, 2020; Hout and Maggio, 2021) suggesting that polarization is at an all-time high since the election of Donald Trump, the 45th president of the United States. To examine whether our framework detects the increased polarization in the general public, we add a pre- and post-Trump dimension to our partisan corpora by further partitioning the 6 pretraining corpora into pre- and post-January 20, 2017. We then pretrain the RoBERTa and GPT-2 checkpoints with the pre- and post-Trump corpora respectively. Figure 2 demonstrates that LMs indeed pick up the heightened polarization present in pretraining corpora, resulting in LMs positioned further away from the center. In addition to this general trend, for RoBERTa and the Reddit-right corpus, the post-Trump LM is more economically left than the pre-Trump counterpart. Similar results are observed for GPT-2 and the News-right corpus. This may seem counter-intuitive at first glance, but we speculate that it provides preliminary evidence that LMs could also detect the anti-establishment sentiment regarding economic issues among right-leaning communities, similarly observed as the Sanders-Trump voter phenomenon (Bump, 2016; Trudell, 2016).

Examining the Potential of Hyperpartisan LMs

Since pretrained LMs could move further away from the center due to further pretraining on partisan corpora, it raises a concern about dual use: training a hyperpartisan LM and employing it to further deepen societal divisions. We hypothesize that this might be achieved by pretraining for more epochs and with more partisan data. To test this, we further pretrain the RoBERTa checkpoint with more epochs and larger corpus size and examine the trajectory on the political compass. Figure 4 demonstrates that, fortunately, this simple strategy is not resulting in increasingly partisan LMs: on economic issues, LMs remain close to the center; on social issues, we observe that while pretraining does lead to some changes, training with more data for more epochs is not enough to push the models’ scores towards the polar extremes of 1010 or −10-10.

2 Political Leaning and Downstream Tasks

We compare the performance of five models: base RoBERTa and four RoBERTa models further pretrained with Reddit-left, News-left, Reddit-right, and News-right corpora, respectively. Table 3 presents the overall performance on hate speech and misinformation detection, which demonstrates that left-leaning LMs generally slightly outperform right-leaning LMs. The Reddit-right corpus is especially detrimental to downstream task performance, greatly trailing the vanilla RoBERTa without partisan pretraining. The results demonstrate that the political leaning of the pretraining corpus could have a tangible impact on overall task performance.

Performance Breakdown by Categories

In addition to aggregated performance, we investigate how the performance of partisan models vary for different targeted identity groups (e.g., Women, LGBTQ+) and different sources of misinformation (e.g., CNN, Fox). Table 4 illustrates a notable variation in the behavior of models based on their political bias. In particular, for hate speech detection, models with left-leaning biases exhibit better performance towards hate speech directed at widely-regarded minority groups such as lgbtq+ and black, while models with right-leaning biases tend to perform better at identifying hate speech targeting dominant identity groups such as men and white. For misinformation detection, left-leaning LMs are more stringent with misinformation from right-leaning media but are less sensitive to misinformation from left-leaning sources such as CNN and NYT. Right-leaning LMs show the opposite pattern. These results highlight the concerns regarding the amplification of political biases in pretraining data within LMs, which subsequently propagate into downstream tasks and directly impact model (un)fairness.

Table 5 provides further qualitative analysis and examples that illustrate distinctive behaviors exhibited by pretrained LMs with different political leanings. Right-leaning LMs overlook racist accusations of “race mixing with asians,” whereas left-leaning LMs correctly identify such instances as hate speech. In addition, both left- and right-leaning LMs demonstrate double standards for misinformation regarding the inaccuracies in comments made by Donald Trump or Bernie Sanders.

Reducing the Effect of Political Bias

Our findings demonstrate that political bias can lead to significant issues of fairness. Models with different political biases have different predictions regarding what constitutes as offensive or not, and what is considered misinformation or not. For example, if a content moderation model for detecting hate speech is more sensitive to offensive content directed at men than women, it can result in women being exposed to more toxic content. Similarly, if a misinformation detection model is excessively sensitive to one side of a story and detects misinformation from that side more frequently, it can create a skewed representation of the overall situation. We discuss two strategies to mitigate the impact of political bias in LMs.

The experiments in Section 4.2 show that LMs with different political biases behave differently and have different strengths and weaknesses when applied to downstream tasks. Motivated by existing literature on analyzing different political perspectives in downstream tasks (Akhtar et al., 2020; Flores-Saviaga et al., 2022), we propose using a combination, or ensemble, of pretrained LMs with different political leanings to take advantage of their collective knowledge for downstream tasks. By incorporating multiple LMs representing different perspectives, we can introduce a range of viewpoints into the decision-making process, instead of relying solely on a single perspective represented by a single language model. We evaluate a partisan ensemble approach and report the results in Table 6, which demonstrate that partisan ensemble actively engages diverse political perspectives, leading to improved model performance. However, it is important to note that this approach may incur additional computational cost and may require human evaluation to resolve differences.

Strategic Pretraining

Another finding is that LMs are more sensitive towards hate speech and misinformation from political perspectives that differ from their own. For example, a model becomes better at identifying factual inconsistencies from New York Times news when it is pretrained with corpora from right-leaning sources.

This presents an opportunity to create models tailored to specific scenarios. For example, in a downstream task focused on detecting hate speech from white supremacy groups, it might be beneficial to further pretrain LMs on corpora from communities that are more critical of white supremacy. Strategic pretraining might have great improvements in specific scenarios, but curating ideal scenario-specific pretraining corpora may pose challenges.

Our work opens up a new avenue for identifying the inherent political bias of LMs and further study is suggested to better understand how to reduce and leverage such bias for downstream tasks.

Related Work

Studies have been conducted to measure political biases and predict the ideology of individual users (Colleoni et al., 2014; Makazhanov and Rafiei, 2013; Preoţiuc-Pietro et al., 2017), news articles (Li and Goldwasser, 2019; Feng et al., 2021; Liu et al., 2022b; Zhang et al., 2022), and political entities (Anegundi et al., 2022; Feng et al., 2022). As extensive research has shown that machine learning models exhibit societal and political biases (Zhao et al., 2018; Blodgett et al., 2020b; Bender et al., 2021; Ghosh et al., 2021; Shaikh et al., 2022; Li et al., 2022; Cao et al., 2022; Goldfarb-Tarrant et al., 2021; Jin et al., 2021), there has been an increasing amount of research dedicated to measuring the inherent societal bias of these models using various components, such as word embeddings (Bolukbasi et al., 2016; Caliskan et al., 2017; Kurita et al., 2019), output probability (Borkan et al., 2019), and model performance discrepancy (Hardt et al., 2016).

Recently, as generative models have become increasingly popular, several studies have proposed to probe political biases (Liu et al., 2021; Jiang et al., 2022b) and prudence (Bang et al., 2021) of these models. Liu et al. (2021) presented two metrics to quantify political bias in GPT2 using a political ideology classifier, which evaluate the probability difference of generated text with and without attributes (gender, location, and topic). Jiang et al. (2022b) showed that LMs trained on corpora written by active partisan members of a community can be used to examine the perspective of the community and generate community-specific responses to elicit opinions about political entities. Our proposed method is distinct from existing methods as it can be applied to a wide range of LMs including encoder-based models, not just autoregressive models. Additionally, our approach for measuring political bias is informed by existing political science literature and widely-used standard tests.

Impact of Model and Data Bias on Downstream Task Fairness

Previous research has shown that the performance of models for downstream tasks can vary greatly among different identity groups (Hovy and Søgaard, 2015; Buolamwini and Gebru, 2018; Dixon et al., 2018), highlighting the issue of fairness (Hutchinson and Mitchell, 2019; Liu et al., 2020). It is commonly believed that annotator (Geva et al., 2019; Sap et al., 2019; Davani et al., 2022; Sap et al., 2022) and data bias (Park et al., 2018; Dixon et al., 2018; Dodge et al., 2021; Harris et al., 2022) are the cause of this impact, and some studies have investigated the connection between training data and downstream task model behavior (Gonen and Webster, 2020; Li et al., 2020; Dodge et al., 2021). Our study adds to this by demonstrating the effects of political bias in training data on downstream tasks, specifically in terms of fairness. Previous studies have primarily examined the connection between data bias and either model bias or downstream task performance, with the exception of Steed et al. (2022). Our study, however, takes a more thorough approach by linking data bias to model bias, and then to downstream task performance, in order to gain a more complete understanding of the effect of social biases on the fairness of models for downstream tasks. Also, most prior work has primarily focused on investigating fairness in hate speech detection models, but our study highlights important fairness concerns in misinformation detection that require further examination.

Conclusion

We conduct a systematic analysis of the political biases of language models. We probe LMs using prompts grounded in political science and measure models’ ideological positions on social and economic values. We also examine the influence of political biases in pretraining data on the political leanings of LMs and investigate the model performance with varying political biases on downstream tasks, finding that LMs may have different standards for different hate speech targets and misinformation sources based on their political biases.

Our work highlights that pernicious biases and unfairness in downstream tasks can be caused by non-toxic data, which includes diverse opinions, but there are subtle imbalances in data distributions. Prior work discussed data filtering or augmentation techniques as a remedy Kaushik et al. (2019); while useful in theory, these approaches might not be applicable in real-world settings, running the risk of censorship and exclusion from political participation. In addition to identifying these risks, we discuss strategies to mitigate the negative impacts while preserving the diversity of opinions in pretraining data.

Limitations

In this work, we leveraged the political compass test as a test bed to probe the underlying political leaning of pretrained language models. While the political compass test is a widely adopted and straightforward toolkit, it is far from perfect and has several limitations: 1) In addition to a two-axis political spectrum on social and economic values (Eysenck, 1957), there are numerous political science theories (Blattberg, 2001; Horrell, 2005; Diamond and Wolf, 2017) that support other ways of categorizing political ideologies. 2) The political compass test focuses heavily on the ideological issues and debates of the western world, while the political landscape is far from homogeneous around the globe. (Hudson, 1978) 3) There are several criticisms of the political compass test: unclear scoring schema, libertarian bias, and vague statement formulation (Utley, 2001; Mitchell, 2007). However, we present a general methodology to probe the political leaning of LMs that is compatible with any ideological theories, tests, and questionnaires. We encourage readers to use our approach along with other ideological theories and tests for a more well-rounded evaluation.

Probing Language Models

For encoder-based language models, our approach of mask in-filling is widely adopted in numerous existing works (Petroni et al., 2019; Lin et al., 2022). For language generation models, we curate prompts, conduct prompted text generation, and employ a BART-based stance detector for response evaluation. An alternative approach would be to explicitly frame it as a multi-choice question in the prompt, forcing pretrained language models to choose from strong agree, agree, disagree, and strong disagree. These two approaches have their respective pros and cons: our approach is compatible with all LMs that support text generation and is more interpretable, while the response mapping and the stance detector could be more subjective and rely on empirical hyperparameter settings; multi-choice questions offer direct and unequivocal answers, while being less interpretable and does not work well with LMs with fewer parameters such as GPT-2 (Radford et al., 2019).

Fine-Grained Political Leaning Analysis

In this work, we "force" each pretrained LM into its position on a two-dimensional space based on their responses to social and economic issues. However, political leaning could be more fine-grained than two numerical values: being liberal on one issue does not necessarily exclude the possibility of being conservative on another, and vice versa. We leave it to future work on how to achieve a more fine-grained understanding of LM political leaning in a topic- and issue-specific manner.

Ethics Statement

The authors of this work are based in the U.S., and our framing in this work, e.g., references to minority identity groups, reflects this context. This viewpoint is not universally applicable and may vary in different contexts and cultures.

Misuse Potential

In this paper, we showed that hyperpartisan LMs are not simply achieved by pretraining on more partisan data for more epochs. However, this preliminary finding does not exclude the possibility of future malicious attempts at creating hyperpartisan language models, and some might even succeed. Training and employing hyperpartisan LMs might contribute to many malicious purposes, such as propagating partisan misinformation or adversarially attacking pretrained language models (Bagdasaryan and Shmatikov, 2022). We will refrain from releasing the trained hyperpartisan language model checkpoints and will establish access permission for the collected partisan pretraining corpora to ensure its research-only usage.

Interpreting Downstream Task Performance

While we showed that pretrained LMs with different political leanings could have different performances and behaviors on downstream tasks, this empirical evidence should not be taken as a judgment of individuals and communities with certain political leanings, rather than a mere reflection of the empirical behavior of pretrained LMs.

Authors’ Political Leaning

Although the authors strive to conduct politically impartial analysis throughout the paper, it is not impossible that our inherent political leaning has impacted experiment interpretation and analysis in unperceived ways. We encourage the readers to also examine the models and results by themselves, or at least be aware of this possibility.

Acknowledgements

We thank the reviewers, the area chair, Anjalie Field, Lucille Njoo, Vidhisha Balachandran, Sebastin Santy, Sneha Kudugunta, Melanie Sclar, and other members of Tsvetshop, and the UW NLP Group for their feedback. This material is funded by the DARPA Grant under Contract No. HR001120C0124. We also gratefully acknowledge support from NSF CAREER Grant No. IIS2142739, the Alfred P. Sloan Foundation Fellowship, and NSF grants No. IIS2125201, IIS2203097, and IIS2040926. Any opinions, findings and conclusions or recommendations expressed in this material are those of the authors and do not necessarily state or reflect those of the United States Government or any agency thereof.

References

Appendix A Probing Language Models (cont.)

We used mask filling to probe the political leaning of encoder-based language models (e.g. BERT (Devlin et al., 2019) and RoBERTa (Liu et al., 2019)). Specifically, we retrieve the top-10 probable token for mask filling, aggregate the probability of positive and negative words, and set a threshold to map them to {\textscstrongdisagree\{\textsc{strong disagree}, disagree, agree, \textscstrongagree}\textsc{strong agree}\}. A complete list of positive and negative words adopted is presented in Table 7, which is obtained after manually examining the output probabilities of 100 examples. We then compare the probability of positive words and negative words to settle agree v.s. disagreee, then normalize and use 0.3 in probability difference as a threshold for whether that response is strongly or not.

A.2 Decoder-Based LMs

We use prompted text generation and a stance detector to evaluate the political leaning of decoder-based language models (e.g. GPT-2 (Radford et al., 2019) and GPT-3 (Brown et al., 2020)). The goal of stance detection is to judge the LM-generated response and map it to {\textscstrongdisagree\{\textsc{strong disagree}, disagree, agree, \textscstrongagree}\textsc{strong agree}\}. To this end, we employed the facebook/bart-large-mnli checkpoint on Huggingface Transformers, which is BART (Lewis et al., 2019) fine-tuned on the multiNLI dataset (Williams et al., 2018), to initialize a zero-shot classification pipeline of agree and disagree, evaluating whether the response entails agreement or disagreement. We further conduct a human evaluation of the stance detector: we select 110 LM-generated responses, annotate the responses, and compare the human annotations with the results of the stance detector. The three annotators are graduate students in the U.S., with prior knowledge both in NLP and U.S. politics. This human evaluation answers a few key questions:

Do language models provide clear responses to political propositions? Yes, since 80 of the 110 LM responses provide responses with a clear stance. The Fleiss’ Kappa of annotation agreement is 0.85, which signals strong agreement among annotators regarding the stance of LM responses.

Is the stance detector accurate? Yes, on the 80 LM responses with a clear stance, the BART-based stance detector has an accuracy of 97%. This indicates that the stance detector is reliable in judging the agreement of LM-generated responses.

How do we deal with unclear LM responses? We observed that the 30 unclear responses have an average stance detection confidence of 0.76, while the 80 unclear responses have an average confidence of 0.90. This indicates that the stance detector’s confidence could serve as a heuristic to filter out unclear responses. As a result, we retrieve the top-10 probable LM responses, remove the ones with lower than 0.9 confidence, and aggregate the scores of the remaining responses.

To sum up, we present a reliable framework to probe the political leaning of pretrained language models. We commit to making the code and data publicly available upon acceptance to facilitate the evaluation of new and emerging LMs.

Appendix B Recall and Precision

Following previous works (Sap et al., 2019), we additionally report false positives and false negatives through precision and recall in Table 12.

Appendix C Experiment Details

We provide details about specific language model checkpoints used in this work in Table 10. We present the dataset statistics for the social media corpora in Table 8, while we refer readers to Liu et al. (2022b) for the statistics of the news media corpora.

Appendix D Stability Analysis

Pretrained language models are sensitive to minor changes and perturbations in the input text (Li et al., 2021; Wang et al., ), which may in turn lead to instability in the political leaning measuring process. In the experiments, we made minor edits to the prompt formulation in order to best elicit political opinions of diverse language models. We further examine whether the political opinion of language models stays stable in the face of changes in prompts and political statements. Specifically, we design 6 more prompts to investigate the sensitivity toward prompts. We similarly use 6 paraphrasing models to paraphrase the political propositions and investigate the sensitivity towards paraphrasing. We present the results of four LMs in Figure 5, which illustrates that GPT-3 DaVinci (Brown et al., 2020) provides the most consistent responses, while the political opinions of all pretrained LMs are moderately stable.

We further evaluate the stability of LM political leaning with respect to minor changes in prompts. We write 7 different prompts formats, prompt LMs separately, and present the results in Figure 6. It is demonstrated that GPT-3 DaVinci provides the most consistent responses towards prompt changes, while the political opinions of all pretrained LMs are moderately stable.

For paraphrasing, we adopted three models: Vamsi/T5_Paraphrase_Paws based on T5 (Raffel et al., 2020), eugenesiow/bart-paraphrase based on BART (Lewis et al., 2019), tuner007/pegasus_paraphrase based on PEGASUS (Zhang et al., 2020), and three online paraphrasing tools: Quill Bot https://quillbot.com/, Edit Pad https://www.editpad.org/, and Paraphraser https://www.paraphraser.io/. For prompts, we present the 7 manually designed prompts in Table 11.

Appendix E Qualitative Analysis (cont.)

We conduct qualitative analysis and present more hate speech examples where pretrained LMs with different political leanings beg to differ. Table 14 presents more examples for hate speech detection. It is demonstrated that pretrained LMs with different political leanings do have vastly different behavior facing hate speech targeting different identities.

Appendix F Hyperparameter Settings

We further pretrained LM checkpoints on partisan corpora and fine-tuned them on downstream tasks. We present hyperparameters for the pretraining and fine-tuning stage in Table 9. We mostly follow the hyperparameters in Gururangan et al. (2020) for the pretraining stage. The default hyperparameters on Huggingface Transformers are adopted if not included in Table 9.

Appendix G Computational Resources

We used a GPU cluster with 16 NVIDIA A40 GPUs, 1988G memory, and 104 CPU cores for the experiments. Pretraining roberta-base and GPT-2 on the partisan pretraining corpora takes approximately 48 and 83 hours. Fine-tuning the partisan LMs takes approximately 30 and 20 minutes for the hate speech detection and misinformation identification datasets.

Appendix H Scientific Artifacts

We leveraged many open-source scientific artifacts in this work, including pytorch (Paszke et al., 2019), pytorch lightning (Falcon and The PyTorch Lightning team, 2019), HuggingFace transformers (Wolf et al., 2020), sklearn (Pedregosa et al., 2011), NumPy (Harris et al., 2020), NLTK (Bird et al., 2009), and the PushShift API https://github.com/pushshift/api. We commit to making our code and data publicly available upon acceptance to facilitate reproduction and further research.