NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
Kristina Gligoric, Myra Cheng, Lucia Zheng, Esin Durmus, Dan Jurafsky
Introduction
The use-mention distinction is the difference between using words (Bananas have a peel) and mentioning them (“Bananas” has 7 letters, or Dan said “Bananas”). The distinction has long been important in the philosophy of language Sperber and Wilson (1981); Saka (1998), where discussions date back to Tarski (1931) and Quine (1940).
The ability to correctly make this distinction is particularly relevant to dealing with problematic language online. While detecting problematic text has become a standard task in NLP, there has been less consideration of text that might mention harmful content without directly using it. Mentioning problematic content is critical to language like counterspeech that challenges or opposes hateful or misleading narratives (Table 1) Wright et al. (2017); Mun et al. (2023); Hangartner et al. (2021); Ecker et al. (2022). And mentioning is similarly crucial in media and academic reporting where researchers and journalists report harmful content Kirk et al. (2022), in educational settings where problematic material is invoked for educational purposes, in disclosures in legal settings where harmful statements need to be quoted Henderson et al. (2022), and in personal testimonies Wexler et al. (2019).
In this work, we focus on the first of these: online counterspeech, defined as speech produced by users of online platforms to counteract harmful speech of others. This effort seeks to stop the spread of harmful speech, mitigate its effects, discourage its recurrence, and provide support to both the targeted individuals and those joining in the counterspeech efforts Garland et al. (2022).
Counterspeech statements often involve referring to or quoting problematic content Vidgen et al. (2021). Errors in distinguishing use from mention might therefore lead to failures in downstream classification, making counterspeech statements more likely to be misclassified as harmful by modern NLP systems, as shown in Figure 1. Yet counterspeech helps curb online abuse Bonaldi et al. (2022) and make online spaces safer Siegel and Badaan (2020). Thus, erroneously classifying counterspeech as problematic leads to content removal with significant implications: misclassification erases opportunities to rectify false narratives, and in doing so, risks further censoring those already most affected by harmful language Sahoo et al. (2022); Rahman (2012); Park et al. (2018).
But addressing the use-mention distinction and assessing its impact on downstream tasks is challenging. First, reasoning about the distinction itself has been hard, due to the lack of datasets, resources, and quantitative measurement methods. The few prior studies have been small and limited to linguistic features like particular mention verbs Wilson (2011b). Second, online counterspeech generally occurs in informal contexts, where markings that formally indicate mention, such as quotation marks or italics, are often missing Wilson (2011a). Finally, since mentioned language is less frequent than use Wilson (2011b), it is easy for researchers to overlook the downstream performance of NLP systems on mentioned language.
Motivated by these technical challenges, failure cases in counterspeech (Figure 1 and Table 1), and literature on lexical and topic biases in harmful content detection Dixon et al. (2018); Ethayarajh et al. (2022), we make the following hypotheses:
NLP models fail to distinguish use from mention in counterspeech.
H2
Failure in distinguishing use from mention impacts downstream tasks like hate speech and misinformation detection.
H3
What is considered ‘permissible’ by hate speech or misinformation classifiers is influenced by the presence of identity terms and other targeted entities, as well as the strength of the stance expressed toward mentioned language.
H4
Downstream performance can be improved by teaching the use-mention distinction to explicitly encode the treatment of mentioned language.
Testing these hypotheses, our contributions include:
(1) Use-mention tasks
We formalize two tasks with challenging use-mention examples: (1) use versus mention classification and (2) downstream hate speech and misinformation detection (Sec.3).
(2) Failure analyses
We identify use-mention distinction failures, show that errors propagate to downstream tasks, and trace failures to specific target entities (misinformation-related and identity terms) and to the strength of the expressed stance (Sec.4–4.3).
(3) Mitigations
We investigate prompting mitigations and implications for downstream tasks. We show that our interventions lead to a significant reduction in error (Sec. 4.4).
To support the further evaluation and development of models, our code is available at https://github.com/kristinagligoric/use-mention. Datasets are publicly available and can be used for research purposes.
Background
Building upon social science literature showing that counternarratives are effective against hate speech (Andrews, 2002; Benesch, 2014; Schieb and Preuss, 2016; Garland et al., 2022), previous work in NLP and HCI has studied counterspeech from various perspectives. Scholars have curated datasets of counterspeech from different sources, including experts, NGO workers, and social media comments (Mathew et al., 2019; Chung et al., 2019; Qian et al., 2019; Garland et al., 2020; Fanton et al., 2021). Others have investigated different types of counterspeech strategies and performed user studies to evaluate the effectiveness of machine-generated counterspeech (Mun et al., 2023; Fraser et al., 2023). Our work also builds upon existing models for counterspeech detection (Garland et al., 2020).
Content moderation policies on counterspeech mentions
Several online platforms acknowledge the importance of counterspeech in their content policies. At the time of writing, TikTok’s Community Principles states “We do not allow language or behavior that harasses, humiliates, threatens, or doxxes anyone. This also includes responding to such acts with retaliatory harassment (but excludes non-harassing counter speech)” TikTok (2023). Similarly, Facebook publisher and creator guidelines state “We know that many publishers use Facebook to challenge ideas, institutions and practices. Such discussion can promote debate and greater understanding”. The guidelines also provide advice on how to write counterspeech to avoid mislabeling as hate speech Meta (2023). Development of models that enable enforcement of such policies is thus pressing.
While previous work has not explored the impact that the use-mention distinction has on online text classification tasks, mentions of toxic phrases were described as one class of false positive errors in toxic comment classification Van Aken et al. (2018), e.g., “I deleted the
The use-mention distinction
The use-mention distinction has been studied in philosophy Sperber and Wilson (1981), computational linguistics Wilson (2010, 2011b); Behzad et al. (2023), and HCI Anderson et al. (2002). In general, mention is defined as follows:
“for a token or a set of tokens in a sentence , if refers to a property of the token , then is an instance of mention.” Wilson (2010)
While many facts about a token can be a mention property, in our domain of counterspeech we focus on two properties of mentioned language Sandhan et al. (2023); Wilson (2011b): attributed language (e.g., mentions to refer to the quotes of original source or stance towards the source) and words or phrases as themselves (e.g., mentions of words to refer to stance towards their use).See the Limitations section on other types of mentioned language which do not lead as clearly to practical harms.
The use-mention distinction and related tasks
Like the use-mention distinction, other NLP tasks bear on speaker intent, such as those related to factuality Saurí and Pustejovsky (2012, 2009); Murzaku et al. (2022) and committed belief Prabhakaran et al. (2010, 2015); De Marneffe et al. (2019). In the context of mentioned language, such tasks aim to directly take into account the speaker’s intention when mentioning specific phrases, and in particular, whether the speech act commits the speaker to the truth or factuality of the expressed proposition. However, in both used and mentioned language, statements do not necessarily commit the speaker to the truth of the source statement or to its factuality. Moreover, harmful language is often implicit, formalized as questions and nuanced statements that need not constitute a committed belief or be factual either. The distinction between using and mentioning is thus related to but distinct from factuality and committed belief.
Lastly, mentioned language is also an instance of metalanguage Behzad et al. (2023); Perlis et al. (1998); Wilson (2012, 2013), and our task builds upon work on similar tasks like distinguishing whether personal names are being used to mention or address Prabhakaran et al. (2023).
Methods
We focus on language indicative of hateful speech or misinformation, two frequent and connected types of harmful online content Mosleh et al. (2024), which can be used or mentioned to express disapproving attitude towards it (as illustrated in Fig 1). We hypothesize that use-mention distinction failures cause harmful misclassifications of counterspeech on downstream tasks. To test this hypothesis, we operationalize two tasks: the use-mention distinction and downstream classification.
For a given text, the task is to classify whether hateful/misinformative language is used or whether it is mentioned. True use are statements which use hate or misinformation, and are counterspeech mentions. For each text , , we classify the text as either use (positive class) or mention (negative class). Metrics we report are false positive rate and false negative rate in detecting use, and average error rate, capturing the average of the two rates. False positives are mentions misclassified as uses, while false negatives are defined as uses misclassified as mentions. On Task 1, false positives (mistaking mention for use) are the errors of primary interest due to their hypothesized impact on downstream tasks.
Task 2: Downstream classification where use-mention distinction matters
We address two important downstream sub-tasks: hate speech detection and misinformation detection, with challenging use-mention examples. Similarly, for each text , , we classify each statement as either the positive (“misinformation” and “hate speech”) or negative class (“not misinformation” and “not hate”) on the downstream task. Since the standard metric is false positive rate capturing how often non-harmful content is misclassified as harmful Dixon et al. (2018); Markov and Daelemans (2021), we report false positive rate (mentions misclassified as “misinformation” or “hate”). We also report false negative rate (uses misclassified as “not misinformation” or “not hate”), and average error rate, capturing the average of the two rates. An ideal model would classify all as “misinformation” or “hate speech” (0% false negative rate), while all counterspeech statements would be classified as “not misinformation” or “not hate” (0% false positive rate). On Task 2, false positive rate on is the pragmatic concern central to our investigation as it captures the censorship rate.
2 Datasets
For this task we rely on two datasets: Knowledge-grounded hate countering Chung et al. (2021) and Multi-Target Counternarratives Fanton et al. (2021). The datasets contain pairs of (hateful statement, counterspeech), illustrated in Table 1. We select counterspeech statements written by human experts.
Countering misinformation
For this task we leverage misinformation counternarratives He et al. (2023). The dataset contains pairs of (misinformative statement, counterspeech), illustrated in Table 1. Misinformative statements were posted on social media, while counter-responses are a mix of naturally occurring social media posts and counterspeech statements written by recruited human participants.
Focal tokens
To confirm that mentioned language is indeed relevant to the practical case of counterspeech, we verified that counterspeech mentions contain language from the original true use sample it addresses, which we refer to as focal tokens. Across pairs , we computed the length of the longest common substring using a dynamic programming algorithm Suzgun et al. (2023b). We found substantial overlap (as in examples in Table 1, focal tokens), with average words in the longest common substring. Additionally, for both datasets, we manually verified that focal tokens are used in true use statements (and not mentioned), and that counterspeech is not using, but mentioning (see Appendix, Sec. A for details).
3 Models
For the use-mention task, we tested the two best performing GPT models at the time of writing (gpt-3.5-turbo, gpt-4), as well as gpt-3.5 instruct, the non-RLHF legacy variant, using zero-shot prompting. For downstream tasks, we tested these three models as well as four widely-used models for hate speech and misinformation detection: Perspective Perspective (2023), Toxigen Hartvigsen et al. (2022), RoBERTa Liu et al. (2019), and RoBERTa fake news Ahmed et al. (2017).
For prompting, we use default parameters (temperature=1) and max output token length 1 (outputting either A or B). The classification prompt includes the instruction and a definition of the classes (for prompt variants and complete prompt text see Appendix, Section D).
Results
How well do the models distinguish whether problematic language is used or mentioned? Across the models (Table 2), average error rates are high (between 12.22% and 16.38% for hate speech and between 13.64% and 37.22% for misinformation), suggesting that state-of-the-art large language models struggle to distinguish use from mention in domains where the distinction matters. The best-performing model, gpt-4, still has a very high false positive rate, mistaking mention for use; mentions are misclassified as use in 20.00% of hate speech and 23.44% of misinformation counternarratives. In these settings, mistaking mention for use leads to the consequential harm of censoring useful counterspeech.
2 Downstream content classification (H2)
The prior section showed that state-of-the-art systems often fail at the use-mention distinction. Here we test the impact on hate speech and misinformation detection, two downstream classification tasks where the use-mention distinction matters (Table 3).
First, on the hate speech detection task, we find that misclassification of counterspeech as hateful is relatively frequent: the popular Toxigen and Perspective API have a false positive rate on counterspeech of over 20%, and recent models still have many false positives, although somewhat reduced, e.g., gpt-3.5-turbo has FPR 11.11%. Regarding the average error rate, gpt-4 is the best-performing model (false positive rate on counterspeech 8.89%).
Second, on the misinformation detection task, misclassification of counterspeech as misinformative is relatively frequent: gpt-3.5 models have over 20% false positive rate on counterspeech, and substantial errors persist even with the best system, gpt-4 (10.21% false positive rate classifying counterspeech mentions as misinformation). We also note that two investigated models, Togixen and RoBERTa fake news, have average error rates above 50%.
Error propagation to downstream tasks
As a further test of the hypothesis that errors in distinguishing use from mention cause these failures in downstream tasks, we assess whether examples that cause errors downstream (Task 2) are also misclassified in use-mention classification (Task 1). We do this by testing how the downstream error rate on Task 2 (hate speech detection and misinformation detection) associates with the error rate on Task 1 (use vs. mention). For gpt-4, for example, we contrast the error rate on Task 2 of 15.78% (aggregated over hate speech and misinformation detection) for statements in which the Task 1 classifier failed at distinguishing use from mention, versus the error on Task 2 of 4.54% (aggregated) for statements in which the Task 1 classifier correctly distinguished use from mention, comparing these rates using the chi-squared test.
We find that, among samples where mentioning counterspeech is misclassified as use, downstream misclassification is higher across the three tested models (all ; Table 4). Results disaggregated by hate speech and misinformation detection are listed in the Appendix (Table 10).
In summary, we find that errors in use-mention distinction do propagate to downstream tasks.
3 Why is the distinction hard (H3)?
Counterspeech that opposes harmful narratives is by definition not harmful. What linguistic aspects of this counterspeech cause downstream NLP tools to incorrectly label counterspeech as harmful? Or, viewed from the other perspective, what terms do NLP tools make permissible as mentions?
We know that toxicity detection algorithms rely heavily on the presence of identity mentions to make their predictions Zhou et al. (2021). For the hate speech task, we therefore hypothesize that identity terms will impact when mentioned language is permissible. To test this hypothesis, we stratify error rate in classifying counterspeech as hate speech by target identity (as labeled in the metadata of the dataset Chung et al. (2021)).
We find that errors vary widely depending on the targeted identity (e.g., gpt-3.5-turbo: Jewish 14.15%, people of color 9.09%, Muslims 6.80%, LGBT+ 6.77%, while for the other groups it is less than 5%; Table 5). Even the most recent large language models’ treatment of counterspeech similarly varies depending on the mentioned identity. Our results suggest that systems treat the mere mention of certain identity terms as impermissible (see Sec.5).
Misinformation: COVID-19 terms and strength in stance towards the embedded language
To understand why counterspeech mentioning misinformation is detected as misinformation, we use the Fightin’ Words method Monroe et al. (2008) to measure statistically significant differences in tokens between two sets—counterspeech mentions classified as misinformation vs. not classified as misinformation—after controlling for variance in words’ frequencies.
We compute the top differentiating words for versus , where
respectively, where are all counterspeech statements and is misinformation classification by classifier .
We find that terms relating to controversial topics are often misclassified; top terms for include “gene therapy”, “mRNA”, “vaccine”, “DNA”, and “CDC”, suggesting that counterspeech mentioning COVID vaccination is often misclassified. Certain terms associated with COVID-19 misinformation are not permissible even when mentioned.
We hypothesize that due to interactions with safety features that prevent large language models from generating health disinformation Menz et al. (2024), the mere presence of specific terms related to COVID vaccination is linked with misclassification of counterspeech as misinformative. This finding that terms related to COVID are treated as ‘impermissible’ when mentioned parallels our finding that mentions of certain demographic identities are impermissible to hate speech detectors.
We also find that top terms for are related to expressing a strong stance against misinformation in the surrounding context and the strength of meta language (as indicated by terms such as “fake news”, “lying”, “lies”, “misleading”). Downstream classification has fewer errors when the disagreement in mentioning statements is not subtle.
Distancing by using quotation marks
Counterspeech that uses verbatim quotes typically involves more severe language, as quotes enable distancing Wilson (2011a). We hypothesize that such texts are more censored. Indeed, we find for both hate speech and misinformation that the counterspeech containing quotation marks is more frequently misclassified as harmful (Table 7).
Summary of Error Analysis
In summary, misclassification of counterspeech mentions as harmful is influenced by over-reliance on (a) surface terms such as identity words and specific COVID-19-related terms, and (b) notions of strength in stance towards mentioned language.
4 How can we teach the distinction (H4)?
Informed by the analyses revealing that errors propagate from the inability to distinguish use from mention into misclassification on the downstream tasks, we explore a set of prompting mitigations to reduce downstream mistakes. In particular, we explore ways to teach the use-mention distinction through controlled prompting.
To that end, we designed and tested CoT prompting mitigation. In CoT mitigation, we (1) embed the definition of use-mention distinction and an instruction specifying that mention of hateful or misinformative language does not imply that the text is hateful or misinformative. We use prompt formats inspired by BigBench CoT Suzgun et al. (2023a) and (2) follow the process prompting the LLM to “think step-by-step” Wei et al. (2022) and use the answer extraction prompt “so the answer is” to the generated rationale to extract the final answer. We also (3) include few-shot examples of mentioned and used language where the first step considers whether potentially problematic language is used or mentioned, before making the classification in the second step (for complete prompt text, see Appendix, Table 13).
Furthermore, we perform an ablation study isolating the impact of use-mention examples alone (Few shot) and embedding the instruction to make the use-mention distinction alone (Mitigation).
For the best-performing model (gpt-4), we tested the prompting mitigation on counterspeech mentions and true uses. Intuitively, the mitigation should reduce the misclassification of counterspeech mentions (false positive rate on counterspeech) while not reducing correct positive classification of true uses (true positive rate on uses).
Results
We find that the CoT mitigation reduces false positive rate among counterspeech mentions by 82.61% for hate speech and 59.06% for misinformation (Table 8). Among true use statements, CoT mitigation reduces true positive rates only marginally (2.99% for hate speech and 2.62% for misinformation). Ablated mitigations reduce false positive rate among counterspeech mentions independently, but CoT + mitigation is the best performing condition. With gpt-3.5-turbo, similar patterns were observed (see Appendix B).
In summary, encoding the use-mention distinction reduces downstream misclassification with otherwise minimal reductions in performance on true use of hate speech and misinformation.
Discussion
Existing datasets of harmful content are typically collected by sampling keywords that co-occur with harmful content and performing annotation. For example, during hate speech annotation, a typical question might ask: “Does the above text contain rude, hateful, aggressive, disrespectful, or unreasonable language?” Rae et al. (2021). Previous work has documented that existing datasets gathered through such a process contain mentioned language misclassified as harmful Van Aken et al. (2018). This implies that researchers may not typically consider the special case of mentioning statements (including counterspeech) when collecting annotations, or that annotators with varying backgrounds and positionality may not agree on how to treat mentioning statements Santy et al. (2023).
By classifying hate speech and misinformation, NLP tools make implicit judgments regarding when mentioned language is permissible. However, counterspeech should be permissible by definition, as it challenges and opposes harmful narratives Mun et al. (2023); Hangartner et al. (2021). We suggest that in downstream content classification, treatment of mentioned language should instead be explicitly encoded. More broadly, efforts to use LLMs for the design of conversational socio-technical systems should take into account the use-mention distinction and strive to embed values regarding how content should be treated. To that end, our prompting mitigation serves as an example of how one societal value can be encoded by leveraging a linguistic construct to specify the treatment of counterspeech.
Beyond surface features
A substantial body of literature has examined shortcomings of online content classification Garg et al. (2023); Van Aken et al. (2018). Our results suggest that many such error cases are linked to the inability to sufficiently distinguish between use and mention.
For instance, previous work has suggested that the mere fact that a text contains swear words Ethayarajh et al. (2022), identity terms Dixon et al. (2018) related to ethnicity Ghosh et al. (2021), religion Sheth et al. (2022), gender, race, and disability Dias Oliva et al. (2021); Díaz and Hecht-Felella (2021), or dialect markers Sap et al. (2019); Halevy et al. (2021) increases toxic classifications Zhou et al. (2021). Such lexical biases where surface features spuriously influence classification are indicative of inadequate ability to make a use-mention distinction.
Surprisingly, we find that even the latest models (i.e., gpt-4) are very sensitive to shallow lexical biases. Although non-literal language understanding beyond surface features is essential to human communication, large language models tend to struggle with interpretations of subtle utterances Hu et al. (2023); Ocampo et al. (2023); Yu et al. (2022). Our work thus extends prior work aiming to perform more nuanced online content classification Pamungkas et al. (2020); Goyal et al. (2022); Sap et al. (2020). Future work needs to attend to these subtle and implicit meanings to keep misinformation and hate speech classifiers from being mere surface topic classifiers.
Beyond counterspeech
While we focus on online speech, the use-mention distinction might be important in other contexts like education or tutoring (where mention language occurs frequently to specify spelling or translation), law Henderson et al. (2022), and human-AI interaction Shaikh et al. (2023). In general, our work offers promising directions for embedding meta-linguistic reasoning to improve performance on challenging downstream tasks, such as eliciting metaphorical meanings, or dog whistle classifications, where existing methods are not reliable Wachowiak and Gromann (2023); Mendelsohn et al. (2023) and teaching meta-linguistic skills might offer a possible way forward.
Conclusion
This work highlights the theoretical and practical importance of the use-mention distinction in NLP and CSS. We also provide guidelines and directions for future research in modeling mentioned language and mitigating the impact on downstream tasks where failure to distinguish mention leads to harmful misclassifications.
Ethical implications
The central implication of our work is that downstream tasks in applications that involve occurrences of mentioned language should be handled with caution, especially when misclassification of mention could be harmful. However, content categorized as counterspeech might not be universally beneficial. For example, humans can demonstrate bias in use of counterspeech to preferentially challenge content from those with whom they disagree politically Allen et al. (2022). In that way, counterspeech can be leveraged as a tool to harass others with opposing views. Consequently, our mitigation to prevent censorship of counterspeech might allow harmful content to proliferate due to taking a form of counterspeech. We note that incorporating specific guidelines into platforms requires further testing.
We also note that strategies involving mentioned language can equally be used to veil the speaker’s true intention. Using specific terms might be socially acceptable for a person of a given identity, while unacceptable for someone else. For instance, acceptability of mentions of certain slurs is debated, especially if a non-derogatory version is readily available Green (2023). Addressing such complexities in the mentioned language largely remains an open question.
Lastly, despite known limitations, natural language systems are widely used for online content moderation Welbl et al. (2021); Gehman et al. (2020). It remains unclear how well they are expected to perform and whether it is appropriate to deploy models that might produce harmful errors Fortuna et al. (2022). Our work echoes the need for more cautious development and deployment of such systems that accounts for the conversational context, identities, and intentions. Our work also builds on existing literature that challenges the very concepts of binary classification Pachinger et al. (2023); Davani et al. (2022) into constructs such as hate speech.
Limitations
We limit the scope of our work to testing publicly available out-of the box classifiers. Similarly, we do not investigate all the possible mitigation strategies. For example, fine-tuning with more examples could help further decrease the error rates. We also note that we do not take into account how humans would rate studied texts. Previous work studying online discourse has found that people too struggle in disentangling mentioning a fact from stating an opinion, which makes the subsequent conversation more likely to derail into uncivil behavior Chang et al. (2020). Nonetheless, counterspeech examples examined in this work were written by social media users and expert humans who deemed them appropriate for responding to harmful narratives.
Finally, the present study is limited to specific types of mentioned language related to counterspeech. We do not solve all the problems of use-mention in the broader linguistics literature Sandhan et al. (2023); Wilson (2011b), including attributed language, words or phrases as themselves, proper names, and translations and transliterations. In this work, we are primarily interested in attributed language and words or phrases as themselves, as they closely relate to counterspeech which needs to mention phrases stated by others. Future work should consider other forms of mentioned language, and their impact on NLP systems more broadly. For instance, grammatical error correction might be impacted by failures in use-mention distinction Arand (2022). Our work is a first step toward analyzing the language of mention in different tasks, contexts, languages, cultures, and times.
Acknowledgments
Kristina Gligorić is supported by Swiss National Science Foundation (Grant P500PT-211127). Myra Cheng is supported by an NSF Graduate Research Fellowship (Grant DGE2146755) and Stanford Knight-Hennessy Scholars graduate fellowship. This work is also funded by the Hoffman–Yee Research Grants Program and the Stanford Institute for Human-Centered Artificial Intelligence.
References
Appendix A Dataset statistics
Two annotators annotated a random sample of statements (160), uniformly distributed across hate speech and misinformation. Half of the instances were uses, and half counterspeech mentions. Each statement was annotated to indicate if the focal tokens are used or mentioned (following Def. 1). Given that we leveraged datasets with curated counterspeech statements Fanton et al. (2021); Chung et al. (2021); He et al. (2023), as expected, within the selected sample, true use original posts contained no mentions, while counterspeech contained no uses of hate speech and misinformation. Similarly, counterspeech mentions the same focal tokens from the original statement (as opposed to mentioning harmful language not in the original use statement).
Appendix B Mitigations
In Figure 2, we illustrate true positive rate on true use and false positive rate for counterspeech statements across the tested models. Mitigation reduces false positive rate on counterspeech, with a small decrease in true positive rate for true use statements. In Table 9, for completeness, we list the results from mitigation study with gpt 3.5-turbo (ChatGPT 3.5).
Appendix C Additional statistics
Table 10 lists propagation statistics separately by task. Table 11 lists recall in hate speech classification stratified by target identity. We note that groups that have higher counterspeech false positive rates have a higher recall on true use, while groups with lower counterspeech false positive rates have lower recall too. Some discrepancies exist, as, for instance, false positive rate in mentioning counterspeech for Jewish is higher than for disabled identities, despite similar recall on true use (96% and 98.11%). These patterns are likely associated with biases in training data which implicitly encode the treatment of different groups, and the respective counterspeech.
Appendix D Prompt text
For reproducibility, in Tables 12 and 13, we list the complete prompts tested in our studies.