Chain of Explanation: New Prompting Method to Generate Higher Quality Natural Language Explanation for Implicit Hate Speech

Fan Huang, Haewoon Kwak, Jisun An

Introduction

Warning: This paper contains offensive content and may be upsetting.

Tremendous hateful speeches are created and spread every second online, which can lead to various social problems (Hine et al., 2017). Natural language processing has shown to be a powerful tool to accurately and efficiently detect hateful speech on online social platforms (Salminen et al., 2018).

When a text explicitly contains hateful speech and obvious discrimination words, feature attribution approaches, such as LIME (Ribeiro et al., 2016) or SHAP (Lundberg and Lee, 2017), can provide reliable information on why the text should be classified as harmful by highlighting specific words.

To obtain plausible and faithful explanations for implicit hate speech, various natural language explanation (NLE) generation methods have been proposed. However, most previous studies have focused on autoregressive generative language models (GLMs) to generate NLE for hate speech without prompting methods (ElSherief et al., 2021). The potential of sequence-to-sequence (Seq2Seq) models and prompting methods has not been fully explored (Ding et al., 2021). Moreover, traditional evaluation metrics, such as BLEU (Papineni et al., 2002) and Rouge (Lin, 2004), applied in NLE generation for hate speech, may also not be able to comprehensively capture the quality of the generated explanations because they heavily rely on the word-level overlaps (Clinciu et al., 2021). To fill those gaps, we propose a Chain of Explanations (CoE) prompt method to generate high-quality NLE distinguishing the implicit hate speech from non-hateful tweets. We then benchmark the mainstream Natural Language Generation models through comprehensive auto-evaluation metrics and human evaluation.

Related Work

Hate Speech Detection and Explanation. Due to its societal importance, researchers have actively studied and proposed various models to detect hate speech (ElSherief et al., 2021; Huang et al., 2023). Providing explanations to the AI system users would help improve user experience and system efficacy (Epstein et al., 2022). Furthermore, providing implied meanings of the text before it is posted has proven to help prevent the potentially harmful posts (ElSherief et al., 2021). For those explicit hate speech, highlighting words or phrases, which can be done via feature attribution-based explainable techniques like LIME (Ribeiro et al., 2016) or SHAP (Lundberg and Lee, 2017), post-hoc explanations techniques that rely on input perturbations, could be an effective way to explain. Thus, more recent works were proposed to apply the Generative Pre-trained models to create hate explanations (ElSherief et al., 2021).

Prompt Learning. In recent years, prompt learning has become a widely-accepted paradigm for pre-trained language models (Ding et al., 2021). The prompt learning could mine knowledge from the pre-trained models in multiple manners through manually selected prompt designs (Gao et al., 2021). With the help of prompt learning, the text generation task could be fostered through a prompt with task-specific information. Still, challenges for prompt learning are: (1) the well-performed prompts are highly task-specific, and (2) the existing well-performed prompt may not be suitable for all data instances when facing different data instances (Gao et al., 2021). While promising, it is under-explored whether prompt learning can generate quality explanations for implicit hate speech.

Chain of Explanation

Following the text generation task formulation used in (ElSherief et al., 2021), we construct our baseline task that only generates NLE for a given text. During the training process, the generation model uses the following token sequence as the input:

where [STR] is the start token, [SEP] is the separate token, and [END] is the end token. Inside, t1,...,tnt_{1},...,t_{n} represents an input tweet, while tE1,...,tEmt_{E1},...,t_{Em} represents the NLE.

2. Chain of Explanation Prompt

Inspired by (Wei et al., 2022) that proves a chain of thought prompting is effective for various tasks involving a complex reasoning process, we propose the Chain of Explanation prompting method for implicit hate speech explanation generation. Our prompt is based on the following guidelines: (1) heuristic words to inform the expected information in the prompting structure, (2) demonstration of the hateful intention of the given text, and (3) demonstration of the target group of the hateful intention.

For the prompt design, the input token sequence is as follows:

where HtextH_{text} stands for the heuristic words for a given text, which is the tokenized sentence of “Given Text: ”; DhateD_{hate} stands for the demonstration information for hateful intention, which is the tokenized sentence of “Is the text hateful? Yes”; DtargetD_{target} stands for the demonstration information for target group of hateful intention, which is the tokenized sentence of “The target group is: {target}”; and HNLEH_{NLE} stands for the heuristic words for NLE, which is the tokenized sentence of “It is hateful because: ”.

NLE Generation

To train and test the NLE generation for hate speech, we use the LatentHatred dataset (ElSherief et al., 2021). The LatentHatred dataset includes 6,358 implicit hateful tweets, and each tweet is annotated with its hate category, target group (a particular group of people (e.g., Asian or women) targeted by hate speech), and implied statement (the implications of the implicit hateful intention).

2. Models

The GPT and GPT-2 models are widely used in the generation tasks for complex reasoning. As the accessibility to the GPT-3 model is limited compared to its preceding models, we use GPT-NEO (Gao et al., 2020) and OPT (Zhang et al., 2022) model instead. The Seq2Seq models, generating the sequence conditioned on the input sequence, could also generate the NLE based on the provided information. The most widely used Seq2Seq models are T5 (Raffel et al., 2020) and BART (Lewis et al., 2020). As for the generation settings, we choose greedy decoding for all the above models. For our experiment, we use the basic version of those models: “gpt2”, “gpt-neo-125m”, “opt-125m”, “bart-base”, and “t5-base”.

3. NLE Evaluation Metrics

The BLEU (Papineni et al., 2002) and ROUGE (Lin, 2004) metrics have been commonly used to measure the quality of generated texts by using the word overlaps. The Meteor (Banerjee and Lavie, 2005) metric calculates the score based on the harmonic mean of unigram recall and precision values considering synonym and stemming. The NIST metric evaluates the informativeness of the n-grams (Doddington, 2002). The SARI metric measures the goodness of words by comparing generated NLE with the ground truth NLE (Xu et al., 2016). With the help of Transformers, semantic-based metrics, such as BERTScore (Zhang* et al., 2020) and BLEURT (Sellam et al., 2020), make use of the word embedding similarity. The NUBIA metric, known as the most advanced metric, reflects how interchangeable the sentences are using various neural models (Kane et al., 2020). All metrics ranges from 0 to 100 except NIST (0 to 300) and BERTScore (-100 to 100). For all scores, the higher, the better.

Evaluation

Table 1 shows an example of generated explanations given a implicit hateful tweet along with ground truth explanation. We see that the results of our proposed models (GPT-2,CoE and BART,CoE) are comparable with the baseline model (GPT-2,base).

We evaluate our results in comparison with the reported baseline (ElSherief et al., 2021) and our replication of the baseline results. ElSherief et al. (2021) use the GPT-2 model to generate both the target group and implied statement at once without a prompt method. We follow a similar training setting and tuning process of hyper-parameters in (ElSherief et al., 2021). We fine-tune for e∈{1,2,3,4,5}e\in\{1,2,3,4,5\} with the batch size per device of 2 and learning rate of 5×10−55\times 10^{-5}, also with 100 steps of linear warm up. Our split portion is 75:12.5:12.5 for training, testing, and validation.

Table 2 shows the results across various automatic metrics explained in §\S4.3. Our replications show comparable results with the reported baselines. For the baseline method, we find the OPT model performs the best among the autoregressive models, and the BART model performs the best among Seq2Seq models. Generally, BART outperforms other models.

At the row of the chain of explanation (CoE) in the table, we can see the significant improvement across all automatic evaluation metrics. The automatic evaluation metric scores improve significantly from 44.2 to 61.8 for BLUE-1 and from 33.7 to 52.6 for Rouge-L. Similarly, the OPT remains the best model among autoregressive models, and the BART remains the best among Seq2Seq models. Between them, BART outperforms in most of the metrics.

Ablation Study We further study the importance of each part of our CoE prompt design by conducting an ablation study. Table 3 present the ablation test results for BART, the best performing Seq2Seq model, inspecting three variations of the chain of explanation prompting design. The ablation design is to remove the heuristic words and re-run the test to see if there would be a difference. We conduct a similar ablation study for OPT, the autoregressive model, and find consistent results.

(1) The heuristic words. We find that the BLEU-1 score drops by 0.6 when removing the heuristic works. One explanation for why the chain of explanation prompting method performs well is that the heuristic words added to the text contributed to guiding the PLMs to generate the explanation with the clear purpose of answering why the given tweets would be considered hateful.

(2) The hate label demonstrations. The BLEU-1 score drops by 1.4 for the BART model. Another possible explanation is that the hate label helps the model to confirm the hateful nature of the given text and thus gives a hint to the model that the given text should not be understood as the non-hateful text so that the model would not generate the meaningless NLE.

(3) The target group demonstrations. We observe the target group is the most important information in generating high quality hate explanation. Removing the target group results in dropping 15.5 of the BLEU-1 score.

Human Evaluation - Informativeness and Clarity Human perceptions have always been considered one of the most important evaluation criteria for Natural Language Generation tasks (Evans and Grefenstette, 2018; Gkatzia and Mahamood, 2015). Thus, we investigate the quality of the generated NLE from the perspective of human perception by using two metrics—Informativeness and Clarity. Informativeness (Dušek et al., 2020) considers how relevant the information in NLE explains why the tweet should be perceived as hateful by human readers (e.g., 1 = Not Informative and 7 = Very Informative). Clarity (Belz and Kow, 2009) measures how clear the meaning of the NLE is expressed (e.g., 1 = Unclear and 7 = Very Clear).

For human annotation, we first tried Amazon Mechanical Turk (AMT). However, we found that the raters do not agree with each other—the inter-rater reliability (Krippendorff’s Alpha value) was 0.14. Thus, we hired three experienced Research Assistants to annotate our data. For 100 randomly selected NLE samples, we collect at least three annotations for each NLE. We then remove the instances that are too hard to reach a consensus based on the same rule in (Clinciu et al., 2021). This results in annotations for 68 NLE samples. The inter-rater reliability score is 0.35, indicating that the raters are in relatively fair agreement. For the ground truth NLE, the average Informativeness is 5.20 (95% CI: 4.95—5.45), and the average Clarity is 4.53 (95% CI: 4.30—4.77). Our best generated NLE shows slightly lower performance on both metrics—its average Informativeness is 4.48 (95% CI: 3.95—5.00), and the average Clarity is 4.34 (95% CI: 3.95—4.72). Still, these values are comparable with the results of existing work (Clinciu et al., 2021). In interpreting Bayesian Network graphical representations, which could be a less subjective task than hate explanation, the human written explanations achieved, on average, 4.66 (95% CI: 4.44—4.88) and 4.65 (95% CI: 4.45—4.86) for Informativeness and Clarity, respectively.

We further analyze how those two human-annotated metrics are correlated with automatic evaluation metrics to show the gap between them. For comparison, we use the median scores of the human annotations, as proposed in (Clinciu et al., 2021), and the scores of the automatic metrics by our best model (i.e., BART). Table 4 presents the Spearman correlation coefficients between each of all automatic metrics and the two human annotated scores.

First, the BLEURT, BERTScore, and NUBIA metrics correlate better with the two human evaluation metrics. The BLEURT metric shows the highest correlations. Second, the SARI metric shows the lowest correlations with both Informativeness and Clarity. Lastly, for the widely used BLEU-1 and ROUGE-L metrics, the correlation with Informativeness is significantly higher than Clarity, which is not aligned with (Clinciu et al., 2021). Understanding the origin of the differences would be a good future research direction.

Discussion and Conclusion

In our work, we proposed the Chain of Explanation prompt design to generate high-quality Natural Language Explanations for implicit online hate speech. The performance is evaluated comprehensively through various auto-evaluation metrics and human evaluation.

As the prompting method aims to generate Natural Language Explanations to illustrate why the given text is toxic or hateful, model’s output thus could contain hateful expressions, which may produce additional stress to the end users. The PLMs would also learn the implicit hateful or conspiracy-based logic and expressions (Levy et al., 2021), making the generation results less accountable. The potential solution is to apply a shepherding system to filter out hateful statements or transfer them to non-hateful expressions.

Ethical Considerations

The Institutional Review Board at Singapore Management University has approved this study (Approval No.: IRB-22-076-A043(622)). For the annotation process, we included a warning in the instructions informing that the content might be offensive or upsetting. Annotators were also encouraged to stop the labeling process at any time if they felt overwhelmed or unwell.

References