Human-Machine Collaboration Approaches to Build a Dialogue Dataset for Hate Speech Countering

Helena Bonaldi, Sara Dellantonio, Serra Sinem Tekiroglu, Marco Guerini

Introduction

While hate towards vulnerable groups or individuals is not a new phenomenon, the upsurge of hate speech and its proliferation is relatively recent and it is enabled by the fast spread of information in online platforms. The rise in hate speech online can even provoke violent actions offline. Consequently, fighting online Hate Speech (HS) has become a vitally important “job for everyone”“Hatred is a danger to everyone – and so fighting it must be a job for everyone.” António Guterres, United Nations Secretary-General, 2021 especially for the NLP researchers. The contrast to HS and haters on social media platforms is usually carried on via user suspension, content removal or shadow banning, which can be mapped to a classification task in NLP terms. However, AI and NLP can play even a more crucial role that is not limited to classification. In fact, recently, NLG models have started to be proposed as an effective tool to counter HS by providing relevant responses. In particular, the idea is to imitate the operators of Non-Governmental Organizations (NGO) that are actually intervening in online discussions by replying to hateful content using so-called Counter Narratives (CN), defined by Schieb and Preuss (2016) as “communicative actions aimed at refuting hate speech through thoughtful and cogent reasons, and true and fact-bound arguments”. Through automatically generating CNs, it is possible to aid NGO operators in their day-to-day manual activities, and therefore to partially countervail the sheer amount of hateful content posted online Chung et al. (2021b).

Despite the invaluable attempts to create HS/CN datasets and systems (Mathew et al., 2019; Qian et al., 2019; Chung et al., 2019; Fanton et al., 2021), up to now only datasets containing 2-turn interactions have been proposed (i. e. a hate speech and a responding counter narrative), while in real scenarios, such as on social media platforms, multi-turn dialogues are the norm. In Figure 1 an example of such dialogues is providedThis paper includes examples of hateful content, which may be upsetting for the readers. However, they do not represent the views of the authors.. Therefore, multi-turn dialogue datasets are necessary for training models that can better handle online hate phenomenon.

Still, obtaining expert-written quality data to train such models on is not trivial. To ameliorate this problem, a recently proposed approach is the use of hybrid data collection strategies where a human and a machine collaborate to build data starting from a seed dataset of expert based examples (Fanton et al., 2021). In this paper we follow this line of research and investigate novel strategies and algorithms that are specifically designed for multi-turn dialogues collection.

In particular, we test 19 different hybrid strategies obtaining a novel dataset of more than 3K dialogical interactions between two interlocutors, one acting as the hater and the other as the NGO operator, for a total of more than 16K turns. We call this dataset DIALOCONAN (DIALOgical COunter-NArratives collectioN). This is the first and most comprehensive multi-target dataset that addresses expert-based counter narrative generation in fully dialogical scenarios, and it can be downloaded at the following link: https://github.com/marcoguerini/CONAN.

Related Work

In this work, we consider four main research areas as relevant: in particular (i) available datasets for hate speech detection, (ii) available datasets for CN generation, (iii) CN generation approaches, and (iv) hybrid data collection methodologies.

Many benchmarks for automatic HS detection are currently available (Mathew et al., 2021; Cao et al., 2020; Kumar et al., 2018; Hosseinmardi et al., 2015; Waseem, 2016; Burnap and Williams, 2016). Regarding the systems built on top of these benchmarks, we refer the readers to the surveys by Poletto et al. (2020); Schmidt and Wiegand (2017); Fortuna and Nunes (2018) for detailed reviews. Other reviews include the analysis of ethical implications (Kiritchenko et al., 2021) and of problems such as bias replication (Binns et al., 2017; Davidson et al., 2019; Vidgen and Derczynski, 2020; Sap et al., 2019; Tsvetkov, 2020).

CN data collection.

Since CNs have been shown to be effective in reducing linguistic violence (Benesch, 2014; Gagliardone et al., 2015; Schieb and Preuss, 2016; Silverman et al., 2016; Mathew et al., 2019) and in changing the viewpoints of bystanders (Allison and Bussey, 2016; Anderson et al., 2014), they are beginning to be collected as training data for supervised NLG models. The investigated approaches for data collection can be listed as crawling (Mathew et al., 2018, 2019; Yu et al., 2022), crowdsourcing (Qian et al., 2019), nichesourcing (Chung et al., 2019) and hybrid approaches (Tekiroğlu et al., 2020; Fanton et al., 2021).

The most relevant datasets for our work are (i) Fanton et al. (2021) in terms of quality and target diversity, even if it only includes HS/CN pairs, and (ii) Qian et al. (2019) that hints at the issue of multi-turn dialogues. However, in the latter the CN is only the last turn of a forum-style dialogue among more than 2 interlocutors, rather than a HS/CN multi-turn dialogue between two opposing actors.

CN generation.

Neural approaches to generate CNs have started to be studied along with available datasets (Fanton et al., 2021; Tekiroğlu et al., 2020; Qian et al., 2019). Tekiroglu et al. (2022) present a thorough comparison of several pre-trained LMs for this task. Zhu and Bhat (2021) propose an entirely automated 2 stage pipeline where several CN candidates are generated and then filtered. Other lines of work include CN generation for under-resourced languages (Chung et al., 2020), or the generation of knowledge-bound CNs, to avoid hallucination phenomena (Chung et al., 2021a). Finally, Ashida and Komachi (2022) studied CN generation with LLMs, using few-shots prompting.

Hybrid models for data collection.

A recently emerged data collection methodology is based on hybrid models, where humans and machines work together to collect better quality data in a more efficient way. Wallace et al. (2019) propose using model output to guide humans in the writing of adversarial examples for question-answering systems. Dinan et al. (2019) and Vidgen et al. (2020) perform a data collection for offensive language detection with repeated model-human interactions where the classifier output drives annotators in example creation at each round. A more recent study proposes a hybrid approach where an LM is trained to generate HS/CN pairs that are validated and post-edited by annotators (Tekiroğlu et al., 2020). Fanton et al. (2021) further expand this approach by making it iterative using several LM configurations.

Methodology

Data collection can be very difficult and time consuming when high quality data from experts are necessary. Given that we need to collect whole HS/CN dialogues and not just pairs, the problem is even harder. Moreover, scraping NGO operators’ real interactions is not a viable solution, considering that this data can be used for account “doxing”. In fact, malicious users could reverse-search the text included in a dataset to identify the operators’ accounts. This would undermine their work, since they usually operate undercover, and would expose them to possible attacks.

Therefore, we decided to resort to hybrid approaches and run 3 different data collection sessions based on the aspects of the dialogue augmentation we want to address (either the structure, in terms of turns order, or the wording of the turns).

In total we tested 19 different dialogue collection strategies. All the strategies are inserted in an author-reviewer pipeline as described by Tekiroğlu et al. (2020), where the author is a single dialogue creation strategy at a time, and the reviewer is represented by a team of trained annotators, who are tasked with post-editing the dialogues generated by the given author strategy.

Each of the 3 data collection sessions we perform has different input data and author tasks, in particular:

Session 1: same wording, new dialogue structure. 7 strategies based on concatenating pre-existing material (HS/CN pairs) to obtain new dialogues.

Session 2: new wording, same dialogue structure. 6 strategies to modify the wording of pre-existing dialogues via paraphrasing.

Session 3: new wording, new dialogue structure. 6 strategies using generative Language Models (LMs) for complete dialogue generation.

Author - Seed datasets.

Since each author configuration needs some textual input, we employ (i) a dataset, created ad hoc, consisting of 222 fictitious dialogues and (ii) HS/CN pairs coming from the dataset presented in Fanton et al. (2021).

The ad hoc fictitious dialogues (DIALOgold henceforth) are written by two expert NGO operators, who have been working for over 10 years in writing CNs on social media platforms. They were asked to write dialogues between a hypothetical hater and an NGO operator, following their real expertise in the task. The dialogues can have 4, 6, or 8 turns (these are typical lengths according to their experience) and cover the following 6 targets of hate, defined beforehand: LGBT+, MIGRANTS, MUSLIMS, JEWS, POC and WOMEN.

Given the small size of DIALOgold, we also use part of the dataset presented in Fanton et al. (2021) as an additional resource. This dataset consists of 5000 HS/CN pairs covering, among others, the 6 targets of hate present in DIALOgold. Therefore, we extracted the pairs labeled with these 6 targets so that the two resources can be ‘aligned’ by topic, and we named it PAIRSgold, since also this dataset was created with the help of expert NGO operators.

Reviewers - Training.

For post-editing the output of the various author configurations, three annotators were recruited from a pool of internship students. They have been extensively trained using the methodology of Fanton et al. (2021), in order to become “experts” on HS/CN post-editing. In particular, we first explained the aim of the task. Then, they had to read NGO guidelines and documentation on CN writingSee https://getthetrollsout.org/stoppinghate as a reference., together with all the dialogues present in DIALOgold, which were provided as examples of the material they would have to work with. We detailed the methodology, explaining that the main focus was to make the dialogues natural, with the minimum intervention possible and keeping the seed dataset as a reference for naturalness. General instructions about the post-editing procedure were also provided, pointing out that for each session specific guidelines would have been given.

Reviewers - Mitigation procedure.

Finally, we also implemented a mitigation procedure similar to the one presented by Vidgen et al. (2019). This procedure is implemented to safeguard the annotators’ well-being while working with abusive content and it includes: (i) explaining to the annotators the pro-social nature of the research and the purpose of their post-editing activity, (ii) advising the annotators to work few hours per day and to take regular breaks (iii) having weekly meetings to let possible problems or distress emerge.

Data collection procedure.

For each session we applied the following procedure: (i) generate dialogue candidates according to session specific strategies, (ii) adapt the annotation guidelines to the specific session, (iii) let the annotators practice the task on a small “training” set of dialogue candidates, and (iv) update the guidelines with respect to their feedback. Lastly, (iv) annotators complete the post-editing on the remaining dialogues following the updated guidelines (the order of the dialogues was randomized to avoid comparison or primacy/recency effects over session strategies).

Metrics

We use several metrics to assess the performance of each strategy. These metrics are aimed to assess either the efficiency of the procedure or the quality of the obtained data.

HTER is an efficiency metric used to measure the post-editing effort of the annotator, and it is usually employed for sentence level translations Specia and Farzindar (2010). A value above 0.4 is generally used to account for low quality outputs, where rewriting from scratch is on par with correcting it Turchi et al. (2013).

Turn deletion is the percentage of turns that are discarded by the reviewers since their quality is too low and/or they do not fit in the current dialogue structure. The more content needs to be deleted, the less efficient the procedure is.

Turns swap is the percentage of turns that are moved by the reviewers from the original position they were in, to another position in the final edited dialogue. Usually turns of this kind have a good quality but they do not fit the current position.

Novelty is utilized to check the quality of a generated dialogue by measuring its lexical difference with respect to a reference set of dialogues, and it is grounded on Jaccard similarity (Dziri et al., 2019; Wang and Wan, 2018).

Repetition Rate (RR) measures the language diversity within a corpus using the rate of non-singleton ngram types (Cettolo et al., 2014; Bertoldi et al., 2013). It is used in our experiments to evaluate each strategy in terms of its ability to provide diverse and varied examples.

Session 1: Dialogue structure

In Session 1 we started from the HS/CN pairs in PAIRSgold and concatenated them in order to produce dialogue candidates with different structures.

We employ 7 strategies to connect HS/CN pairs from PAIRSgold to create 4, 6, and 8 turns examples (consistently with the DIALOgold characteristics).

During the concatenation, each pair is used only once in a dialogue. The connection strategies are: random concatenation (1 strategy), similarity concatenation (4 strategies), and keyword matching concatenation (2 strategies). In order to obtain a balanced dataset, for each strategy, each target, and each dialogue length combination we created 10 connected dialogues. In-detail descriptions of the 7 concatenation strategies we utilized are as follows:

For the random connection (RND), the selected pairs for each target are randomly concatenated to form dialogues. This strategy represents a baseline to which we compare against while analysing the other strategies.

Similarity connection.

To connect pairs depending on to their similarity, we utilize (i) the Jaccard similarity and (ii) the cosine similarityCosine Similarity is computed on their embeddings obtained with mpnet-base. The Sentence Transformer library (https://www.sbert.net/) has been employed.. Both for the Jaccard and cosine similarity, we perform pair matching via two approaches to form the HSi,CNi,HSi+1,CNi+1HS_{i},CN_{i},HS_{i+1},CN_{i+1} concatenation:

SIMHS-HS{}_{HS\text{-}HS} = the similarity between HSiHS_{i} and HSi+1HS_{i+1};

SIMCN-HS{}_{CN\text{-}HS} = the similarity between CNiCN_{i} and HSi+1HS_{i+1};

For each pair, we randomly select 1 among the 10 most similar pairs according to the chosen similarity (either Jaccard or cosine) and concatenation elements (either HS-HS or CN-HS). The procedure is repeated until the desired number of turns for each dialogue is reached.

Keywords connection.

We employ the YAKE keyword extractor (Campos et al., 2020) to extract two keywords from each HS and CN of PAIRSgold and perform a concatenation similar to the previous strategies. We connect HSi,CNiHS_{i},CN_{i} and HSi+1,CNi+1HS_{i+1},CN_{i+1} according to the following criteria:

KWHS-HS{}_{HS\text{-}HS} = if HSiHS_{i} and HSi+1HS_{i+1} share two keywords;

KWCN-HS{}_{CN\text{-}HS} = if CNiCN_{i} and HSi+1HS_{i+1} share two keywords;

We decided on a 2-keywords match since according to our preliminary manual analysis we found that the first keyword is often target-related; by considering two keywords we aim to include also a topic-related keyword.

As a final note, we should highlight that the two groups of connection strategies (HS-HS and CN-HS) represent either (i) a global semantic coherence across turns (all HS being similar) or (ii) a local semantic coherence (only CN-HS of adjacent turns being similar) both for SIM and KW. By using a global semantic coherence via HS-HS matching we attempted to simulate the attitude of the hater which is convinced of their own ideas and do not accept any external input, while with the local connections, we aimed to recreate a “linguistic alignment” phenomenon Doyle and Frank (2016). Details on the matching procedures and the description of the algorithms for SIM and KW we employed are reported in Appendix A.1.

2 Reviewing phase and guidelines

In order to obtain natural dialogues, the annotators in this session received specific post-editing instructions:

Since CNs are gold, it is strongly suggested to post-edit only the HSi+1 to “align” it with the CNi belonging to the previous turn.

If a pair is in an unnatural position of the dialogue it should be moved to a better position.

If a pair is not fitting with the flow of the dialogue and cannot be moved elsewhere, it should be deleted.

If the whole dialogue makes no sense, or is too difficult to fix, it should be deleted.

A characteristic example of the post-editing done in Session 1 is shown in Table 10 in Appendix C.

3 Results

Results of this session in terms of efficiency and quality are reported in Table 1. In general, we observe that strategies using any HS-HS connection are less efficient, having higher HTER scores as compared to the CN-HS ones. HS-HS connections also have a high rate of deleted turns, in particular KWHS-HS{}_{HS\text{-}HS} and J-SIMHS-HS{}_{HS\text{-}HS}. The KWHS-HS{}_{HS\text{-}HS} strategy is even more inefficient than the random connection baseline (it reaches the highest number of deleted turns and the highest HTER), and it is the most repetitive before post-editing, as showed by the RRgen. These results are also confirmed by the annotators’ feedback, who noted the presence of dialogues which were particularly difficult to edit since they contained the same HS repeated multiple times (see example in Table 11, Appendix C). A posteriori analysis showed that these dialogues were mainly obtained through the KWHS-HS{}_{HS\text{-}HS} connection. Moreover, each HS-HS connection strategy achieves a higher RRgen score than its CN-HS counterpart, showing that connecting through a global similarity generates a higher overall repetitiveness than using a local similarity. The particular high scores reached by the RRgen of both the keywords connection strategies can be explained by the procedure employed for connection: for keywords we performed an exact matching, whereas with the cosine or Jaccard similarity, the connection was selected from the 10 most similar candidates.

After post-editing, all the strategies achieve a lower RR, between 4.5 and 9, indicating a more diversified content. The novelty is calculated against DIALOgold: the scores are similar for all the strategies, and they are hardly affected by the post-editing, showing that each strategy managed to add a consistent novelty to the already present gold data. Finally, it is worth noting that the strategies employing HS-HS connections have less turn swaps than C-SIMCN-HS{}_{CN\text{-}HS} and J-SIMCN-HS{}_{CN\text{-}HS}. The most probable explanation is that CN-HS strategies require less deletion, but this comes at the cost of more turn swaps.

Session 2: Dialogue Wording

In the second session we focused on strategies aiming to obtain a new wording, given a structured dialogue. In particular, we tested 6 paraphrasing approaches on DIALOgold{gold} and on a part of the dialogues resulting from the first session. In this session, our overall aim is to obtain novel and diverse responses to hate. Therefore, we chose to paraphrase only the CNs belonging to a subset of the data collected in Session 1, while keeping the corresponding HS as it is.

We carried out two exploratory studies to test different paraphrasing configurations and we selected the 6 most promising ones, as described in Appendix A.2. We use both paraphrasers with no specific style and with style transfer in order to attain a diverse data collection. The selection has been performed by assessing the aspects of dialogue wording that we deem the most relevant for our scenario.

We use 2 paraphrasing tools as a ‘baseline’ where we do not impose any specific style to the paraphrases: the Protaugment paraphraser (Dopierre et al., 2021) and the Style paraphraser (Krishna et al., 2020) with basic style.

Style paraphrasing.

This group includes 4 strategies in which we aimed to generate paraphrases with specific styles, in order to enhance the diversity of our data collection. Specifically, we focused on a style similar to that present in social media or in dialogues (Style paraphraser (Krishna et al., 2020) with Twitter and Switchboard style), and formal or casual (Style former paraphraser https://github.com/PrithivirajDamodaran/Styleformer with casual and formal style).

For each CN, 3 different paraphrases are generated using the same paraphrasing strategy.

2 Reviewing phase and guidelines

In order to obtain more natural examples, the post-editing instructions given to the annotators are adapted accordingly, emphasizing the significance of novel wording.

The annotator should keep the gold HS as it is, while post-editing the most promising among the 3 CN paraphrasis suggestions, i. e. the one introducing the least errors and the most different one from the original.

Turn swap in this case is not allowed, since turns order was already validated in these dialogues and paraphrasing would not affect it.

For the same reason, turn and dialogue deletion are not allowed.

An example of a typical intervention of the annotators in Session 2 is shown in Table 12 in Appendix C.

3 Results

We report the results in terms of efficiency and quality in Table 2. All the paraphrasers employed reach similar HTER scores, which are below the 0.4 threshold, but higher than Session 1 results. Regarding the quality, generated paraphrases are highly novel with respect to the dialogues present in DIALOgold, but not as high if compared to the dialogues resulting from the connection of the gold pairs in Session 1. In addition, the annotators’ intervention enhances the novelty of the generated paraphrases in almost all the cases, and reduces the RR for all the paraphrasers, with lower scores than in the first session (3,739 vs. 5,814 on average).

To sum up, we conclude that it is better to concatenate PAIRSgold if we have a high number of pairs available, while paraphrasing is a viable solution if there is no pairs availability, since it implies a higher HTER and it is not justified by higher novelty.

Session 3: Generation

In this session we follow the overall configuration presented in Fanton et al. (2021), where the author is an LM fine-tuned on the DIALOgold together with the dialogues resulting from Session 1We did not include the dialogues resulting from Session 2 since they would have added little novelty. In particular, Session 2 CNs are paraphrases of those present in Session 1, and the HS of the dialogues in the two sessions are identical..

An autoregressive model specific for dialogue generation (Zhang et al., 2020). We choose DialoGPT since it is proven to be effective in CN generation as well (Tekiroglu et al., 2022);

T52m.

Two T5 (Raffel et al., 2020) models conversing with each other: one fine-tuned to produce only HS and one to produce CNs. This configuration allows to completely decouple CN production from HS production.

T51m.

One T5 model able to produce both HS and CN. We test it as a comparison to the two T5 models conversing with each other.

For each configuration, we test a baseline model, fine-tuned on DIALOgold only, and a model fine-tuned on both DIALOgold and the post-edited dialogues resulting from Session 1. For each model we employed the Top-pp decoding mechanism (Holtzman et al., 2020) with p=0.9p=0.9Training details are reported in Appendix B.. In all cases, we split the employed dataset into training, development, and test sets with a ratio of 8:1:1. For the generation phase, we use as a prompt the initial HS of the test set dialogues. Then, we generate a single turn at a time by feeding the model with the context generated so far, until we reach 8 turns dialogues.

2 Reviewing phase and guidelines

The data generated with the LMs include both HS and CN, therefore the annotators are allowed to post-edit both, unlike the previous session. The reviewing guidelines are similar to those for Session 1, with the following changes:

it is possible to swap single turns and not only pairs, since the connection between HS/CN is not granted a-priori as in the previous sessionsFor example, a model can introduce hateful content when it is supposed to generate a CN, or viceversa (as showed in Table 13, Appendix C)..

if some turns in a dialogue have a clearly different target than the labeled one, they should try to change turns wording to fit the original target.

the annotators should check the veracity of fact-based statements since they might derive from LM hallucinations.

3 Results

Results of this session, in terms of efficiency and quality, are reported in Table 3. There are two major conclusions we can drawWhile our main focus is on dataset creation, the results of this session offer also a form of simple benchmarking and some useful insights for the development of new models. In fact, the various metrics that we employed (post-editing, turn deletion, etc.) already provide a good indication of the LMs performance, especially for an open-ended scenario..

Firstly, adding the post-edited dialogues obtained concatenating PAIRSgold to the training data (DIALOgold) strongly increases the efficiency. In fact, these models require much less deletion from the annotators with respect to the baselines, reaching a lower HTER (<=0.4<=0.4). Also, even if the dialogues generated with the baselines have a higher novelty with respect to the training data, they are also extremely repetitive in almost all cases.

Secondly, as already shown by Tekiroglu et al. (2022), autoregressive models are producing more varied and relevant content as compared to seq2seq models. In fact, even if DialoGPT requires more post-editing than T5 configurations (with comparatively higher HTER scores), its output dialogues require a lower number of deletion. This indicates that that the DialoGPT generation is suboptimal but rarely unsuitable, while this often is not achieved by T5. In particular, turns swaps are present only for the DialoGPT models. According to the annotators, this is explained by the characteristics of T5 dialogues, which are more stereotypical, vague but have a better structure (see Table 14 in Appendix C). This is also confirmed by the quality results: T5 models generate content with similar novelty scores to DialoGPT, but they also tend to be more repetitive.

Session comparison & Data description

Finally, Table 4 compares the results for each session over the main metrics of interest. We observe that concatenating pre-existing material that is already verified (i.e. HS/CN from PAIRSgold) requires less effort than generating new data from scratch or paraphrasing gold material, as Session 1 reaches a lower HTER than both Session 2 and Session 3. On the other hand, in terms of the structure of the dialogues, Session 1 requires the highest effort as shown by the high swap rate. Meanwhile, Session 2 is the least repetitive, but also the least novel, providing dialogues with a good wording, even if this is not accompanied with a novel content. In general, all the sessions reach an HTER lower or equal to 0.4, and similar novelty scores with respect to the gold data. Therefore, in all cases it was possible to enhance the novelty of the initial seed dataset, with a reasonable post-editing effort.

Table 4 also shows a syntactic analysis of the data collected with each session, calculated at turn-level. The dialogues generated with the Language Models achieve the most balanced distribution in terms of number of turnsWe collected dialogues with 4, 6 and 8 turns, so a perfect balance would be of 6 turns., at the cost of simpler turns, as shown by the low maximum syntactic depth (MSD) and average syntactic depth (ASD) reached by Session 3. Paraphrasing instead provides the shortest generations both in terms of average turns length and of number of sentences (NST).

By comparing the results of the different sessions, we can conclude that the choice of the preferable data collection strategy firstly depends on the available input data, e. g. we might not always have gold HS/CN pairs available or multiple turns dialogues. Secondly, depending on the desired output, if the priority is to obtain novel content, Session 2 strategies would be the least favorite. Also, the concatenation of existing pairs as in Session 1 is a more cautious approach than the generation of completely new dialogues through LMs. Thus, Session 1 strategies can be preferred for a more conservative approach, whereas Session 3 strategies are better suited for a more creative data collection that comes at the cost of higher human correction effort.

As a last step, we performed a sanity check in which a senior NGO expert conducted a qualitative evaluation by reading a random sample of the post-edited dialogues from each session. Their feedback was positive, no critical issues were raised and all the dialogues were approved both in terms of produced CNs and of their overall structure/naturalness. Our final dataset, DIALOCONAN, includes also the dialogues collected through the various training phases and exploratory studies of the annotators. Table 5 shows the distribution of targets in terms of number and percentage of dialogues. The distribution is reasonably balanced, with the LGBT+ target being the most represented. Overall, we collected 3059 dialogues for a total of 16625 turns.

Conclusion

In this paper we have presented a hybrid approach for dialogue data collection in the realm of hate speech countering. These dialogues have been obtained starting from two expert-based seed datasets and then combining the intervention of human annotators over machine generated dialogues. We tested 19 different strategies for generation, focusing on two crucial aspects of dialogue, i.e. structure and wording. We analysed all these strategies in terms of efficiency of the procedure and quality of the data obtained. The result of this work is DIALOCONAN, the first dataset comprising over 3000 fictitious multi-turn dialogues between a hater and an NGO operator, covering 6 targets of hate.

Acknowledgements

We are deeply thankful to Stop Hate UK and its volunteers for the help in writing the dialogues for the DIALOgold seed dataset and for sharing their expertise, fundamental to this work.

Limitations

The datasets currently available for CN generation are mainly for the English language and this one is no exception. The problem is that getting in contact with NGO operators for other languages is not easily solvable. The alternatives, such as translating this dataset into other languages represent a suboptimal solution. In fact each language and country has (i) its own peculiar canards against minorities, (ii) even if HS can be ported across countries (e.g. “Migrants steal our jobs.”), the arguments to counter such HS may vary (e.g. different laws, different socioeconomic situations, different statistical data). When translating dialogues all these nuances get lost reducing the possible effectiveness and introducing possible unnatural answers making reference to the country for which the original CN was written.

Although we tried to keep the overall quality of the final output as high as possible, since the dataset is created through a human-machine collaboration paradigm, it can still not be on par with the data that we can potentially obtain with niche sourcing the whole dataset to skilled NGO operators. Additionally, the number of turns in dialogues are strictly controlled and might not reflect the more natural number of turns that would have occurred under those circumstances.

As previously stated, scraping NGO operators real online intervention is not desirable (we need to protect their identity). Still, even when collecting DIALOgold we encountered some problems, i. e. even if the annotators were trained and used to the task, they told us that the simulation was really frustrating (e.g. “repeating over and over again the same hateful content”).

Ethics Statement

Counter Narrative generation task and corresponding datasets have been proposed as a contribution of scientific research to a more ethical world. However, even the best intentions in the minefield of online hate can still bring along certain risks of undesired impacts on data curators (i.e., expert/non-expert annotators), on researchers, and on society. Therefore, in this study, we took meticulous precautions in order to avoid such effects.

As the most important stakeholders of this research, the annotators were constantly supported in terms of mental welfare. In particular, we put in practice a mitigation procedure similar to the one proposed by Vidgen et al. (2019), as described in Section 3.

Dataset.

Since the dataset is created from scratch via an expert-machine collaboration schema (rather than scraping the dialogues among individuals online) it does not pose any threat to personal privacy or individual rights. Additionally, we avoid to model inappropriate CNs (e. g. containing abusive language) that could be produced by scraping non-expert users in their online activity (Mathew et al., 2018).

Generation Task.

We consider the generation task as an aid to boost the data collection in terms of time, quantity, and certain quality aspects. Therefore, the models we trained are not meant to be deployed as part of a live system. Moreover, our main focus is clearly on the counter narrative generation part and the corresponding CN quality/diversity. For this reason, and to limit possible misuses, in our dialogues we tried to keep the HS as simple and stereotypical as possible and we always left ‘the last word’ to a CN turn. We encourage other researchers to conduct the generation tasks in a similar manner and for this reason the dialogue dataset will be made available for research purposes together with the code/models used to generate it.

References

Appendix A Appendix

The matching procedures over HS/CN pairs we employed are slightly different according to whether we performed a similarity (algorithm 1) or a keywords connection (algorithm 2). The main difference is that for similarity metrics it was always possible to choose among the 10 most similar pairs to the one of our interest, while when concatenating through keywords we put in practice an exact matching of pairs containing the same 2 keywords.

A.2 Session 2: exploratory studies

We select three settings for paraphrasing with no style transfer. We employ two paraphrasers: the Protaugment paraphraser and the Style transfer paraphraser. The tested configurations are the following:

Setting 1: Protaugment paraphraser with default parameters but drop_chance is set to 0.10.1 and lower_is_better=False.

Setting 2: Protaugment paraphraser with default parameters but lower_is_better=False.

Setting 3: Style transfer paraphraser with basic stye and p=0.6p=0.6.

In total, we select 36 dialogues to be paraphrased: 12 for each setting, with 4 dialogues for 4, 6, and 8-turns dialogues. We generate 3 candidate paraphrases for each CN while the HS is not paraphrased, since our interest is to enlarge the CN data, and not the HS data. One expert annotator is given instructions of reading all the dialogues and, for each CN, to select the most appropriate paraphrasis and modify it to make it fit in the dialogue. The chosen paraphrasis should be the one which requires at the same time the least editing to fit in the dialogue naturally and to be as much different as possible from the original CN.

high values for the HTER between CN and original paraphrasis (HTER CN-pselp_{sel}) and between the CN and post-edited paraphrasis (HTER CN-pedp_{ed});

low HTER between original and post-edited paraphrasis (HTER pselp_{sel}-pedp_{ed}).

From the results in Table 6 we can notice that the first setting (Protaugment paraphreser with default setting but lower_is_better = False and drop_chance = 0.1) is achieving the lowest values on the HTER between CN and pselp_{sel} and between CN and pedp_{ed}, while the second lowest with the HTER between pselp_{sel} and pedp_{ed}. The second setting (Protaugment paraphreser with default setting but lower_is_better = False) has the highest values on all the HTER results. The third setting (Style transfer paraphraser with default settings and basic style) has medium values on the HTER between CN and pselp_{sel} and between CN and pedp_{ed}, but the lowest value on the HTER between pselp_{sel} and pedp_{ed}, thus representing a good compromise for the characteristics of our interest. All the paraphrasers are making the original text shorter. From the results of the Δ\Delta length between CN and pedp_{ed} we can notice that the setting 3 is the one that after post-editing is making the paraphrasis closer to the original length, whereas this is more difficult to achieve with the other settings (same effort, paraphrasis closer to the original CN length).

As shown in Table 7, setting 1 is the one with less extreme results but for the HTER between CN-pselp_{sel} and between pselp_{sel}-pedp_{ed} the situation is the opposite than the one we aim for; setting 2 achieves the most extreme results. Despite setting 3 has a high percentage of examples reaching a high HTER between CN-pselp_{sel} and between CN-pedp_{ed}, still the results for HTER between pselp_{sel}-pedp_{ed} are not the worst.

For all these reasons, we decide to employ both the settings 1 and 3, while leaving out the setting 2.

Exploratory study 2

In order to test the paraphrasis with style transfer, we use the following configurations:

Setting 1: Style former from casual to formal.

Setting 2: Style former from formal to casual.

Setting 3: Style transfer with Tweets style (split + 1 step pipeline)

Setting 4: Style transfer with Tweets style (split + 2-steps pipeline)

Setting 5: Style transfer with Switchboard style (no split + 1 step pipeline)

Setting 6: Style transfer with Switchboard style (split + 1 step pipeline)

Once again, we select 12 dialogues for each setting, with 3 candidate paraphrases generated for each CN. The instructions given to the expert annotator are the same as in the first exploratory study.

Results are showed in Table 8 and can be summed up as follows:

Tweets: setting 4 is achieving the highest HTER CN-pselp_{sel} and HTER CN-pedp_{ed} while having a HTER pselp_{sel}-pedp_{ed} in the middle. We would prefer it to setting 3 which instead has almost the same HTER pselp_{sel}-pedp_{ed} but a much lower HTER CN-pselp_{sel} and HTER CN-pedp_{ed};

Formal and informal: both setting 1 and setting 2 achieve high HTER CN-pselp_{sel} and HTER CN-pedp_{ed} while low HTER pselp_{sel}-pedp_{ed} with formal performing slightly better;

Switchboard: setting 5 is preferable to setting 6 since it has a lower HTER pselp_{sel}-pedp_{ed}.

According to these results, we choose to employ setting 1, 2, 4 and 5 for the paraphrasis session.

Appendix B Session 3: Training details

For reproducibility purposes, we report here the parameters employed for fine-tuning the LMs used in Session 3. For each model, we used a version smaller than the largest available, i. e. the medium version for DialoGPT and the base version of T5. We used Optuna to conduct a hyperparameters search with 10 trials, and we selected the trial achieving the lowest evaluation loss. The search space for the parameters of our interest was the following: learning-rate: {1e−5, 2e−5, 3e−5, 4e−5, 5e−5}\{1e-5,\ 2e-5,\ 3e-5,\ 4e-5,\ 5e-5\}, warmup-ratio: {0, 0.1}\{0,\ 0.1\}, batch size: {1, 2, 4}\{1,\ 2,\ 4\}, number of epochs: {2, 3, 5}\{2,\ 3,\ 5\}. The selected parameters for each model are summed up in Table 9.

Appendix C Reviewing examples

Table 10 shows an example of turns swap from Session 1: CN2 is a question that can be answered with HS0, so it is moved at the beginning. At the same time, concluding the dialogue with the most substantial CN, i. e. CN1, makes the dialogue stronger. HS1 is modified by the addition of ‘because’ in order to be linguistically aligned with the preceding turn, which is a question.

In table 11 an example of a dialogue resulting from the concatenation of similar HS is showed. The high repetitiveness makes it necessary to remove the pair HS2, CN2 and to modify HS1.

In table 12 an example of CN post-editing coming from Session 2 is showed: the selected paraphrasis is modified in order to be as much different from the original as possible, while keeping the dialogue flow naturally. The paraphrases of CN1 and CN2 are highly similar to the original text, and require a major intervention from the annotator.

Table 13 and Table 14 show two peculiar cases of the annotators’ intervention in Session 3. In table 13, HS3 and CN3 are swapped because CN3 contains hateful content , while HS3 is a CN. For the same reason it was necessary to post-edit CN2.

In table 14, an example of a dialogue generated with T5, characterised by a poorly varied content. Both HS and CN are edited a lot to make the dialogue more natural.