Everyone's Voice Matters: Quantifying Annotation Disagreement Using Demographic Information

Ruyuan Wan, Jaehyung Kim, Dongyeop Kang

Introduction

Supervised AI systems are trained on annotated datasets with labels determined by consensus among multiple annotators. The subjective opinions of different annotators often bring annotation disagreement in the decision of the final labels. Most commonly, this disagreement is addressed by ignoring highly-disagreed cases and only including those whose opinions were voted on by the majority as the final label. When the labeling tasks become more subjective and require the annotator’s own interpretation and judgment, such as detecting offensiveness and judging social dilemmas (Reidsma and op den Akker 2008), this majority-based aggregation often fails to learn the true distribution of annotators’ voices. The increasing subjectivity of NLP problems in modern NLP will cause the annotators’ disagreement not only due to the potential random error in the process but also because annotators may interpret the text with different views and make judgments based on their own connotations.

Different demographics, cultural backgrounds, and living experiences influence how people receive and interpret information. This difference is more visible in subjective tasks. For example, Sap et al. found more consecutive annotators who have higher scores on racist beliefs are more likely to label African American English as toxic rather than label anti-Black language as toxic (Sap et al. 2022). In these cases, the aggregated singular labels can bring bias by using less inclusive and societal-representative labels to accommodate everyone’s voices in subjective studies.

This paper assumes that annotators’ disagreement potentially comes from the limited representations of the annotator group assigned or controversy of the text in nature. This study focuses on exploring the relationship between the annotator group and natural controversy in text by developing a disagreement predictor with and without the information about the annotator group, as depicted in Figure 1. In particular, we analyze annotators’ disagreement from five subjective task datasets to answer the following research questions:

Is it possible to predict the level of annotators’ disagreement with text using language models? Does knowing annotators’ identities, like demographic information in addition to text, help predict annotation disagreement?

Is the disagreement caused by the natural controversy of the text or by the biased distribution of the assigned annotators?

Our research demonstrates that disagreement is predictable in subjective annotation tasks by using Roberta model (Liu et al. 2019) to predict both hard disagreement (binary label) and soft disagreement (continuous label). We further design two demographic augmentation experiments and find that bringing the annotator-level demographic information can significantly improve disagreement prediction performance. Finally, based on the findings, we simulated artificial annotators’ backgrounds to predict disagreement to check whether the disagreement score will be changed in a wider annotator population. In short, we propose a disagreement measurement that can efficiently suggest the optimal number of annotators and assign an appropriate demographic group of annotators per text, possibly helping improve the fairness and quality of subjective annotation.

Related Works

Tasks like toxicity detection (Sap et al. 2020; Yu 2022), sentiment analysis (Potts et al. 2021), and social, ethical labeling (Forbes et al. 2020; Hendrycks et al. 2021) are highly subjective and controversial. One can think one post is offensive, but others may consider it acceptable. There is no one objective ground truth. Röttger et al. summarized three key challenges of descriptive annotation in subjective NLP tasks: interpretation of disagreement, label aggregation, and representativeness of annotators.

Due to the absence of absolute ground truth, the interpretation of disagreement becomes complicated (Alm 2011). For instance, the disagreement may result in considerably different reliabilities: whether the annotators disagree on the most critical or least crucial instances (Foley 2018). In addition, researchers commonly use some notion of the agreement to measure the task’s subjectiveness, such as using inter-annotator agreement metrics Cohen’s Kappa (Cohen 1960) or Fleiss’ Kappa (Fleiss 1971) to measure annotations’ reliability. But when presenting the final results in the downstream task, people usually use the aggregated labels that can conceal informative disagreement and evaluation metrics that are unaware of the task’s subjectiveness (Röttger et al. 2022).

Rather than ’correct’ or ’wrong,’ Alm pointed out the concept of acceptability. There might be multiple acceptable answers in subjective tasks. However, aggregating labels through major voting will increase the risk of discarding minority voices. To address the problem, Davani, Díaz, and Prabhakaran treats predicting each annotator’s judgments as separate subtasks, which achieved the same or better performance than aggregating labels in the data before training. They also further evaluate the model uncertainty using the variance of the predicted annotation label. However, this still concerns that recognizing the aggregated major votes as the final targets does not always represent all acceptable answers. On the other hand, Uma, Almanea, and Poesio used posterior calibration with a soft-loss approach to learning from data containing disagreement. They noticed that temperature scaling only functions with data where disagreements are caused by label overlap and not with data where disagreements are caused by annotator subjective judgment or language ambiguity. This aligns with Foley ’s finding for tasks with subjective labels: without collecting additional labels, models reach the ceiling of performance given the small dataset size and the inherent disagreement between annotators on which documents are controversial.

In previous research, annotators’ demographics have shown importance in improving the annotation quality in subjective tasks. For example, Prabhakaran, Davani, and D’iaz demonstrated that the agreement scores could be very significant among different socio-demographic groups annotators identified when certain individual annotators disagree with the majority labels. Further, Gordon et al. proposed jury learning, a recommender system approach defining which people or groups, in what proportion, determine the classifier’s prediction. For instance, a jury learning model would recommend women and black jurors for online hate speech detection, who are mainly targeted in online harassment. However, many public datasets didn’t collect annotators’ demographics with their annotations. Further, the datasets that reported the annotator’s demographics also have imbalanced representative concerns. For example, the race is often skewed, and dominant with the white race (Sap et al. 2020; Forbes et al. 2020; Hendrycks et al. 2021; Sap et al. 2022).

As shown above, researchers have implicitly resolved the label disagreement using majority votes, annotator selection, and the soft-loss approach(Uma et al. 2021). Different from the interrater disagreement resolution, which defines disagreement as a sign of poor quality or mistakes to be resolved(Oortwijn, Ossenkoppele, and Betti 2021), our research explicitly quantifies disagreement as our task target and further distinguishes the nuance among various socio-demographic groups.

Methods

This section presents our method for quantifying subjective annotation disagreements. Our main idea is modeling the annotation disagreement using demographic information of each annotator as additional inputs, with the pre-trained language model, e.g., RoBERTa (Liu et al. 2019). In Section 3.1, we first introduce the mathematical notations. Then, we elaborate on the details of the proposed method in Section 3.2. Finally, we provide a way to simulate the annotators’ demographic information in Section 3.3.

2 Disagreement Prediction with Demographic Information

Our goal is to predict the disagreement rˉ(x)\bar{r}(\mathbf{x}) of given text x\mathbf{x} because it provides an effective way to understand which content is controversial or not. To this end, our first idea is to utilize the pre-trained language model, e.g., RoBERTa (Liu et al. 2019), for training a predictor fθf_{\theta} of the disagreement of given text. Specifically, we train the model by minimizing a mean square error (MSE) loss as follow:

However, the annotators’ disagreement is not only from the controversy of the text in nature but also from the limited representations of the assigned annotator group. Hence, more than just using text as input is needed to capture the disagreement fully.

Incorporation of demographics: Group vs Personal. To this end, our key idea is incorporating the demographic information of annotators {d(t)(x)}t=1T\{\mathbf{d}^{\tt(t)}(\mathbf{x})\}_{t=1}^{T} to train the model fθf_{\theta}. Intuitively, it is expected to encode the valuable information of the disagreement of the text x\mathbf{x}, especially related to limited representations of the annotator group assigned. To be specific, we propose two different ways to incorporate the demographic information: (1) Text with group demographic information and (2) Text with personal demographic information.

Text with group demographic information x~group\widetilde{\mathbf{x}}_{\tt group} is constructed by listing all NN annotators’ information d(t)(x)\mathbf{d}^{\tt(t)}(\mathbf{x}) in one string and then concatenating with the targeted text x\mathbf{x}:

Therefore, the group demographics supplemented text also has the same number of instances as the original dataset.

On the other hand, text with personal demographic information x~person\widetilde{\mathbf{x}}_{\tt person} is constructed by concatenating only one annotator’s demographic with text:

where j=1,…,Nj=1,\dots,N and hence it results in NN times larger dataset with NN different annotators.

Format: Templated vs Sentence. For combining the demographic information and text, we further propose two different ways with specific templates: (1) Templated format and (2) Sentence format. Templated format represents the category and value of each demographic information in a separate sentence, then concatenate all of them with the given text. For example, if one annotator is 36 years white woman, this demographic information is converted to ”Age: 36, Color: white, Gender: women”, then concatenated with the original sentence in case of the text with person demographic. On the other hand, sentence format represents the demographic information with a natural sentence, e.g., the annotator is a 36 years old white woman., then concatenate it with the original sentence.

With these demographics supplemented text x~\widetilde{\mathbf{x}} (x~group(\widetilde{\mathbf{x}}_{\tt group} or x~person)\widetilde{\mathbf{x}}_{\tt person}), we train our model similar to the case with the original sentence x\mathbf{x} in Equation (1):

An illustration of the proposed demographic-based disagreement predictor is presented in Figure 2.

3 Simulation of Demographic Information

In addition, we propose a simulation of demographic information, which is a novel approach to analyze how the different annotator groups impact disagreement prediction. It is expected to separately reveal the inherent disagreement of annotators from the controversy of the text in nature. Specifically, instead of ground-truth {d(t)(x)}t=1T\{\mathbf{d}^{(t)}(\mathbf{x})\}_{t=1}^{T}, we combine the artificial demographic information {dˉ(t)(x)}t=1T\{\bar{\mathbf{d}}^{(t)}(\mathbf{x})\}_{t=1}^{T} with the given text x\mathbf{x} and annotations y(x)\mathbf{y}(\mathbf{x}), to simulate the scenario with different annotators. Such as, the gender demographic type has four possible options: woman, man, transgender, non-binary; and ethnicity with seven options: white, black or African American, American Indian or Alaska Native, Asian, Native Hawaiian or other pacific islanders, Hispanic, or some other race. Overall, we have a total 28=4×728=4\times 7 different combinations of the annotator’s demographic information for the simulation, while the ground-truth demographic information is one of them; hence, it offers an opportunity to explore the more extensive range of demographic information with the increased number of instances. Then, we obtain a predicted disagreement using fθf_{\theta}, which is trained with x\mathbf{x} and {d(t)(x)}t=1T\{\mathbf{d}^{(t)}(\mathbf{x})\}_{t=1}^{T} as introduced in Section 3.2.

Then, we evaluate whether the predicted disagreement is easily or hard to be changed among the simulated demographic profiles so that we can distinguish whether the disagreement comes from the controversy of text or uncertainty from annotators for the disagreement label. For example, if the variation of predicted disagreement among the simulated combinations is high and the average change of the predicted disagreement between the simulated combinations and real disagreement is large, it might reveal that disagreement is highly related to the uncertainty of annotators. In contrast, the lower variation and smaller change between real disagreements indicate the disagreement is based on the controversy in the text, which is stable disagreement among various kinds of people.

Experiments

To obtain the annotators’ disagreement, we choose the following five datasets of subjective tasks that include annotators’ voting records in the raw format.Note that only the SBIC and SChem101 datasets report annotators’ demographic information, so we used these two datasets to evaluate the effect of including demographic information in disagreement prediction.

Social Bias Inference Corpus (SBIC) (Sap et al. 2020) contains 150k structured annotations of social media posts. Each post has three different annotators. Annotators indicated whether the post could be considered “offensive to anyone.” The offensiveness is a categorical variable with three possible answers (yes, maybe, no).

Social Chemistry 101 (SChem101) (Forbes et al. 2020) is a corpus of cultural norms via free-text rules-of-thumb created by crowd workers. A rule-of-thumb is a judgment of action which is further broken down into 12 theoretically-motivated dimensions of people’s judgments. Our study focuses on the anticipated agreement category. It reflects workers’ opinion on what portion of people probably agree with the judgment given the action. The category has five possible answers: almost no one believes, people occasionally think this, controversial, common belief, universally true. Each rule of thumb is annotated by five workers.

Scruples-dilemmas (Lourie, Bras, and Choi 2021) is a resource for normative ranking actions. Each instance pairs two unrelated actions and identifies which action crowd workers found less ethical. Each instance is annotated by five different annotators.

Dyna-Sentiment (Potts et al. 2021) is an English language benchmark task for ternary sentiment analysis. Each Yelp review is validated by five crowd workers into three possible sentiment results: positive, negative, and neutral.

Wikipedia Politeness (Danescu-Niculescu-Mizil et al. 2013) is a collection of requests from Wikipedia Talk pages, annotated with politeness. Each Wikipedia request is annotated by five annotators on a 1 to 25 scale. As Danescu-Niculescu-Mizil et al. ignored neutral cases for politeness prediction, we extracted the disagreement between the binary classes of request, i.e., polite and impolite.

The Figure 3 shows the distributions of disagreement scores among five datasets. For dynasent dataset, since the majority of the dataset has disagreement between 0.3 to 0.6. The prediction concentrate around 0.4 to 0.5. The comparison among multiple datasets reflects that the subject topics influence the crowd annotators’ disagreement. For example, most texts regarding offensiveness had consensus opinions from the annotators, while most annotators disagreed regarding sentiment.

2 Experimental Details

All the experiments are conducted by fine-tuning RoBERTa-base (Liu et al. 2019) using Adam optimizer (Kingma and Ba 2015) with a fixed learning rate 1e-5 and the default hyperparameters of Adam. For the text classification tasks, the model is fine-tuned with batch size 8 for 15 epochs.

To the best of our knowledge, we couldn’t find any existing disagreement predictors to be used as baselines. As a result, we compare our predictors with different input types and disagreement labeling setups. Different versions of pre-trained language models were tested, but RoBERTa always performed better. For the evaluation of the performance of the trained disagreement predictor, we use both 1) hard score F1 and 2) soft score Mean Square Error (MSE), and compare the measurement effect of binary disagreement label and continuous disagreement rate.

3 Main Results

From Table 2, we notice that continuous disagreement achieves better prediction than binary disagreement for most of the datasets. Among the datasets, the disagreement prediction models work the best in the SBIC dataset. The binary label prediction are close to continuous prediction for SBIC and Politeness datasets. But SChem and Dilemmas have 0 F1 scores which only give 0 outputs. That means the binary label is not reliable for the two datasets.

For Dynasent, the binary label has an inconsistent performance based on hard score F1 and soft score MSE. We think one potential reason is that the binary disagreement is highly unbalanced while converting a continuous prediction to categorical labels like 0, 0.33, 0.67, and 1 is easy to accidentally assign an intermediate value to a wrong group. Therefore, even though we used both F1 and MSE metrics, they are used to have a parallel comparison between the binary label and continuous label setup. Among the binary classification, we consider F1 as the metric of model goodness, on the opposite, we use MSE to evaluate the regression fitness.

Disagreement prediction with text and demographic information

Further, by comparing different experiment setups for disagreement with demographic information in Table 3, we focus on the different effects of a group of demographics or the personal level of demographics. The results show that personal-level demographics improve the disagreement prediction more than group-level demographics. One potential reason is that the annotator’s level of demographics may imitate the annotation process that each annotator labels the text without knowing each other. And also because concatenating personal level demographics can be considered as oversampling that group-level setup can not.

Qualitative Results Analysis

Lastly, we categorize prediction into four types and provide an example per each in Table 4. Using Local Interpretable Model-Agnostic Explanations (LIME) (Ribeiro, Singh, and Guestrin 2016), we found demographics have been used in prediction with the text. Blue text is important feature to predict agreement while orange text is used to predict disagreement.

4 Simulation of Everyone’s Voices with Artificial Demographics

One remaining question is how to reflect everyone’s diverse opinions on such subjective and socially sensitive annotation tasks. To explore this aspect, we run additional experiments with the simulated demographics introduced in Section 3.3. Namely, we simulate a different combination of all possible artificial demographic groups, rather than using the real annotators’ demographics used in model training (Section 4.3). Then, the disagreement of the simulated demographic information and the text is predicted using the fine-tuned disagreement predictor introduced in Section 3.2.

Our study is motivated by the Intersectionality theory (Crenshaw 1990), assuming that people’s perspectives are shaped by the intersection of all available demographic categories. We set four gender types, seven ethnicity types, and five age ranges (see A.1 for details), and thus we have 140 (4×\times7×\times5) artificial annotators’ unique demographic characteristics. Since we only trained our disagreement predictor with demographic information on SBIC and SChem101 datasets, the simulation experiments are also applied to these two datasets. We randomly sampled 600s text instances in each dataset and concatenated them with 140 artificial annotators’ demographic information in the colon template to predict continuous disagreement.

To visualize the simulation result of 140 artificial annotators per text, we made a scatter plot based on the mean and variance of 140 disagreement prediction as shown in Figure 4. The color and shape denoted at the legend shows the text’s disagreement label in the original dataset. The higher points in the plot means higher predicted disagreement rate. The more rightward point implies a greater variance in the disagreement prediction among the 140 artificial annotators. The difficulty of disagreement prediction is related to the dataset’s topic, quality etc. SBIC is collected from social media data while SChem is created by crowdsourcing, which might explain why the clusters are more clear in the Figure 4(b) than in the Figure 4(a). From 4(b), most text are predicted into corresponding disagreement clusters. But some outliers are predicted to be more controversial or agreeable. For example, the circled outlier has an original 0.5 disagreement label but ends up with a 0.04 disagreement prediction among 140 artificial annotators. The text is ”It is understandable to want to spend holidays with family.” Those outliers in the simulation experiment show the disagreement rate would change if the annotator change. Other than the outliers, the disagreement clusters shows they are less influenced by annotator change. With this simulation, we can distinguish disagreements caused by the natural controversy of the text or by the biased distribution of the assigned annotators.

Discussion and Future Work

We could think of potential applications in NLP data annotation pipeline using our disagreement prediction model:

Annotator number estimation. We could potentially use the predicted disagreement score in order to decide the appropriate number of annotators in a cost-efficient manner, e.g., we may not need three or five annotators for the text being predicted zero disagreements. For instance, we may need one or two annotators if a text is predicted to have lower disagreement scores. Other than that, we can assign five or even more annotators to those texts being predicted as highly disagreeable.

Annotator group assignment. Additionally, we suggest considering the annotation disagreement as a critical factor in finding the optimal group of annotator pools. This can be used as a novel annotator assignment supporting system for the data annotation pipeline. In the current annotator recruiting process, there is usually some uncontrollable randomness from annotators, either from skewed representatives or individual variations. We present a low-cost approach to simulate as diverse as possible artificial annotation pools to identify the controversial samples that maximize the disagreement. Thus, we avoid ignoring human bias and listening to opinions from a more diverse group of people to avoid polarized analysis. We hope our study can evoke others’ attention in designing a more fair and representative annotation pipeline.

Potential risk of using demographic information. Last but not least, though our research shows that annotators’ demographics help disagreement prediction, we should be careful about collecting private and personal information. Also, we admit that NLP or AI systems trained on demographic information might make another bias toward certain demographic groups.

Conclusion

Overall, we propose a disagreement prediction framework that measures annotators’ disagreement in subjective tasks, predicts disagreement with/without demographic information and simulates 140 artificial annotators to build a relatively fair annotation pool. Our results show that the annotators’ disagreement could be fairly predictable from the text and even better performs when we know the demographic information of the annotators. With our disagreement predictor, we believe we could shed light on various applications of data annotation in a more effective and inclusive manner.

Acknowledgments

We thank Dr. Maxwell Forbes for sharing the demographic information data for Social Chemistry 101 dataset. We also thank the anonymous reviewers and Minnesota NLP members for their insightful comments and suggestions.

References

Appendix A Appendices

We set gender with four possible options: woman, man, transgender, non-binary; ethnicity with seven options: white, black or African American, American Indian or Alaska Native, Asian, Native Hawaiian or other pacific islanders, Hispanic, or some other race. Also, we set the age with five ranges: 18 to 29, 30 to 39, 40 - 49, 50-59, and 60 to elder.

A.2 Annotators Distributions

Our analysis finds that the annotators’ pool in the SBIC dataset was relatively gender-balanced and age-balanced (55% women, 42% men, 1% non-binary; 36±10 years old), but racially skewed (82% White, 4% Asian, 4% Hispanic, 4% Black). And it was also politically skewed (63% liberal, 20% conservative). Overall, workers agreed on a post being offensive at a rate of 76%. Later, Sap et al. showed that annotator identity and beliefs are highly related to their toxicity ratings in their annotators with attitudes paper (Sap et al. 2022). Similar to the demographic distribution in the SBIC dataset, the crowd worker pool in SChem101 is also gender-balanced and race-skewed: 55% were women and 45% men. 89% of workers identified as white, 7% as Black. 39% were in the 30-39 age range, 27% in the 21-29, and 19% in the 40-49 age range. Regarding education, 44% had a bachelor’s degree, and 36% had some college experience or an associate’s degree. However, even though some people consider one rule as a common belief, other people may think no one believes it.

A.3 Group v.s. Personal Demographics Setup

Table 5 shows one example of text with individual annotators’ demographics or with the group of annotators’ demographics.

A.4 Disagreement Prediction Given Only Or Partial Demographics

To further evaluate how annotators’ demographics influence disagreement prediction, we also tested the inputs of only demographics, which performed much worse than the inputs including text. Notably, this experimental input setup might mislead, assuming people from certain social groups always have the kind of opinion regardless of the text context.

Based on our above study, we controlled the demographics in the templated format of individual annotators and the label in the continuous format, which is the optimal setup. And we evaluated the input of text with partial demographic information as shown in Table 6. It shows that the predictions are given input of text with a single demographic factor, or only demographics perform worse than predicting with text and intersectional demographic information. We also tried using random forests given only demographic features to predict annotation disagreement. The age feature was the most important.

A.5 Results of Other Language Models on Disagreement Prediction

We only reported Roberta in our main paper, which showed the best performance. But we have also conducted experiments with other language models like BERT(Devlin et al. 2018), XLNet(Yang et al. 2019), and AlBERTa(Lan et al. 2019). Table 7 shows the other language models’ prediction results on SChem as an example.