Scruples: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes

Nicholas Lourie, Ronan Le Bras, Yejin Choi

Introduction

State-of-the-art techniques excel at syntactic and semantic understanding of text, reaching or even exceeding human performance on major language understanding benchmarks (Devlin et al. 2019; Lan et al. 2019; Raffel et al. 2019). However, reading between the lines with pragmatic understanding of text still remains a major challenge, as it requires understanding social, cultural, and ethical implications. For example, given “closing the door in a salesperson’s face” in Figure 1, readers can infer what is not said but implied, e.g., that perhaps the house call was unsolicited. When reading narratives, people read not just what is stated literally and explicitly, but also the rich non-literal implications based on social, cultural, and moral conventions.

Beyond narrative understanding, AI systems need to understand people’s norms, especially ethical and moral norms, for safe and fair deployment in human-centric real-world applications. Past experiences with dialogue agents, for example, motivate the dire need to teach neural language models the ethical implications of language to avoid biased and unjust system output (Wolf, Miller, and Grodzinsky 2017; Schlesinger, O’Hara, and Taylor 2018).

However, machine ethics poses major open research challenges. Most notably, people must determine what norms to build into systems. Simultaneously, systems need the ability to anticipate and understand the norms of the different communities in which they operate. Our work focuses on the latter, drawing inspiration from descriptive ethics, the field of study that focuses on people’s descriptive judgements, in contrast to prescriptive ethics which focuses on theoretical prescriptions on morality (Gert and Gert 2017).

As a first step toward computational models that predict communities’ ethical judgments, we present a study based on people’s diverse ethical judgements over a wide spectrum of social situations shared in an online community. Perhaps unsurprisingly, the analysis based on real world data quickly reveals that ethical judgments on complex real-life scenarios can often be divisive. To reflect this real-world challenge accurately, we propose predicting the distribution of normative judgments people make about real-life anecdotes. We formalize this new task as Who’s in the Wrong? (WHO), predicting which person involved in the given anecdote would be considered in the wrong (i.e., breaking ethical norms) by a given community.

Ideally, not only should the model learn to predict clean-cut ethical judgments, it should also learn to predict if and when people’s judgments will be divisive, as moral ambiguity is an important phenomenon in real-world communities. Recently, Pavlick and Kwiatkowski (2019) conducted an extensive study of annotations in natural language inference and concluded that diversity of opinion, previously dismissed as annotation “noise”, is a fundamental aspect of the task which should be modeled to accomplish better language understanding. They recommend modeling the distribution of responses, as we do here, and found that existing models do not capture the kind of uncertainty expressed by human raters. Modeling the innate ambiguity in ethical judgments raises similar technical challenges compared to clean-cut categorization tasks. So, we investigate a modeling approach that can separate intrinsic and model uncertainty; and, we provide a new statistical technique for measuring the noise inherent in a dataset by estimating the best possible performance.

To facilitate progress on this task, we release a new challenge set, Scruples:Subreddit Corpus Requiring Understanding Principles in Life-like Ethical Situations a corpus of more than 32,000 real-life anecdotes about complex ethical situations, with 625,000 ethical judgments extracted from reddit.https://reddit.com: A large internet forum. The dataset proves extremely challenging for existing methods. Due to the difficulty of the task, we also release Dilemmas: a resource of 10,000 actions with normative judgments crowd sourced from Mechanical Turk. Our results suggest that much of the difficulty in tackling Scruples might have more to do with challenges in understanding the complex narratives than lack of learning basic ethical judgments.

Define a novel task, Who’s in the Wrong? (WHO).

Release a large corpus of real-life anecdotes and norms extracted from an online community, reddit.

Create a resource of action pairs with crowdsourced judgments comparing their ethical content.

Present a new, general estimator for the best possible score given a metric on a dataset.Try out the estimator at https://scoracle.apps.allenai.org.

Study models’ ability to predict ethical judgments, and assess alternative likelihoods that capture ambiguity.Demo models at https://norms.apps.allenai.org.

Datasets

Scruples has two parts: the Anecdotes collect 32,000 real-life anecdotes with normative judgments; while the Dilemmas pose 10,000 simple, ethical dilemmas.

The Anecdotes relate something the author either did or considers doing. By design, these anecdotes evoke norms and usually end by asking if the author was in the wrong. Figure 1 illustrates a typical example.

Each anecdote has three main parts: a title, body text, and label scores. Titles summarize the story, while the text fills in details. The scores tally how many people thought the participant broke a norm. Thus, after normalization the scores estimate the probability that a community member holds that opinion. Table 2 provides descriptions and frequencies for each label. Predicting the label distribution from the anecdote’s title and text makes it an instance of the WHO task.

In addition, each story has a type, action, and label. Types relate if the event actually occurred (historical) or only might (hypothetical). Actions extract gerund phrases from the titles that describe what the author did. The label is the highest scoring class.

Scruples offers 32,766 anecdotes totaling 13.5 million tokens. Their scores combine 626,714 ethical judgments, and 94.4% have associated actions. Table 1 expands on these statistics. Each anecdote exhibits high lexical diversity with words being used about twice per story. Moreover, most stories have enough annotations to get some insight into the distribution of ethical judgments, with the median being eight.

To study norms, we need representative source material: real-world anecdotes describing ethical situations with moral judgments gathered from a community. Due to reporting bias, fiction and non-fiction likely misrepresent the type of scenarios people encounter (Gordon and Van Durme 2013). Similarly, crowdsourcing often leaves annotation artifacts that make models brittle (Gururangan et al. 2018; Poliak et al. 2018; Tsuchiya 2018). Instead, Scruples gathers community judgments on real-life anecdotes shared by people seeking others’ opinions on whether they’ve broken a norm. In particular, we sourced the raw data from a subforum on reddithttps://reddit.com/r/AmItheAsshole, where people relate personal experiences and then community members vote in the comments on who they think was in the wrong.Scruples v1.0 uses the data from 11/2018–4/2019. Each vote takes the form of an initialism: yta, nta, esh, nah, and info, which correspond to the classes, author, other, everyone, no one, and more info. Posters also title their anecdotes and label if it’s something that happened, or something they might do. Since all submissions are unstructured text, users occasionally make errors when providing this information.

Each anecdote derives from a forum post and its comments. We obtained the raw data from the Pushshift Reddit Dataset (Baumgartner et al. 2020) and then used rules-based filters to remove undesirable posts and comments (e.g. for being deleted, from a moderator, or too short). Further rules and regular expressions extracted the title, text, type, and action attributes from the post and the label and scores from the comments. To evaluate the extraction, we sampled and manually annotated 625 posts and 625 comments. Comments and posts were filtered with an F1 of 97% and 99%, while label extraction had an average F1 of 92% over the five classes. Tables 3 and 4 provide more detailed results from the evaluation.

Each anecdote’s individual components are extracted as follows. The title is just the post’s title. The type comes from a tag that the subreddit requires titles to begin with (“AITA”, “WIBTA”, or “META”).Respectively: “am I the a-hole”, “would I be the a-hole”, and “meta-post” (about the subreddit). We translate AITA to historical, WIBTA to hypothetical, and discard META posts. A sequence of rules-based text normalizers, filters, and regexes extract the action from the title and transform it into a gerund phrases (e.g. “not offering to pick up my friend”). 94.4% of stories have successfully extracted actions. The text corresponds to the post’s text; however, users can edit their posts in response to comments. To avoid data leakage, we fetch the original text from a bot that preserves it in the comments, and we discard posts when it cannot be found. Finally, the scores tally community members who expressed a given label. To improve relevance and independence, we only consider comments replying directly to the post (i.e., top-level comments). We extract labels using regexes to match variants of initialisms used on the site, and resolve multiple matches using rules.

2 Dilemmas

Beyond subjectivity (captured by the distributional labels), norms vary in importance: while it’s good to say “thank you”, it’s imperative not to harm others. So, we provide the Dilemmas: a resource for normatively ranking actions. Each instance pairs two actions from the Anecdotes and identifies which one crowd workers found less ethical. See Figure 2 for an example. To enable transfer as well as other approaches using the Dilemmas to solve the Anecdotes, we aligned their train, dev, and test splits.

For each split, we made pairs by randomly matching the actions twice and discarding duplicates. Thus, each action can appear at most two times in Dilemmas.

We labeled each pair using 5 different annotators from Mechanical Turk. The dev and test sets have 5 extra annotations to estimate human performance and aid error analyses that correlate model and human error on dev. Before contributing to the dataset, workers were vetted with Multi-Annotator Competence Estimation (MACE)Code at https://github.com/dirkhovy/MACE (Hovy et al. 2013).Data used to qualify workers is provided as extra train. MACE assigns reliability scores to workers based on inter-annotator agreement. See Paun et al. (2018) for a recent comparison of different approaches.

Methodology

Ethics help people get along, yet people often hold different views. We found communal judgments on real-life anecdotes reflect this fact in that some situations are clean-cut, while others can be divisive. This inherent subjectivity in people’s judgements (i.e., moral ambiguity) is an important facet of human intelligence, and it raises unique technical challenges compared to tasks that can be defined as clean-cut categorization, as many existing NLP tasks are often framed.

In particular, we identify and address two problems: estimating a performance target when human performance is imperfect to measure, and separating innate moral ambiguity from model uncertainty (i.e., a model can be certain about the inherent moral ambiguity people have for a given input).

For clean-cut categorization, human performance is easy to measure and serves as a target for models. In contrast, it’s difficult to elicit distributional predictions from people, making human performance hard to measure for ethical judgments which include inherently divisive cases. One solution is to ensemble many people, but getting enough annotations can be prohibitively expensive. Instead, we compare to an oracle classifier and present a novel Bayesian estimator for its score, called the Best performance,Bayesian Estimated Score Terminus available at https://scoracle.apps.allenai.org.

To estimate the best possible performance, we must first define the oracle classifier. For clean-cut categorization, an oracle might get close to perfect performance; however, for tasks with innate variability in human judgments, such as the descriptive moral judgments we study here, it’s unrealistic for the oracle to always guess the label a particular human annotator might have chosen. In other words, for our study, the oracle can at best know how people annotate the example on average.This oracle is often called the Bayes optimal classifier. Intuitively, this corresponds to ensembling infinite humans together.

Formally, for example ii, if NiN_{i} is the number of annotators, YijY_{ij} is the number of assignments to class jj, pijp_{ij} is the probability that a random annotator labels it as class jj, and Yi:Y_{i:} and pi:p_{i:} are the corresponding vectors of class counts and probabilities, then the gold annotations are multinomial:

The oracle knows the probabilities, but not the annotations. For cross entropy, the oracle gives pi:p_{i:} as its prediction, p^i:\hat{p}_{i:}:For hard-labels, we use the most likely class. This choice isn’t optimal for all metrics but matches common practice.

We use this oracle for comparison on the evaluation data.

The Best Performance

Even if we do not know the oracle’s predictions (i.e., each example’s label distribution), we can estimate the oracle’s performance on the test set. We present a method to estimate its performance from the gold annotations: the Best performance.

Since Yi:Y_{i:} is multinomial, we model pi:p_{i:} with the conjugate Dirichlet, following standard practice (Gelman et al. 2003):

In particular, for cross entropy on soft labels:

Simulation Experiments

To validate Best, we ran three simulation studies comparing its estimate and the true oracle score. First, we simulated the Anecdotes’ label distribution using a Dirichlet prior learned from the data (Anecdotes). Second, we simulated each example having three annotations, to measure the estimator’s usefulness in typical annotation setups (3 Annotators). Last, we simulated when the true prior is not a Dirichlet distribution but instead a mixture, to test the estimator’s robustness (Mixed Prior). Table 5 reports relative estimation error in each scenario.

2 Separating Controversiality from Uncertainty

Most neural architectures confound model uncertainty with randomness intrinsic to the problem.Model uncertainty and intrinsic uncertainty are also often called epistemic and aleatoric uncertainty, respectively (Gal 2016). For example, softmax predicts a single probability for each class. Thus, 0.5 could mean a 50% chance that everyone picks the class, or a 100% chance that half of people pick the class. That singular number conflates model uncertainty with innate controversiality in people’s judgements.

To separate the two, we modify the last layer. Instead of predicting probabilities with a softmax

We make activations positive with an exponential

and use a Dirichlet-Multinomial likelihood:

In practice, this modification requires two changes. First, labels must count the annotations for each class rather than take majority vote; and second, a one-line code change to replace the loss with the Dirichlet-Multinomial one.Our PyTorch implementation of this loss may be found at https://github.com/allenai/scruples.

With the Dirichlet-Multinomial likelihood, predictions encode a distribution over class probabilities instead of singular point estimates. Figures 4 and 4 visualize examples from the Dilemmas and Anecdotes. Point estimates are recovered by taking the mean predicted class probabilities:

Which is mathematically equivalent to a softmax. Thus, Dirichlet-multinomial layers generalize softmax layers.

3 Recommendations

Synthesizing results, we propose the following methodology for NLP tasks with labels that are naturally distributional:

Rather than evaluating hard predictions with metrics like F1, experiments can compare distributional predictions with metrics like total variation distance or cross-entropy, as in language generation. Unlike generation, classification examples often have multiple annotations and can report cross-entropy against soft gold labels.

Many models are poorly calibrated out-of-the-box, so we recommend calibrating model probabilities via temperature scaling before comparison (Guo et al. 2017).

Human performance is a reasonable target on clean-cut tasks; however, it’s difficult to elicit human judgements for distributional metrics. Section 3.1 both defines an oracle classifier whose performance provides the upper bound and presents a novel estimator for its score, the Best performance. Models can target the Best performance in clean-cut or ambiguous classification tasks; though, it’s especially useful in ambiguous tasks, where human performance is misleadingly low.

Softmax layers provide no way for models to separate label controversiality from model uncertainty. Dirichlet-multinomial layers, described in Section 3.2, generalize the softmax and enable models to express uncertainty over class probabilities. This approach draws on the rich tradition of generalized linear models (McCullagh and Nelder 1989). Other methods to quantify model uncertainty exist as well (Gal 2016).

Table 6 summarizes these recommendations. Some recommendations (i.e., targeting the Best score) could also be adopted by clean-cut tasks.

Experiments

To validate Scruples, we explore two questions. First, we test for discernible biases with a battery of feature-agnostic and stylistic baselines, since models often use statistical cues to solve datasets without solving the task (Poliak et al. 2018; Tsuchiya 2018; Niven and Kao 2019). Second, we test if ethical understanding challenges current techniques.

The following paragraphs describe the baselines at a high level.

Feature-agnostic baselines use only the label distribution, ignoring the features. Prior predicts the class probability for each label, and Sample assigns all probability to one class drawn from the label distribution.

Stylistic baselines probe for stylistic artifacts that give answers away. Style applies a shallow classifier to a suite of stylometric features such as punctuation usage. For the classifier, the Anecdotes use gradient boosted decision trees (Chen and Guestrin 2016), while the Dilemmas use logistic regression. The Length baseline picks multiple choice answers based on their length.

These baselines apply classifiers to bag-of-n-grams features, assessing the ability of lexical knowledge to solve the tasks. BinaryNB, MultiNB, and CompNB (Rennie et al. 2003) apply Bernoulli, Multinomial, and Complement Naive Bayes, while Logistic and Forest apply logistic regression and random forests (Breiman 2001).

Lastly, the deep baselines test how well existing methods solve Scruples. BERT (Devlin et al. 2019) and RoBERTa (Liu et al. 2019) fine-tune powerful pretrained language models on the tasks. In addition, we try both BERT and RoBERTa with the Dirichlet-multinomial likelihood (+ Dirichlet) as described in Section 3.2.

2 Training and Hyper-parameter Tuning

All models were tuned with Bayesian optimization using scikit-optimize (Head et al. 2018).

While the feature-agnostic models have no hyper-parameters, the other shallow models have parameters for feature-engineering, modeling, and optimization. These were tuned using 128 iterations of Gaussian process optimization with 8 points in a batch (Chevalier and Ginsbourger 2013), and evaluating each point via 4-fold cross validation. For the training and validation metrics, we used cross-entropy with hard labels. All shallow models are based on scikit-learn (Pedregosa et al. 2011) and trained on Google Cloud n1-standard-32 servers with 32 vCPUs and 120GB of memory. We tested these baselines by fitting them perfectly to an artificially easy, hand-crafted dataset. Shallow baselines for the Anecdotes took 19.6 hours using 32 processes, while the the Dilemmas took 1.4 hours.

Deep models’ hyper-parameters were tuned using Gaussian process optimization, with 32 iterations and evaluating points one at a time. For the optimization target, we used cross-entropy with soft labels, calibrated via temperature scaling (Guo et al. 2017). The training loss depends on the particular model. Each model trained on a single Titan V GPU using gradient accumulation to handle larger batch sizes. The model implementations built on top of PyTorch (Paszke et al. 2017) and transformers (Wolf et al. 2019).

Most machine learning models are poorly calibrated out-of-the-box. Since cross-entropy is our main metric, we calibrated each model on dev via temperature scaling (Guo et al. 2017), to compare models on an even footing. All dev and test results report calibrated scores.

3 Results

Following our goal to model norms’ distribution, we compare models with cross-entropy. RoBERTa with a Dirichlet likelihood (RoBERTa + Dirichlet) outperforms all other models on both the Anecdotes and the Dilemmas. One explanation is that unlike a traditional softmax layer trained on hard labels, the Dirichlet likelihood leverages all annotations without the need for a majority vote. Similarly, it can separate the controversiality of the question from the model’s uncertainty, making the predictions more expressive (see Section 3.2). Tables 7 and 8 report the results. You can demo the model at https://norms.apps.allenai.org.

Label-only and stylistic baselines do poorly on both the Dilemmas and Anecdotes, scoring well below human and Best performance. Shallow baselines also perform poorly on the Anecdotes; however, the bag of n-grams logistic ranker (Logistic) learns some aspects of the Dilemmas task. Differences between shallow models’ performance on the Anecdotes versus the Dilemmas likely come from the role of lexical knowledge in each task. The Anecdotes consists of complex anecdotes: participants take multiple actions with various contingencies to justify them. In contrast, the Dilemmas are short with little narrative structure, so lexical knowledge can play a larger role.

Analysis

Diving deeper, we conduct two analyses: a controlled experiment comparing different likelihoods for distributional labels, and a lexical analysis exploring the Dilemmas.

Unlike typical setups, Dirichlet-multinomial layers use the full annotations, beyond just majority vote. This distinction should especially help more ambiguous tasks like ethical understanding. With this insight in mind, we explore other likelihoods leveraging this richer information and conduct a controlled experiment to test whether training on the full annotations outperforms majority vote.

In particular, we compare with cross-entropy on averaged labels (Soft) and label counts (Counts) (essentially treating each annotation as an example). Both capture response variability; though, Counts weighs heavily annotated examples higher. On the Anecdotes, where some examples have thousands more annotations than others, this difference is substantial. For datasets like the Dilemmas, with fixed annotations per example, the likelihoods are equivalent.

Tables 9 and 10 compare likelihoods on the Anecdotes and Dilemmas, respectively. Except for Counts on the Anecdotes, likelihoods using all annotations consistently outperform majority vote training in terms of cross-entropy. Comparing Counts with Soft suggests that its poor performance may come from its uneven weighting of examples. Dirichlet and Soft perform comparably; though, Soft does better on the less informative, hard metric (F1). Like Counts, Dirichlet weighs heavily annotated examples higher; so, re-weighting them more evenly may improve its score.

2 The Role of Lexical Knowledge

While the Anecdotes have rich structure—with many actors under diverse conditions—the Dilemmas are short and simple by design: each depicts one act with relevant context.

To sketch out the Dilemmas’ structure, we extracted each action’s root verb with a dependency parser.en_core_web_sm from spaCy: https://spacy.io/. Overall, the training set contains 1520 unique verbs with “wanting” (14%), “telling” (7%), and “being” (5%) most common. To identify root verbs significantly associated with either class (more and less ethical), we ran a two-tailed permutation test with a Holm-Bonferroni correction for multiple testing (Holm 1979). For each word, the likelihood ratio of the classes served as the test statistic:

We tested association at the 0.050.05 level of significance using 100,000 samples in our Monte Carlo estimates for the permutation distribution. Table 13 presents the verbs most significantly associated with each class, ordered by likelihood ratio. While some evoke normative tones (“lying”), many do not (“causing”). The most common verb, “wanting”, is neither positive nor negative; and, while it leans towards more ethical, this still happens less than 60% of the time. Thus, while strong verbs, like “ruining”, may determine the label, in many cases additional context plays a major role.

To investigate the aspects of daily life addressed by the Dilemmas, we extracted 5 topics from the actions via Latent Dirichlet Allocation (Blei, Ng, and Jordan 2003), using the implementation in scikit-learn (Pedregosa et al. 2011). The main hyper-parameter was the number of topics, which we tuned manually on the Dilemmas’ dev set. Table 12 shows the top 5 words from each of the five topics. Interpersonal relationships feature heavily, whether familial or romantic. Less apparent from Table 12, other topics like retail and work interactions are also addressed.

Related Work

From science to science-fiction, people have long acknowledged the need to align AI with human interests. Early on, computing pioneer I.J. Good raised the possibility of an “intelligence explosion” and the great benefits, as well as dangers, it could pose (Good 1966). Many researchers have since cautioned about super-intelligence and the need for AI to understand ethics (Vinge 1993; Weld and Etzioni 1994; Yudkowsky 2008), with several groups proposing research priorities for building safe and friendly intelligent systems (Russell, Dewey, and Tegmark 2015; Amodei et al. 2016).

Beyond AI safety, machine ethics studies how machines can understand and implement ethical behavior (Waldrop 1987; Anderson and Anderson 2011). While many acknowledge the need for machine ethics, few existing systems understand human values, and the field remains fragmented and interdisciplinary. Nonetheless, researchers have proposed many promising approaches (Yu et al. 2018).

Efforts principally divide into top-down and bottom-up approaches (Wallach and Allen 2009). Top-down approaches have designers explicitly define ethical behavior. In contrast, bottom-up approaches learn morality from interactions or examples. Often, top-down approaches use symbolic methods such as logical AI (Bringsjord, Arkoudas, and Bello 2006), or preference learning and constraint programming (Rossi 2016; Rossi and Mattei 2019). Bottom-up approaches typically rely upon supervised learning, or reinforcement and inverse reinforcement learning (Abel, MacGlashan, and Littman 2016; Wu and Lin 2018; Balakrishnan et al. 2019). Beyond the top-down bottom-up distinction, approaches may also be divided into descriptive vs. normative. Komuda, Rzepka, and Araki (2013) compare the two and argue that descriptive approaches may hold more immediate practical value.

In NLP, fewer works address general ethical understanding, instead focusing on narrower domains like hate speech detection (Schmidt and Wiegand 2017) or fairness and bias (Bolukbasi et al. 2016). Still, some efforts tackle it more generally. One body of work draws on moral foundations theory (Haidt and Joseph 2004; Haidt 2012), a psychological theory explaining ethical differences in terms of how people weigh a small set of moral foundations (e.g. care/harm, fairness/cheating, etc.). Researchers have developed models to predict the foundations expressed by social media posts using lexicons (Araque, Gatti, and Kalimeri 2019), as well as to perform supervised moral sentiment analysis from annotated twitter data (Hoover et al. 2020).

Moving from theory-driven to data-driven approaches, other works found that word vectors and neural language representations encode commonsense notions of normative behavior (Jentzsch et al. 2019; Schramowski et al. 2019). Lastly, Frazier et al. (2020) utilize a long-running children’s comic, Goofus & Gallant, to create a corpus of 1,387 correct and incorrect responses to various situations. They report models’ abilities to classify the responses, and explore transfer to two other corpora they construct.

In contrast to prior work, we emphasize the task of building models that can predict the ethical reactions of their communities, applied to real-life scenarios.

Conclusion

We introduce a new task: Who’s in the Wrong?, and a dataset, Scruples, to study it. Scruples provides simple ethical dilemmas that enable models to learn to reproduce basic ethical judgments as well as complex anecdotes that challenge existing models. With Dirichlet-multinomial layers fully utilizing all annotations, rather than just the majority vote, we’re able to improve the performance of current techniques. Additionally, these layers separate model uncertainty from norms’ controversiality. Finally, to provide a better target for models, we introduce a new, general estimator for the best score given a metric on a classification dataset. We call this value the Best performance.

Normative understanding remains an important, unsolved problem in natural language processing and AI in general. We hope our datasets, modeling, and methodological contributions can serve as a jumping off point for future work.

Acknowledgements

We would like to thank the anonymous reviewers for their valuable feedback. In addition, we thank Mark Neumann, Maxwell Forbes, Hannah Rashkin, Doug Downey, and Oren Etzioni for their helpful feedback and suggestions while we developed this work. This research was supported in part by NSF (IIS-1524371), the National Science Foundation Graduate Research Fellowship under Grant No. DGE 1256082, DARPA CwC through ARO (W911NF15-1- 0543), DARPA MCS program through NIWC Pacific (N66001-19-2-4031), and the Allen Institute for AI.

Ethics Statement

Ethical understanding in NLP, and machine ethics more generally, are critical to the long-term success of beneficial AI. Our work encourages developing machines that anticipate how communities view something ethically, a major step forward for current practice. That said, one community’s norms may be inappropriate when applied to another. We urge practitioners to consider the norms of their users’ communities as well as the consequences and appropriateness of any model or dataset before deploying it. The code, models, and data in this work engage in an active area of research, and should not be deployed without careful evaluation. We hope our work contributes towards developing robust and reliable ethical understanding in machines.

References

Appendix A Estimating Oracle Performance

As discussed in Section 3.1, given class label counts, YiY_{i}, for each instance, we can model them using a Dirichlet-Multinomial distribution:

where NiN_{i} is the fixed (or random but independent) number of annotations for example ii.

First, we estimate α\alpha by minimizing the Dirichlet-multinomial’s negative log-likelihood, marginalizing out the θi\theta_{i}’s:

Pushing the log⁡\log inside the products leaves only log-gamma terms. For implementation, calling a log-gamma function rather than a gamma function is important to avoid overflow.

Given our estimate of α\alpha, α^\hat{\alpha}, we compute the posterior for each example’s true label distribution using the fact that the Dirichlet and multinomial distributions are conjugate:

Then, we can sample a set of class probabilities from this posterior:

Finally, we can use θ^i\hat{\theta}_{i} as the prediction for example ii. Repeating this process many times (2,000–10,000) and averaging the results yields the Best performance, estimating the oracle’s performance on the evaluation data.

It’s worth noting that this procedure will work with most metrics or loss functions, as long as the dataset has multiple class annotations per example. The main assumptions are that the annotations are independent, the true label distributions are roughly Dirichlet distributed, and the number of annotations is independent from the labels.

Appendix B Dataset Construction

Section 2 presents the Anecdotes and Dilemmas. This appendix further describes the Anecdotes’ construction.

The AITA subreddit is an ideal source for studying communal norms. On it, people post stories and comment with who they view as in the wrong. Figure 5 offers a screenshot of the interface. While users can view others’ comments before responding, the comments only appear after the response form on the page.

B.2 Extraction

Section 2.1 evaluates and describes how we filter and extract information from posts and comments at a high level. The following paragraphs detail each component’s extraction.

Titles come directly from the posts’ titles. Generally, they summarize the posts; however, some users editorialize or choose humorous titles.

The subreddit requires titles to begin with a tag categorizing the post (“AITA”, “WIBTA”, or “META”).Respectively: “am I the a-hole”, “would I be the a-hole”, and “meta-post” (about the subreddit). Using regexes, we match these tags, convert AITA to historical, WIBTA to hypothetical, and discard all META posts.

The stories’ titles often summarize the main thing the author did. So, we extracted an action from each using a sequence of rules-based text normalizers, filters, and regular expressions. Results were then transformed into gerund phrases (e.g. “not offering to pick up my friend”). 94.4% of stories have successfully extracted actions.

Posts have an attribute providing the text; however, users can edit their posts in response to comments. The subreddit preserves stories as submitted, using a bot that comments with the original text. To avoid data leakage, we use this original text and discard posts when it cannot be found.

The scores tally community members who expressed a given label. To improve relevance and independence, we only considered comments replying directly to the post (i.e., top-level comments). We extracted labels using regexes to match variants of initialisms used on the site (like \m(?i:ESH)\M) and textual expressions that correspond to them (i.e. (?i:every(?:one |body) sucks here){e<=1}). For comments expressing multiple labels, the first label was chosen unless a word between the two signified a change in attitude (e.g., but, however).

Appendix C Baselines

In Section 4, we evaluate a number of baselines on the Anecdotes and the Dilemmas. This appendix describes each of those baselines in more detail.

The feature-agnostic models predict solely from the label distribution, without using the features.

predicts label probabilities according to the label distribution.

assigns all probability to a random label from the label distribution.

C.2 Stylistic Baselines

These models probe if stylistic artifacts like length or lexical diversity give away the answer.

picks multiple choice answers based on their length. Reported results correspond to the best combination of shortest or longest and word or character length (i.e. fewest / characters).

applies classifiers to a suite of stylometric features. The features are document length (in tokens and sentences), the min, max, mean, median, and standard deviation of sentence length (in tokens), document lexical diversity (type-token ratio), average sentence lexical diversity (type-token ratio), average token length in characters (excluding punctuation), punctuation usage (counts per sentence), and part-of-speech usage (counts per sentence). The Anecdotes use gradient boosted decision trees (Chen and Guestrin 2016); while the Dilemmas use logistic regression on the difference of the choices’ scores.

C.3 Lexical and N-Gram Baselines

These baselines measure to what degree shallow lexical knowledge can solve the tasks.

apply bernoulli and multinomial naive bayes to bag of n-grams features from the title concatenated with the text. Hyper-parameter tuning considers both character and word n-grams. Complement naive bayes (CompNB) classifies documents based on how poorly they fit the complement of the class (Rennie et al. 2003). This often helps class imbalance.

scores answers with logistic regression on bag of n-grams features. Hyper-parameter tuning decides between tf-idf features, word, and character n-grams. The Anecdotes’ linear model considers both one-versus-rest and multinomial loss schemes; while the Dilemmas’ model uses the difference of the choices’ scores as the logit for whether the second answer is correct.

trains a random forest on bag of n-grams features (Breiman 2001). Hyper-parameter tuning tries tf-idf features, pure counts, and binary indicators for vectorizing the n-grams.

C.4 Deep Baselines

These baselines test whether current deep neural network methods can solve Scruples.

achieves high performance across a broad range of tasks (Devlin et al. 2019). BERT pretrains its weights with masked language modeling, a task that predicts masked out tokens from the input. The model adapts to new tasks by fine-tuning problem-specific heads end-to-end along with all pretrained weights.

improves upon BERT with better hyper-parameter tuning, more pretraining, and by removing certain model components (Liu et al. 2019).

uses a Dirichlet (Gelman et al. 2003) likelihood in the last layer instead of the traditional softmax. The Dirichlet likelihood allows the model to leverage all the annotations, instead of the majority label, and to separate a question’s controversiality from the model’s uncertainty. Section 3.2 discusses the model in more detail.

C.5 Alternative Likelihoods

In addition to the deep baselines, Section 5 explores alternative likelihoods that, like the Dirichlet-multinomial layer, leverage all of the annotations rather than training on the majority vote.

uses a softmax in the last layer with a categorical likelihood, similarly to standard training setups; however, instead of computing cross-entropy against the (hard) majority vote label, it computes cross-entropy against the (soft) average label from the annotations. Thus, the loss becomes:

uses a softmax in the last layer with a categorical likelihood, similarly to standard setups and the soft labels baseline; however, the loss treats each annotation as its own example or, equivalently, uses unnormalized counts:

Again, using the notation from Section 3.1. This loss is the same as maximum likelihood estimation on the annotations.

Appendix D The Role of Lexical Knowledge

This appendix gives more details on the two analyses presented in Section 5.2. The first measured association between root verbs and their actions’ labels. The second extracted topics describing the Dilemmas.

The first analysis used the likelihood ratio:

to measure association between root verbs and whether their actions were judged as less ethical. Using a two-tailed permutation test with a Holm-Bonferroni correction for multiple testing (Holm 1979), we selected only the verbs with statistically significant associations at the 0.050.05 level. Even using 100,000 samples in our Monte Carlo estimates of the permutation distribution, some p-values were computed as zero due to the likelihood ratio in the original data being higher than any of the sampled permuations. Since the Holm-Bonferroni correction doesn’t account for this approximation error, the final p-values remain zero even though they would be larger if we’d used more samples; however, each zero would still be below the next smallest p-value from the test. Table 13 presents all of the verbs significantly associated with each class, ordered by likelihood ratio.

Table 14 shows the top 25 words from each of the five topics.

Appendix E Latent Trait Analysis

Formalizing norms as classifying right versus wrong behavior has the advantages of simplicity, intuitiveness, ease of annotation, and ease of modeling. In reality, norms are much more complex. Often, decisions navigate conflicting concerns, and reasonable people disagree about the right choice. Given the goal of reproducing these judgments, it’s worth asking whether binary labels provide a sufficiently precise representation for the phenomenon.

Motivated by moral foundations theory, a psychological theory positing that people make moral decisions by weighing a small number of moral foundations (e.g. care/harm or fairness/cheating) (Haidt and Joseph 2004; Haidt 2012), we explored the degree to which binary labels reduce a more complex set of concerns. To investigate this hypothesis empirically on the Dilemmas, we conducted an exploratory latent trait analysis.A good introduction may be found in chapter 8, “Factor Analysis for Binary Data”, of Galbraith et al. (2002). Latent trait analysis models the dependence between categorical variables by approximating the data distribution with a linear latent variable model. Concretely, the model represents the distribution as logistic regression on a Gaussian latent variable. In other words, if YY is the vector of binary responses from a single annotator, then:

The parameters, W\mathbf{W} and b\mathbf{b}, are then fitted via maximum likelihood estimation, marginalizing out ZZ.

often measures the goodness-of-fit, where fθf_{\theta} is the probability density function and θs\theta_{s} is the parameters for the saturated model, i.e. the model that can attain the best possible fit to the data. In our case, the saturated model assumes full dependence and assigns to each possible vector of responses XX the frequency with which it was observed in the data. We can then use the percentage of deviance in the null (independent) model explained by the current model under consideration to assess goodness of fit:

Where θ0\theta_{0} is the fully independent model.

For this analysis, we created a densely annotated set of questions by randomly sampling 20 from the benchmark’s development set and crowdsourcing labels for them from 1000 additional annotators. Figure 8 plots the deviance for models with varying numbers of traits. The high degree of unexplained deviation in the models suggests that deviation between annotators’ responses is not explained well by the linear model.

In addition to goodness-of-fit statistics, we can use the two dimensional latent trait model to visualize the data. Figure 7 plots the questions based on their weights, while Figure 7 shows the responses projected into the latent space.

Summarizing the analysis, the linear latent variable model was not able to explain very much of the deviance beyond the independent model. One challenge in trying to explain annotator’s decisions is the fact that the Dilemmas randomly pairs items together to form moral dilemmas. Anecdotally, this random pairing makes comparisons more difficult since actions may come from unrelated contexts, and often neither is clearly worse. Future work may want to investigate annotating more nuanced information about the actions or how annotators arrive at their conclusions.

Appendix F Hyper-parameters

Section 4.2 describes our overall training, hyper-parameter tuning, and evaluation methodology. This appendix details the hyper-parameter spaces searched and final hyper-parameters used for each baseline. Full code for each baseline, including the hyper-parameter search spaces, is available at https://github.com/allenai/scruples.

In addition to the information in Section 4.2, all servers used to run experiments had Ubuntu 18.04 for the operating system. The server used to train the deep models had 32 Intel(R) Xeon(R) Silver 4110 CPUs and 126GB of memory. Titan V GPUs have 12GB of memory.

F.1 Feature-agnostic Baselines

Both the Prior and Sample baselines have no hyper-parameters.

F.2 Stylistic Baselines

searched for the best combination of shortest or longest sequence and words or characters as the measure of length. The reported baseline chooses options with the fewest characters.

used a gradient boosted decision tree (Chen and Guestrin 2016) on the Anecdotes and searched the following hyper-parameters: a max-depth of 1 to 10, log-uniform learning rate from 1e-4 to 1e1, log-uniform gamma from 1e-15 to 1e-1, log-uniform minimum child weight from 1e-1 to 1e2, uniform subsampling from 0.1 to 1, uniform column sampling by tree from 0.1 to 1, log-uniform alpha from 1e-5 to 1e1, log-uniform lambda from 1e-5 to 1e1, uniform scale positive weight from 0.1 to 10, and a uniform base score from 0 to 1. For the Dilemmas, Style used logistic regression searching over a log-uniform C from 1e-6 to 1e2, and a class weight of either "balanced" or None. The hyper-parameters for the best model on the Anecdotes were a max-depth of 2, learning rate of 0.1072, gamma of 1.726e-7, minimum child weight of 63.05, subsampling of 0.6515, column sampling by tree of 0.8738, alpha of 4.902e-4, lambda of 6.750e-5, scale positive weight of 2.544, and a base score of 0.3413. The hyper-parameters for the best model on the Dilemmas were a C of 0.1714 and a class weight of "balanced".

F.3 Lexical and N-Gram Baselines

had the following hyper-parameters for vectorizing the text: strip accents of either "ascii", "unicode", or None, lowercase of either True or False, stop words of either "english" or None, an n-gram range from 1 to 2 for the lower bound and up to 5 more for the upper bound, analyzer of either "word", "char", or "char_wb" (characters with word boundaries), uniform maximum document frequency from 0.75 to 1, and a uniform minimum document frequency from 0 to 0.25. All three naive Bayes baselines additionally searched alpha (additive smoothing) uniformly from 0 to 5, while CompNB also tried a norm (normalization) of either True or False. The final hyper-parameters for BinaryNB were a strip accents of "unicode", lowercase of True, stop words of None, n-gram range of 1 to 1, analyzer of "char_wb", maximum document frequency of 0.9298, minimum document frequency of 0.25, and an alpha of 5. For MultiNB, the final hyper-parameters were a strip accents of "ascii", lowercase of True, stop words of "english", n-gram range of 1 to 3, analyzer of "word", maximum document frequency of 1, minimum document frequency of 0.25, and an alpha of 5. Lastly, the best CompNB model had for its hyper-parameters a strip accents of "ascii", lowercase of False, stop words of "english", n-gram range of 1 to 4, analyzer of "word", maximum document frequency of 1, minimum document frequency of 0.1595, alpha of 0, and a norm of False.

used the following hyper-paramter search space: the same text vectorization and tf-idf featurization hyper-parameters as the Logistic model on the Anecdotes, and for the random forest classifier a splitting criterion of either "gini" or "entropy", minimum samples for splitting between 2 and 500, minimum samples per leaf from 1 to 250, uniform minimum weight fraction per leaf from 0 to 0.25, either with bootstrap resampling or without, and a class weight of either "balanced", "balanced_subsample", or None. The best Forest model had the following hyper-parameters: strip accents of "ascii", lowercase of True, stop words of None, n-gram range of 1 to 3, analyzer of "word", maximum document frequency of 1, minimum document frequency of 0.25, binary of False (i.e., it used counts), tfidf normalization of "l2", use idf of False, sublinear term frequency of False, splitting criterion of "entropy", minimum samples for splitting of 230, minimum samples per leaf of 1, minimum weight fraction per leaf of 0, bootstrap of False (i.e., it did not use bootstrap resampling), and a class weight of None.

F.4 Deep Baselines

searched a log-uniform learning rate from 1e-8 to 1e-2, log-uniform weight decay from 1e-5 to 1e0, uniform warmup proportion from 0 to 1, and a training batch size from 8 to 1024. For the Anecdotes, the number of epochs was explored from 1 to 10, and the sequence length was 512. For the Dilemmas, the number of epochs was explored from 1 to 25, and the sequence length was 92. On the Anecdotes, the best BERT model had a learning rate of 9.113e-7, weight decay of 1, warmup proportion of 0, batch size of 8, and 10 epochs. On the Dilemmas, the best BERT model had a learning rate of 7.615e-6, weight decay of 0.3570, warmup proportion of 0.1030, batch size of 32, and 7 epochs.

searched the same space of hyper-parameters as BERT, except that the sequence length for the Dilemmas was 90 rather than 92. On the Anecdotes, the best RoBERTa model had a learning rate of 3.008e-6, weight decay of 1.0, warmup proportion of 9.762e-2, batch size of 8, and 10 epochs. On the Dilemmas, the best RoBERTa model had a learning rate of 1.130e-6, weight decay of 1, warmup proportion of 0.2193, batch size of 16, and 25 epochs.

requires no additional hyper-parameters beyond the base neural network. On the Anecdotes, the best BERT + Dirichlet model had a learning rate of 4.403e-6, weight decay of 1, warmup proportion of 0, batch size of 8, and 4 epochs and the best RoBERTa + Dirichlet model had a learning rate of 5.147e-5, weight decay of 4.341e-3, warmup proportion of 0.3269, batch size of 64, and 2 epochs. On the Dilemmas, the best BERT + Dirichlet model had a learning rate of 8.202e-5, weight decay of 1e-5, warmup proportion of 1, batch size of 128, and 11 epochs and the best RoBERTa + Dirichlet model had a learning rate of 7.132e-6, weight decay of 2.009e-3, warmup proportion of 0.2113, batch size of 8, and 7 epochs.

F.5 Alternative Likelihoods

has no hyper-parameters beyond those of the base neural network. On the Anecdotes, the best BERT + Soft model had a learning rate of 7.063e-6, weight decay of 1e-5, warmup proportion of 0, batch size of 8, and 10 epochs and the best RoBERTa + Soft model had a learning rate of 2.408e-5, weight decay of 2.517e-5, warmup proportion of 0.3215, batch size of 256, and 8 epochs. On the Dilemmas, the best BERT + Soft model had a learning rate of 4.637e-5, weight decay of 4.139e-5, warmup proportion of 0.4407, batch size of 128, and 22 epochs and the best RoBERTa + Soft model had a learning rate of 4.279e-6, weight decay of 1e-5, warmup proportion of 0, batch size of 8, and 22 epochs.

requires no hyper-parameters beyond those of the base neural network. On the Anecdotes, the best BERT + Counts model had a learning rate of 7.953e-6, weight decay of 1.461e-4, warmup proportion of 0.2682, batch size of 32, and 10 epochs and the best RoBERTa + Counts model had a learning rate of 4.353e-6, weight decay of 1e-5, warmup proportion of 1, batch size of 8, and 10 epochs. On the Dilemmas, the best BERT + Counts model had a learning rate of 6.099e-6, weight decay of 1e-5, warmup proportion of 0.2659, batch size of 8, and 20 epochs and the best RoBERTa + Counts model had a learning rate of 4.585e-6, weight decay of 1e-5, warmup proportion of 0.5714, batch size of 8, and 25 epochs.

Appendix G Examples

This appendix provides examples from the Anecdotes and Dilemmas. Some examples illustrate the diversity, variety, and specificity found in these real-world anecdotes, while others highlight how the ethical judgments can be divisive.

The examples often touch on sensitive, sometimes troubling, topics. We present these examples to illustrate the complexity and distributional nature inherent in predicting a community’s ethical judgments.