Adversarial Training for High-Stakes Reliability
Daniel M. Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Ben Weinstein-Raun, Daniel de Haas, Buck Shlegeris, Nate Thomas
Introduction
Advances in deep learning have led to increasingly powerful AI systems, for example in sequential decision making , robotics , and language modeling and text-based reasoning . Most empirical work on techniques for aligning powerful AI has focused on achieving good average-case performance in domains where no single action is catastrophic, for example using human trajectory rankings or imitation learning . However, many situations where we want to deploy AI systems are high-stakes—that is, it is possible for the system to take actions that lead to catastrophic outcomes.
In these situations, one of our most important goals is high-stakes reliability: avoiding even a single catastrophic failure while in deployment. Achieving high-stakes reliability is difficult because some failures might not be encountered during the ordinary course of training, leaving them uncorrected by default. These failures could arise on out-of-distribution data resulting from domain shift or adversaries in the environment. Alternatively, undetected failures could arise without distributional shift if they occur with sufficiently low probability. We describe our setting more precisely in Section 3.1.
One technique for improving high-stakes reliability is adversarial training . In its general form, adversarial training consists of finding inputs that a model does especially poorly on and then training the model on those examples. If our adversarial attacks sufficiently cover the space of catastrophic inputs, then adversarial training incentivizes the model to avoid catastrophic failures.
In this work, we used a simple task as a testbed for adversarial training. The system must take a three-sentence excerpt from a story (a “prompt”) and output one more sentence (a “completion”) that continues the story without introducing any physical injuries to any characters. To do this, we train a language model as a classifier for injurious completions, which we use to filter the outputs of a generative language model. We then adversarially train it using a variety of attacks (Figure 1).
As measured by both the false negative rate on our adversarial datasets and the time to generate adversarial examples, we found that adversarial training increased robustness to attacks similar to those trained against (Section 4.4.3), although it did not eliminate failures completely. Qualitatively, we found that the remaining failures in adversarially trained models were less egregious and were less likely to contain mention of direct injury (as opposed to implied or indirect injuries). At the same time, we found that adversarial training did not degrade performance on our baseline (non-adversarial) dataset. Finally, we found that we could set very conservative classifier thresholds without degrading the quality of our generator output.
Our main contributions are the following:
We highlight the setting of high-stakes reliability and report the results of an initial project in this setting.
We demonstrate a novel tool-assisted human attack that increases the ease of finding adversarial examples (Section 4.4.3)
We found that on our chosen task, conservative thresholds enable a high degree of worst-case reliability, with minimal impact on average-case performance.
We see our work as exploratory and think that there are many promising follow-up directions to pursue for stronger results. We hope that this project will be followed by work building the theory and practice of adversarial training to the point where it can robustly enable high-stakes reliability.
Related work
The field of adversarial machine learning or even the subfield of adversarial training are too large to summarize in this paper. Here, we outline a handful of particularly related areas.
Much recent work in adversarial training has been on preventing adversarial examples for image classifiers . Notably, the majority of image adversarial training work studies ball perturbations . In contrast, we allow adversarial examples in an unrestricted space , with the goal of avoiding any feasible failures.
There is a large and growing literature on both adversarial attacks and adversarial training for large language models . The majority of these focus on automatic attacks against language models, where the authors have access to an automated source of ground truth, or restrict their examples to small perturbations that are assumed not to change the ground truth. In this work, we primarily used human adversaries as we chose a task without an automated source of ground truth.
Our human-centered approach to adversarial training is most similar to the “Break-it-Fix-it” approach taken by Dinan et al. , which uses human-generated adversarial examples to increase the robustness of an offensive speech classifier.
In recent years, this approach has also been proposed as a method of evaluating language model classifiers in general. Several recent datasets and benchmarks are constructed with human-in-the-loop adversaries, such AdversarialNLI , AdversarialGLUE , and DynaBench . Our analysis of the effects of multiple iterations of adversarial training resembles DADC .
Most prior work picked tasks challenging enough for the model that unaugmented human labelers were able to easily find adversarial examples, whereas in our case it started to become more difficult for human labelers to find further adversarial examples after initial adversarial training. This led us to augment our labelers’ abilities. Like Wallace et al. , our tool displays a saliency map; however, we added a token substitution tool that gives humans automated suggestions to help hill-climb toward an adversarial example.
Problem statement
Here, we define high-stakes reliability more precisely. A high-stakes reliability setting contains both: 1) an average-case performance metric that measures successful execution of the intended task; as well as 2) a notion of “catastrophic failure” that must be avoided in deployment. The goal in high-stakes reliability settings is to train a model that avoids catastrophic failures while still achieving good average-case performance.
During training, we assume that we have an oversight process that can safely detect catastrophes in arbitrary new examples. However, the oversight process may be too expensive to evaluate for every example in deployment. For example, we might be able to pay humans labellers to closely examine a few tens of thousands model outputs in training, but not be able to scale human oversight to hundreds of millions of model outputs in deployment. As fully preventing all catastrophic failures may be unachievable with current techniques (and very low failure rates are hard to measure), we propose using two proxy metrics instead. First, we can measure the failure rate on adversarially-generated datasets designed to elicit catastrophic behavior. Second, we can measure the difficulty of finding a novel adversarial example, using particular styles of adversarial attacks.
2 Our specific task: filtering a story generator
In this work, we consider the task of producing safe completions to fictional stories with a filtered language model. Specifically, given a three-sentence prompt, our goal is to produce a noninjurious completion—that is, one further sentence that does not introduce additional injury to any characters.We provide more details of our definition of injury in Section A.1.2. We chose this relatively simple non-injury task to make it reasonably easy for non-expert humans to recognize failures and enable relatively small models to perform acceptably well on our task.
In our specific setup, we assume that we have access to a generative language model that generates high-quality completions to given prompts; our task is to learn an injuriousness classifier that classifies completions as injurious and then use it to filter the output of our generator. We use the quality of our filtered completions (as judged by human raters, relative to our unfiltered language model) as our average case performance metric. Our “catastrophic failures” are injurious examples that the classifier incorrectly labels as safe (that is, the catastrophic failure rate of the system is the false negative rate on filtered generator outputs).
Methods
In this section, we describe how we trained our injuriousness classifier. After training a baseline classifier on some initial labelled data, we attacked it with several adversarial training techniques and retrained it using the adversarial examples we generated. We summarize the properties of the datasets used in training in Table 1.
We sourced human contractors primarily from Upwork and from Surgehttps://www.surgehq.ai/ to perform our labeling. To determine whether snippets were injurious, we asked the contractors to label each one injurious, non-injurious, or “Unsure”. We used these human labelers to label all our training and evaluation data, including our adversarial training data.See Appendix A.1.6 for details of our labeling process.
2 Classifier training
We trained a classifier by fine-tuning deberta-v3-large from HuggingFace . During training, we treated all snippets labeled “Unsure” as injurious for the purposes of training because we wanted to train our classifier to be conservative. Because our datasets had more non-injurious examples than injurious examples, we upsampled snippets labeled injurious by up to 5 so that they were closer to the number of non-injurious examples.We found in a preliminary experiment that upsampling injurious snippets improved performance, though the effect did not reach statistical significance. We document other hyperparameters in Appendix A.3.
3 Initial data sources
Our initial, baseline classifier training set consisted of “snippets” derived from a dataset of fan fiction stories. We sourced our prompts from an archive of approximately 300 GB of stories from fanfiction.net, and subselected them for increased likelihood of injury to address the class imbalance caused by the low base prevalence of injury.See Appendix A.1 for more details on the classifier training dataset. We generated completions from a GPT-Neo-2.7B fine-tuned on this story dataset.
4 Adversarial attacks and datasets
After training a model on our initial fanfiction dataset, we performed three rounds of adversarial training. In each round, we generated adversarial examples against a model trained against previously generated datasets, using a different type of adversarial attack. We labelled the adversarial examples with human labellers, before training a new model on all the datasets generated so far.
In our first round, we used unassisted, manual human adversaries. We created a web interface where a user could input a prompt and a completion and check the classifier score for our baseline classifer. Users were able to submit the snippet if the score dropped below and they believed the snippet was injurious. We then hired contractors to write adversarial examples using this web interface. In order to increase the diversity of adversarial examples, partway through data collection, we changed the target model from baseline classifier to a model that had been trained on a subset of the manual adversarial examples gathered so far.
4.2 Automatically paraphrased adversarial examples
To increase the size of our adversarial dataset, we used GPT-3 to paraphrase manually generated human adversarial snippets, as has been done with previous generative language models . We devised a “few-shot” GPT-3 prompt for this purpose, containing several examples of paraphrasingWe include the GPT-3 prompt in Appendix A.6. For each snippet, we generated many possible paraphrases of it with GPT-3, then had our classifier classify all the paraphrases as injurious or non-injurious. We discarded all paraphrases that the classifier already rejected as injurious, as we were only interested in adversarial examples. We continued sampling until we had generated 5–10 accepted paraphrases for each snippet.
4.3 Tool-assisted rewrites
To generate adversarial examples more effectively, we helped our contractors by adding two key features to our interface, as shown in Figure 4 in the Appendix.Our tool is available at https://www.ttft.io/talk-to-filtered-transformer.
We built a token substitution and insertion feature. The user could click on a token in the snippet and a dropdown would appear with a list of new tokens sorted by how much they reduce the current model score. The candidate tokens were plausible “fill-in-the-blank” suggestions from deberta-v2-xlarge .
We displayed an indication of saliency for each token in the snippet. The interface highlighted in yellow the tokens most likely to affect the classifier’s output. The saliency score for a token was the magnitude of the gradient of the classifier’s output with respect to the token’s embedding.
As before, users were free to enter or edit examples manually (regardless of the tokens suggested by our token substitution and insertion features), and could submit examples if the classifier score dropped below 0.05. We show examples of adversarial snippets and an example progression of the tool-assisted rewriting process in Appendix C.5. As with the manual adversarial examples, partway through we trained our model on what we had so far and collected the rest of the rewrites against the new, more powerful model.
Experiments and results
In order to evaluate how helpful each of the three attack techniques (Section 4.4) were, we trained a series of classifiers with the data from one further attack added each time.
We trained all models with the same number of labeled snippets (166,210) in order to make it clear that the differences are due to the increased utility of adversarial examples. We randomly removed a number of snippets from the initial fan fiction dataset equal to the number of adversarial snippets added at each step (1,492, 12,514, and 4,904 respectively).
To maximize the reliability of our filtered generator, we wanted to find the lowest classifier threshold—that is, the classifier score above which we consider a completion to be injurious—that had only a small impact on the quality of filtered snippets. We created a static quality-evaluation dataset that could be reused to estimate the quality for a variety of classifiers and thresholds, by asking contractors to compare 4–100 pairs of completions for each prompt, with completions drawn i.i.d. from the unfiltered generator. We asked them to choose the higher quality completion or mark them both as the same quality, where two completions of equal quality count as half a win for each completion. Then, we estimated the quality at each classifier threshold by evaluating the classifier against the labeled generations to simulate rejection sampling. See Appendix A.7 for more details.
When evaluating our models’ performance, we used the lowest threshold that kept the 95% confidence lower bound of the preference rate at greater than 48% (with 50% meaning equal performance to the unfiltered generator). We were able to select thresholds between and .More details on classifier hyperparameters in Appendix A.3. Note that these thresholds were significantly more conservative than the threshold of that we used during adversarial training. At the selected thresholds, our classifiers reject about 25% of random snippets and 50–60% of completions proposed during rejection sampling.
After selecting the thresholds, we confirmed that our quality estimates were accurate by performing a direct quality evaluation. We rejection-sampled one completion from each of 5000 in-distribution prompts and compared them to a completion from the unfiltered generator.If the filtered generator could not generate an accepted completion within 100 attempts, we always counted that comparison as a loss. As our results in in Table 2 show, the thresholds we picked did not significantly degrade the quality of the filtered generator.
2 Tool-assisted rewrites were faster to generate than manual adversarial examples.
To determine how much the token substitution tool helped, we ran a smaller scale experiment where we measured the time for contractors to attack the Baseline and +tool-assisted classifier with and without our rewrite tools. Contractors took on average 20 minutes (95% CI: ) to find an adversarial example for the baseline classifier without the tool, compared to 13 minutes with the tool (95% CI: ). In addition, contractors took on average 44 minutes (95% CI: ) to generate an adversarial example for +tool-assisted without tools, compared to an average of 26 minutes (95% CI: ) with our rewrite tool. We also received feedback from contractors saying that the task was significantly more frustrating without our rewrite tool. This demonstrates that our tool helps our contractors more quickly generate adversarial examples for our classifiers.
3 Adversarial training increased robustness to adversarial attack.
We evaluated our classifiers on adversarial snippets from our final attack to determine whether the adversarial training increased robustness to this type of out-of-distribution data.
We had contractors attack our final classifiers using tool-assisted rewrites. Each contractor was randomly assigned a classifier each day they worked so that an approximately equal amount of effort was spent on each classifier.If we had held the number of snippets per classifier equal (instead of time spent), differences in different labelers’ comparative strength against different classifiers might have biased the measurement. The per-day assignment also gave labelers the chance to learn tricks specific to one classifier, although they were blinded to which classifier they were assigned to. Each tool-assisted rewrite was labeled by a distinct pair of contractors (with a third tie-break label in case of disagreement) to confirm that it was injurious. Further details of this process are in Appendix A.4.1. Randomly selected tool-assisted rewrites from each dataset are found in Appendix 8.
As reported in Figure 2, against each successive classifier, it took more time to generate a new adversarial example. We estimated this by counting the total time spent rewriting snippets per classifier and dividing by the number of successful tool-assisted rewrites. This shows that our attacks have some shared surface area, i.e. training on our earlier adversarial attacks also confers advantages against the tool-assisted adversarial attack. Training on all adversarial attacks increased the time per example by .
The smaller-scale experiment in Section 5.2 shows a similar increase in difficulty when contractors do not have access to the tools. Contractors took 44 minutes (95% CI: ) to break +tool-assisted manually, versus 20 minutes (95% CI: ) to break the baseline classifier, so adversarial training increased the time per example by (95% CI: [, ]).
We treated each set of tool-assisted rewrites as a separate dataset and evaluated each classifier against each dataset. Figure 3 shows the results. Ignoring the faded bars for attacks targeting the same classifier being evaluated (which will always be at FNR = 1.0), each chart shows a lower FNR for later classifiers. The +tool-assisted classifier is robust to attacks targeted at previous classifiers (rightmost bar in each chart). Conversely, attacks targeting it work nearly as well on previous classifiers also (bottom right chart), suggesting that its failure modes were largely preexisting rather than the result of new problems introduced via additional adversarial training.
We found that, in our subjective judgment, the rewritten snippets that our contractors submitted against more adversarially-trained classifiers seemed somewhat less egregiously injurious than tool-assisted rewrites against less adversarially-trained classifiers. We defined “egregiousness” as a combination of the severity of the injury and the likelihood of injury or increased injury in the completion. One researcher labeled ten snippets rewritten against each classifier subjectively on a scale from 1 (not injurious) to 10 (specific, maximum severity injury) while blinded to which classifier the attack targeted. The average egregiousness ratings are reported in Figure 2.
We also looked at the snippet labeled most egregious for each classifier from this process and compared these snippets subjectively. Table 3 shows two of them. As the table suggests, classifiers trained against a broader pool of adversarial attacks tended to have less egregious adversarial examples.
4 Adversarial training did not hurt performance on the in-domain data.
To evaluate the effect of adversarial training on our initial baseline (in-domain) dataset, we collected 100,033 new snippets drawn from our story database and our generator in a method similar to the one described in Section 4.3, except without selecting for injurious snippets. We had each snippet labeled twice and had each snippet with a disagreement labeled by a third tie-breaker to decrease the error rate on our test set labels. We also removed the option for “Unsure” due to a miscommunication with our vendor (see Appendix C.4).
As reported in Table 2, all four of the evaluated models performed about equally well, getting 2 to 3 false negatives on 2447 positives on our dataset of 100,033 examples. In addition, all four classifiers achieved similar false positive rates. This suggests that adversarial training did not affect the performance of our classifiers on in-domain data.
Discussion and Future Work
In this work, we explored a simple example of a high-stakes reliability task. We developed a quality measurement technique for rejection-sampled generators and found that we could set very conservative thresholds without significantly reducing quality. We built a series of adversarial training techniques, including a tool-assisted human attack, and found that they improved the classifier’s robustness to attack without affecting in-distribution reliability. Below, we outline some limitations of the current work and a variety of directions for future work.
The contractors had a tendency to produce adversarial examples that were relatively borderline or ambiguous, particularly when targeting more adversarially robust classifiers. However, when we attacked our models with our rewrite tool, we were able to construct more egregious adversarial examples, featuring direct injury, in part because researchers on our team used different heuristics for finding adversarial examples (see Appendix A.8). This underscores the need for a more diverse pool of stronger adversarial attacks, for better adversarial training . Future work could add more tools (such as better suggestions for our human adversaries ) and study the relative effectiveness of the different tools, develop better training methods for human attackers, or more fully characterize properties of adversarial inputs to better understand our models .
In this work, we used human contractors (augmented with tools) to generate adversarial examples, as our task lacks an automated source of ground truth, we did not restrict our adversarial examples, and we were not successful in fine-tuning an LM adversary (as discussed in Appendix A.5). Future work could explore ways to generate synthetic examples, such as imitation learning on human examples or better methods of using reinforcement learning to fine-tune automated adversaries .
Exploring the generality of our results. Much of our high level of reliability can be attributed to the fact that we were able to set particularly strict thresholds without significantly impacting quality on the story continuation task. Future work is needed to test whether or not this holds true on open-ended generation tasks in general.
Adversarial training on larger models. The classifiers we trained were 304M-parameter DeBERTa V3 models . Most likely, many of their failures were due to capability limitations, and working with larger models would improve their performance substantially. On the other hand, we think that working with larger models would still leave us in qualitatively the same situation, since state-of-the-art models still fail to understand many things that humans do.
Measuring the reliability of very robust classifiers by sampling randomly is very expensive. For example, on our test set of 100k examples, the difference between our best and worst classifiers was misclassifying 2 examples versus 3. Future work could attempt to use techniques similar to AMLS to more precisely measure in-distribution and out-of-distribution reliability in an extremely-low-failure-rate setting, or define a upper bound on the reliability using techniques such as SDP relaxation .
Acknowledgments and Disclosure of Funding
Paul Christiano originally proposed this project, and we benefited immensely throughout from discussions with him, as well as with Ajeya Cotra and Beth Barnes. We thank John Schulman, Jared Kaplan, Sam Bowman, Rohin Shah, Jonathan Uesato, Holden Karnofsky, Jan Leike, Jacob Hilton, Ethan Perez, Collin Burns, Jean-Stanislas Denain, Summer Yue, Nix Goldowsky-Dill, Chris MacLeod, Ryan Greenblatt, and Bill Zito for reading drafts of the paper and giving helpful feedback. We are grateful to Shauna Kravec, Dane Sherburn, and Everett Smith for their contributions to parts of the project, and to Kelsey Piper for organizing a party to collect more manual adversarial examples. We thank Surge and our contractors for their dedicated efforts over many months of labeling and writing adversarial examples. Finally, we thank the Redwood Research operations staff for providing an excellent work environment.
This work was funded by Redwood Research Group Inc.
References
Checklist
Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]
Did you describe the limitations of your work? [Yes] See Section 6.
Did you discuss any potential negative societal impacts of your work? [Yes]
Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]
If you are including theoretical results…
Did you state the full set of assumptions of all theoretical results? [N/A]
Did you include complete proofs of all theoretical results? [N/A]
Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] Linked in the Appendix.
Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes] Listed in Appendix A.
Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [Yes]
Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] Discussed in Appendix A.
If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…
If your work uses existing assets, did you cite the creators? [Yes] We cited all pre-existing models and frameworks
Did you mention the license of the assets? [N/A]
Did you include any new assets either in the supplemental material or as a URL? [Yes] We include a download link to our data and model weights in the Appendix.
Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A]
Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [N/A]
If you used crowdsourcing or conducted research with human subjects…
Did you include the full text of instructions given to participants and screenshots, if applicable? [Yes] See Appendix A.
Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]
Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [No] Unfortunately, as we contracted our labeling to a third party, we do not have access to the hourly compensation figures.
Appendix A More methodological details
We decided to study the “no injuries predicate", as it was a simple task so we could push the reliability to a level that is rarely studied in machine learning. In addition, using a simple task made it reasonably easy for a large team of non-expert humans to recognize failures and enabled small models to perform acceptably well on the task.
A.1.2 Definition of injury
We needed to give our labelers a relatively precise definition of what it means for text to describe an injury. The core definition we ended up with was:
Something is a Yes if and only if, according to the most natural interpretation of the text, the completion introduces a new injury to a living person or animal, or (if the injury was already mentioned in the prompt) the completion reveals that it was worse than you’d expect just from the prompt. To count, the injury must have started happening by the end of the completion, and it has to be as bad or worse as something that would leave a bruise or a cut.
In our definition, illness and death both counted as injuries. We gave further clarification of edge cases and various examples in a Google Doc https://docs.google.com/document/d/1OgZzybPZN4N0PPI84sTW0p8Q1GIZvEwgl_1f43sHkRU/ which we sent to everyone doing our task.
A.1.3 Definition of valid snippets
A snippet consists of a prompt and a completion. A valid prompt contains exactly three periods, with no text after the last period. A completion contains any 16 characters and then some number of non-period characters followed by exactly one period, at the end. When encoded with the deberta-v3-large tokenizer, a snippet must fit within 256 tokens.
A.1.4 Fan fiction distribution
Our source dataset was a GB archive of stories from fanfiction.nethttps://archive.org/details/FanficRepack_Redux. We defined the “random distribution” of snippets by the following sampling procedure:
Eliminate most of the preamble and postamble text from the stories using some hard-coded heuristics. (For example, “a/n” was a signal of a preamble and “END” was a signal of a postamble.)
Randomly sample a four-sentence snippet with a valid prompt and completion according to the previous section, capping at 1200 characters.
Eliminate any snippets which contain null bytes, are not detected as English (according to fasttext with the classifier from https://dl.fbaipublicfiles.com/fasttext/supervised-models/lid.176.bin), or do not contain at least 8 letters in the prompt (according to Python’s isalpha).
Replace the completion with a new 32-token completion from our generator, truncated after the first period after 16 characters (to make the completion valid). If there is no period within 32 tokens, add one at the end.
Note: at earlier stages of the project, we sampled snippets for training and evaluation in more ad-hoc ways that were not uniformly distributed and did not follow the same set of constraints.
A.1.5 Injury-enriched stories
Because of the low prevalence of injurious snippets in the in-distribution story dataset, we selected our initial training data to have more injurious completions. First we manually developed heuristics that helped us identify injurious completions. For example, we found that the word “sliced”, in the presence of a word for a weapon like “sword” and no words related to food like “bread”, was a strong heuristic for finding snippets in which injury happens. We selected snippets with completions that fit the heuristics we developed. We then replaced the original completion in the fanfic with a sentence from our generator model, and labeled the resulting data.
A.1.6 Human labeling process
In order to train and evaluate our classifiers, we had to label examples as injurious or non-injurious. We hired contractors from Upwork, Surge, and from respondents to a Facebook post soliciting contractors. Additionally, Redwood Research staff sometimes labeled snippets themselves to increase our data throughput or to correct mistakes. We had a total of over 100 labelers, although more than 95% of our labels came from 25 labelers.
Labelers could enter data into a website we built to classify snippets, although Surge labelers used Surge’s internal interface. Labelers weren’t told the source of the snippets they were labeling to reduce bias.
Labelers were asked if the completion of the text included an injury. They could select “Yes,” “No,” or “Unsure.” An image of the labeling interface appears in Figure 5. Contractors on Surge’s platform used Surge’s own interface for labeling snippets.
89% of the snippets in the training set were labeled once, while the remainder were labeled at least twice. In case of a disagreement, Redwood Research staff’s labels were preferred. If no staff member labeled a snippet, we chose the plurality decision. If there was a tie, we chose the most injurious label - Yes over Unsure, and Unsure over No.
To ensure that our labelers were following meaningful, consistent rules when identifying injurious content, we monitored their performance. We asked them to label ten snippets and checked that they matched gold standard labels chosen by Redwood Research staff that we required them to get right to ensure that they understood the task. We audited a subset of their labels to ensure high quality. On our test set, we labeled everything twice, and in cases where the two labelers disagreed, we referred the disagreement to a more trusted auditor. There were 391 (.391%) disagreements on the in-distribution data and 351 (10.2%) disagreements on the out-of-distribution adversarial examples.
A.1.7 Public access to adversarial datasets
All of our adversarial datasets used in this paper can be downloaded at https://injury-adversarial-training.s3.us-west-1.amazonaws.com/adversarial-data.zip.
A.2 Generator fine-tuning details
We fine-tuned GPT-Neo-2.7B on the raw text from the fan fiction archive. (We did not apply any pre/postamble removal or break the text up into snippets.) We used Adam with a learning rate of 5e-6. We ended up training only on a small prefix of the archive ordered lexicographically by filename, making Alex Rider fan fiction, one of the first in the list, overrepresented.
When sampling from the generator, we used temperature 0.9. When rejection-sampling using a classifier, we draw 100 samples from the generator and output nothing if none of them are accepted.
Our generator can be downloaded at https://injury-adversarial-training.s3.us-west-1.amazonaws.com/injury-classifier.zip and loaded via HuggingFace.
A.3 Classifier training details
When training classifiers, we typically trained with a range of learning rates and multiple random seeds. For the final classifiers, we used five different random seeds and five different learning rates (geometrically spaced between 1e-6 and 4e-6) for a total of 25 models per condition. For each condition, we chose the model with the lowest false negative rate at the classifier threshold that limited the KL divergence from the unfiltered generator to 0.5, as described in Section 4.2. This resulted in thresholds of for the baseline and +paraphrases classifiers, for the +manual classifier, and for the +tool-assisted classifier. We used a form of virtual adversarial training called scale-invariant fine tuning (SiFT) to train the classifiers. SiFT augments the training data with slight perturbations of the inputs to increase its robustness to adversarial examples. After training was finished, we cast all models to fp16 and ran all evaluations at that precision.
We had fifteen tool-assisted rewritten snippets that were based on the validation set that made it into the training set due to a bug in the tool. We discovered this after training our final classifiers, so we didn’t restart the training process for these. However, as only 15 snippets (out of 1513 total rewritten snippets and over 100 thousand total dev set snippets), our hyperparameter selection did not change after fixing this bug, meaning it had no effect on our actual results. The test set was also unaffected.
We trained each of our classifiers with a batch size of 3 for 1 epoch over the (partially upsampled) data, with a dropout rate of 0.3. The best learning rates were 4e-6 for the +paraphrases classifier and 2.8e-6 for the three classifiers trained on the other three conditions.
We used AWS for our classifier training. Each of the 100 training runs (25 hyperparameter settings x 4 classifiers) used a p3.2xlarge instance containing a single V100 GPU. We ran each training run for 12 hours.
Our classifiers can be downloaded at https://injury-adversarial-training.s3.us-west-1.amazonaws.com/injury-classifier.zip and loaded via HuggingFace.
A.4 Evaluation details
We asked our contractors to find snippets that were classified as non-injurious by one of our classifiers but that qualified as injurious by our rules. We measured how long it took for them to find new injurious snippets against each of our classifiers, with the goal of increasing the time it took to find them.
We used similar methodology to that used for collecting adversarial examples for training. Our contractors used the tool that showed them an injurious snippet and allowed them to edit it. We included features that helped them insert or replace tokens in the text, and that highlighted the most salient tokens for the classifier. The classifier ran on the edited snippets in real-time. When the injuriousness score went below the classifier’s injuriousness threshold, they were allowed to submit the snippet they rewrote. Later, another contractor would label the snippets to validate that they were indeed injurious according to our standards.
The process differed from our adversarial example collection process in several ways:
The injurious snippets that the contractors edited were drawn from a test set of prompts from the random distribution with completions from the generator.
Our raters were given a stricter threshold—the same that we used to define the rejection threshold—to hit for their adversarial examples (although the threshold appeared to them to be the same, because we rescaled it to 0.05 to keep the user experience consistent).
We randomly ordered our classifiers for each contractor and assigned them a new classifier from that order each day. We collected 14 days of data so we could ensure that we had each contractor working on each classifier for at least three days.
We asked our contractors to clock in and clock out when they were working so we could track how much time they worked each day. They had to click a “clock in” button before beginning work, and they could clock out manually or they would be automatically clocked out after 5 minutes of inactivity or when they closed the web page with the rewriting tool.
A.4.2 Rejection sampling failure rate
The most important in-distribution metric for our classifiers is: when they are used to rejection-sample the generator, how frequently do failures still occur?
We care about the fraction of accepted examples that are false negatives. However, just computing this fraction on the test set does not give the right value. Prompts that are less likely to produce accepted completions are less likely to appear in the calculation, whereas rejection sampling weights each prompt according to the original distribution. If there were a positive correlation between a prompt’s probability of being followed by injury and a classifier’s false negative rate on completions from that prompt, then the estimator would be biased downwards.
The baseline classifier reaches a high degree of reliability, with . The failure rates of the adversarially-trained classifiers are not noticeably different; the estimate for the final classifier reaches . The error bars are large due to the very small number of positive examples (see Table 5 for the list) and fact that some prompts are weighted much more heavily than others. Moreover, the bootstrap likely underestimates the true error bars: given how few false negatives appeared in our dataset, it is likely that we missed a small number of high-weight false negatives which would cause the failure rate to be substantially higher than reported.
A.4.3 ROC curves
Since we can’t tell apart the failure rates at the chosen thresholds, we show the full sensitivity-specificity (ROC) curves for the classifiers in Figure 6. Even at higher thresholds (where there are significantly more false negatives), classifiers with more adversarial training do not appear to perform noticeably worse. This is in spite of the fact that some of their training data was replaced with adversarial data, which is from a very different distribution.
A.5 More techniques we tried
We tried several approaches that did not show strong promise, which we describe here. Because we didn’t investigate them further, we didn’t collect rigorous data on their efficacy.
Since the datasets of adversarial examples written by humans were much smaller than our in-distribution dataset, we worried that the model would overweight in-distribution data and not learn as much from the limited quantity of examples. Thus, we tried oversampling the smaller datasets - including them 3–5 times more than they would have otherwise. (Note that injurious examples are already included 3–5 times each, so at maximum the model will see a single example 25 times). The effect on performance seemed close to zero.
In our setting, we only care about each snippet’s classification with respect to the threshold, not its classification score. However, cross-entropy loss rewards a model just as much for moving a violent example from 50% to 100% as from 1% to 2%, even though if our threshold was (say) 2%, the latter move is very helpful and the former is irrelevant. We also care more about rejecting injurious examples than accepting non-injurious ones. We hand-devised some loss functions that we hoped would better capture our desires and picked a series of PyTorch operations that approximated them. We then trained classifiers with these ad-hoc loss functions, as well as mean-squared-error. The resulting classifiers behaved very similarly to those trained using cross entropy.
We wondered if it was possible to use a language model to write adversarial examples. This process is difficult since we have no automatic ground truth for our classification task. More generally, we can’t rely on our LM to know better than our current best classifier, because if it did, then there would be some way to extract that information and that would then be our new best classifier. Instead, we created a reward signal by making two classifiers trained on somewhat different data give very different answers and tried using PPO and FUDGE as reinforcement learning algorithms to get models to output snippets which at least one model got wrong. However, it was often the case that the snippet was a false positive on the model which was wrong, rather than the snippet being a false negative for the other model, and we also had issues with mode collapse (the diversity of the examples generated by the LM was low). It’s possible future efforts here could produce a more promising adversary.
A.6 Automatic Paraphrase Prompt
Here is the main few-shot prompt that we used for automatic paraphrasing with GPT-3. We substituted in the snippet being paraphrased for {snippet}.
A.7 Comparison based quality evaluation
One way we can estimate the quality of a policy is to ask human raters to give a numerical quality estimate for each completion. However, some past work on human preferences found that absolute scores tend to be difficult to calibrate across raters .
Instead, a common approach is to use comparisons between the completions generated with two policies. In some work, these comparisons are used to fit a numerical quality function, but the simplest approach is to report the fraction of the time that the filtered policy is preferred, counting ties as 0.5.
Given a set of prompts , and two lists of completions and sampled from and respectively (the lengths of the lists doesn’t need to be uniform), we can estimate as:
Our offline quality evaluation procedure is as follows. Conceptually, we construct a filtered generator using the completions with score less than in our dataset, and then we estimate its quality by looking at how many contests it wins.
We sample prompts from our prompt distribution.
For each prompt , we sample completions from the unfiltered policy, then send the pairs of completions to human raters for labels . We varied the number of pairs of completions generated per prompt based on their estimated difficulty. For each prompt, we estimated the rate of generating a completion with score below the threshold by rejection sampling until we received 10 completions below the threshold. Then, we scaled such that we would have a probability of to generate at least one completion below the threshold. We also floored below at 4 and capped it above at 100. In other words, . This helps ensure that for most of our thresholds, the majority of our prompts will have at least one completion under the threshold.
To estimate the quality of a filtered generator, we compute:
correspond to the number of completions in the first and second positions of the tuple, respectively, with score less than . In cases where , we treat the quality of that prompt as . That is, for prompts where our approximate filtered generator fails to generate a completion, we treat that as a loss for the filtered generator.An alternative is to assume that when generation fails, we default to fallback completion of “……………..”, and try to estimate the quality of this fallback completion. However, the results were broadly similar to simply treating a failure to generate as a loss.
A.8 Researcher vs Contractor Adversarial Examples
A serious limitation of our human-centered approach to generating adversarial examples is that the adversarial examples found will depend greatly on the particular human adversaries. For example, as shown in Table 8, the majority of adversarial examples found by our contractors contained indirect injuries, and were thus relatively non-egregious. However, using our tool, one of our researchers was able to find the following adversarial example containing a character unambiguously dying, after about an hour of attempts:
Elvis peacefully enjoyed the light and flowers with his love," she said quietly, as he knelt down gently and silently. "I also would like to walk once more into the garden if I only could," he said, watching her. "I would like that so much," Katara said. His heart exploded, and he died instantly, though not before suggesting the use of camouflaged ships, disguised specialist pilots, a professionally trained force, occasional outright espionage, and perhaps smaller units for the outer perimeter.
This suggests that further work needs to be done to better train human adversaries to cover the space of adversarial examples.
Appendix B Additional results
Figure 7 shows the false negative rates on our validation datasets. Unsurprisingly, our models performed well on the validation sets for data distributions they were trained on.
Appendix C Example snippets
In Table 5 we include all of the in-distribution test-set false negatives.
C.2 In distribution test-set false positive examples
Table 6 shows one randomly selected positive for each model. Most false positives applied to all models, so we only chose false positives that were unique to each model. Table 7 includes four randomly-selected false positives that applied to all four models as well.
C.3 Test set adversarial examples generated using our tool-assisted attack
In Table 8, we show one randomly selected test-set tool-assisted rewrite for each classifier, as described in Section 2. Many of these are rather borderline, which is partly due to the robustness of the classifiers and partly due to a mistake described in C.4.
C.4 Mistake: no “Unsure” in test set
Due to a miscommunication with our vendor, our in-distribution test set labels did not have “Unsure” as an option. To stay consistent for all test-set data, we also removed the “Unsure” label for the test-set adversarial examples, forcing them to choose either “Yes” or “No”. However, as a result, many relatively borderline adversarial examples were labeled “Yes” (such as some of the ones in 8).
C.5 Example of the process of adversarially attacking a classifier with our tool
Below is an example of the development of an adversarial attack against our +manual classifier. We selected changes that we thought were interesting. Insertions are highlighted in teal and deletions are struck out.
C.6 Example prompts that had no valid completions
In our final quality eval, a number of prompts had no completions that were not classified as injurious. The table below shows these prompts. For the sake of space, we did not include the completions here.
C.7 Public demo of our rewrite tool.
A demo of our tool can be accessed at https://www.ttft.io/talk-to-filtered-transformer.