Understanding and Evaluating Racial Biases in Image Captioning
Dora Zhao, Angelina Wang, Olga Russakovsky
Introduction
Computer vision applications have become ingrained in numerous aspects of everyday life, and problematically, so have the societal biases they contain. For example, gender and racial biases are prevalent in image tagging and image search ; visual recognition models have disparate error rates across demographics and geographic regions . The perpetuation and amplification of social biases precipitate the need for a deeper exploration of these systems and of the bias propagation pathways.
We focus on the task of image captioning: the process of generating a textual description of an image . This task serves as an important testbed for visual reasoning and can improve accessibility of digital images for people who are blind or low vision.
In this work, we assess the pathways for bias propagation: from the images, to the manual captions, and finally to the automatically generated captions. We focus our attention on studying the Common Objects in Context (COCO) dataset; it is a widely used image captioning benchmark , thus making any biases especially problematic . We collect both skin color and perceived gender annotations on 28,315 of the people in the COCO 2014 validation dataset after obtaining IRB approval. This data allows us (and future researchers) to analyze disparities in image captioning (and other visual recognition tasks) across different demographics. Concretely, we observe:
The dataset is heavily skewed towards lighter-skinned (7.5x more common than darker-skinned) and male (2.0x more than female) individuals.The gender disparity was previously observed in although with automatically-inferred rather than manually-annotated labels. Further, darker-skinned females are especially underrepresented, appearing 23.1x less than lighter-skinned males.
There are racial terms (including racial slurs) in the manual captions. The racial descriptors are not learned by the older captioning systems , but are learned by the newer transformer-based models – although the slurs do not yet appear to be learned.
Image captioning systems perform slightly better (according to CIDEr and BLEU , although not SPICE ) on images of lighter-skinned people. This is consistent with disparate accuracies on e.g., pedestrian detection and facial recognition .
There are visual differences in the depictions of lighter and darker-skinned individuals. For example, lighter-skinned people tend to be pictured more with indoor and furniture objects, whereas darker-skinned people tend to be more with outdoor and vehicle objects.
Even after controlling for visual appearance, the captions still differ in word choices used to describe images with lighter versus darker-skinned individuals. This is particularly apparent in the manual captions and in modern transformer-based systems.
Our work lays the foundation for studying bias propagation in image captioning on the popular COCO dataset. Data and code is freely available for research purposes at https://princetonvisualai.github.io/imagecaptioning-bias/.
Related Work
Presence of dataset bias. Our work follows a long line of literature identifying, analyzing, and mitigating bias in machine learning systems. One key facet of this discussion is the bias in datasets used to train models. Under the framework of representational harms , there is commonly a lack of representation and stereotyped portrayal of certain marginalized demographic groups. Along with many ethical concerns , these dataset biases are problematic because they can propagate into models . In this work we analyze the biases present in a commonly-used image captioning benchmark, COCO , using our new crowdsourced annotations.
Mitigating dataset bias. The root causes of dataset bias are complex: they stem from bias in image search engines , data collection practices , and real-world disparities. Proposed solutions to dataset bias include new data collection approaches , manual data cleanup , synthetic data generation – or, in extreme cases, even withdrawing the dataset after insurmountable biases have been identified . Researchers have advocated for increased transparency of datasets , including developing tools to steer researcher intervention . Our work does not aim to mitigate dataset bias but instead to articulate its impact on downstream image captioning models.
Algorithmic bias mitigation. In tandem with efforts to reform data collection, a variety of algorithmic bias mitigation techniques have been proposed; see e.g., Hutchinson and Mitchell. for an overview. This work goes along with others that unveil biases present in existing algorithms . One important theme is bias amplification , or social biases in the data getting amplified in the trained models. In this vein, we study how bias in manual image captions propagates into automated captioning systems.
Image captioning models. Image captioning models are increasingly being developed as a more complex way of labeling images . Recent work has discovered biases in these systems, but often with respect to gender ; the study of racial biases in captioning has been limited to analyzing bias in the manual captions . Racial bias has been identified in other automated systems (e.g., speech recognition , facial recognition , pedestrian detection ); here we expand this work to studying racial biases in image captioning. This spurs the important question of whether race should be included in generated image captions at all. Prior works find that, in certain contexts, people who are blind or low vision want racial descriptors to be included. Further, this motivates the need to understand how people prefer their identities labeled by an automated captioning system, a question studied extensively by Bennett et al. .
Crowdsourcing Demographic Annotations
Dataset. To study bias in image captioning systems, we collect annotations on COCO , a large-scale dataset containing images, labels, segmentations, and 5 human-annotated captions per image. COCO is a widely used image captioning benchmark. We focus on the 40,504 images of the COCO 2014 validation set, and look for person instances with sufficiently large bounding boxes (at least 5,500 pixels in area) such that there is a reasonable expectation of being able to infer gender and skin color. This results in 15,762 images and 28,315 person instances.
Annotation setup. Using Amazon’s Mechanical Turk (AMT), we crowdsource race and gender labels. In our interface (Fig. 1), we present workers with a person instance in a COCO image and ask them to provide the skin color using the Fitzpatrick Skin Type scale , ranging from 1 (lightest) to 6 (darkest), and the binary gender expression. We also give workers the option of marking “unsure” for either. Each instance is annotated three times. We compensate the workers at a rate of $10 / hr.
Inferring race and gender. Race and gender annotations are fundamentally imperfect . First, the annotated labels may differ from the person’s identity. Second, the labels are discretized (which enables disaggregated analysis at the cost of collapsing identities). Further, the labels are for social constructs and thus subjective and influenced by the annotators’ perceptions. We follow prior work in formulating our annotation process; we use phenotypic skin color as a proxy for race because of its visual saliency over other conceptualizations of race. However, as noted by Hanna et al. , we are actualizing a particular static conceptualization of observed race here. By operationalizing race this way, we miss differences that may appear in other operationalizations, such as racial identity.
Quality control. To ensure annotation quality to the extent possible, we limit the task to workers who have completed over 1,000 tasks with a acceptance rate. We also construct 57 gold standard images where the gender and light-or-dark labels were agreed-upon by five independent annotators, including one of the authors. We inject 5 of these images randomly in a task with 50 images, and only allow workers who have correctly labeled these images to submit.
2 Gender annotations
We start by analyzing the collected gender annotations, looking at distributions at both the instance and image level.
Instance-level annotations. We analyze the gender annotations of the 28,315 person instances. To determine the label for a person, we use the majority over the three annotations. If majority is not achieved, or there are contradictory gender labels, the instance is labelled as no consensus. We observe that contradictory gender labels are most common when the person is a child, has obscured facial features, or possesses features that contradict social gender stereotypes (e.g. woman with short hair).
Analyzing the distribution, we see that males make up of the instances compared to females who only comprise (see Fig. 2). Most of the remaining instances were annotated unsure (), and a consensus was unable to be reached for only of instances.
Image-level annotations. To analyze the dataset at the granularity of images, which is what the captions refer to, we map individual instance annotations to the image (as there are often multiple people per image). We use the annotations given to the largest bounding box, under the assumption that captions will mainly refer to the largest person in the image . The only exception is if the second largest bounding box contains an individual of the opposite gender, and is more than half the size of the largest bounding box. In this case, we categorize the image as both.
The image-level distribution closely mirrors that of the instance-level (Fig. 2). Again, there are more than twice as many male images () as female images ().
Comparing collected gender annotations with automatically derived ones . Previously, works looking at gender bias in COCO have used gender labels derived from the manual captions: “[if] any of the captions mention the word man or woman we mark it, removing any images that mention both genders.” We compare our annotations with theirs. They label 5,413 images: our labels agree with theirs on 66.3% and disagree on 1.4%; the remaining 32.3% we determine cannot be reliably labeled with one gender, e.g., because the person is too small or there are multiple people of different genders in the image. We successfully label 10,780 images; they only label 3,591 of these correctly (details in Appendix A). This is consistent with the argument of Jacobs and Wallach : gender is operationalized differently in caption-derived versus human-collected annotations.
3 Skin color annotations
For the skin color annotations, we follow a similar process as with our gender annotations. The only difference is that we add a method for dividing skin color into the broader categories of lighter and darker. Using these new categories, we similarly analyze the skin color distribution at both the instance and image level.
Instance-level skin color distribution. Using the same schema as in Sec. 3.2, we obtain instance-level annotations for skin color. The top two most frequently occurring Fitzpatrick Skin Types are 2 () and 1 (). In contrast, Fitzpatrick Skin Types 5 and 6 comprise only and of the instances, respectively. This underrepresentation of darker-skinned individuals is an example of representational harm in and of itself.
We also include a broader skin color breakdown consisting of two categories: lighter and darker. Following previous work , we define the lighter category as all instances rated 1-3 on the Fitzpatrick scale and darker as containing 4-6. We also assign some of the instances that were previously uncategorized by skin color (because of conflicting labels assigned under the more granular 6-point scheme) to these broader categories. Using this skin color breakdown, of the instances are lighter individuals, whereas only are darker individuals. The amount of no consensus instances decreases from to when using this breakdown.
Image-level skin color distribution. At the image-level, we categorize skin color as lighter and darker, employing the same consensus method as for gender in Sec. 3.2. Of the images, are part of the lighter category and are part of darker, meaning there are 9.2x more lighter-skinned images than darker-skinned.
Intersectional analysis. We analyze the skin color and gender labels in tandem. Within lighter images, males are overrepresented at compared to females at . However, this difference is even starker when looking at darker images, where males comprise of the images while females only make up , reflecting the unique intersectional underrepresentation faced by darker-skinned females, as noted by Buolamwini and Gebru . In fact, of the 15,762 images annotated, only 226 of them () are of darker-skinned females.
Worker information. AMT workers were asked to optionally disclose their own race and gender identity. Of the workers asked, provided their gender and provided their race. As seen in Fig. 2, the annotators are predominantly white () and male ().
Prior work has found that annotators describe in-group versus out-group members differently . Thus, there may be a concern that the skew in worker demographics could influence our collected labels. To understand whether a worker’s demographics influences their selection of labels, we explore disagreements in annotations. We do so by comparing the mean difference in annotation when the pair of workers are of the same self-reported demographic group versus when they are of differing groups. If workers from different groups label images differently, we would expect pairs from distinct groups to have a greater disagreement than pairs from the same group. However, we find for skin tone there is not a substantial difference in the disagreement between pairs of the same racial group () and different groups (). For gender, the mean difference for same gender pairs () and different gender pairs () is similar as well. This indicates that there is not a systematic difference between how workers of different self-reported demographic groups label images, suggesting our collected labels would be similar even if the workers came from a different demographic composition.
Experiments
We now discuss the findings from our experiments on understanding what kinds of biases propagate in image captioning systems. First, we examine racial terms (Sec. 4.1) and disparate performance (Sec. 4.2). We then analyze bias in terms of representation, i.e., differences between the lighter and darker images and corresponding captions. To do this we first consider the images in Sec. 4.3, before controlling for these visual differences and studying the captions in Sec. 4.4.
Models. We examine the captions generated by six image captioning models: (1) FC is a simple sequence encoder that takes in image features encoded by a CNN; (2) Att2in is similar but images are encoded using spatial features; (3) DiscCap further adds a loss term to encourage discriminability; (4-6) Transformer , AoANet , and Oscar are transformer-based models representing the current state-of-the-art. In our analysis we particularly focus on contrasting Att2in vs DiscCap, since they differ only in the added discriminability loss, and the older (1-3) vs the newer (4-6) models. We train the models on the COCO 2014 training set using proposed hyperparameters from the respective papers (e.g., the discriminability loss weight is for DiscCap). Oscar is further pre-trained on a public corpus of text-image pairs
Data. Our racial analysis is performed on 10,969 images of the COCO 2014 validation set which were definitively labeled as either lighter or darker (not both or unsure).
We begin by analyzing the presence of racial descriptors and offensive language in the manual as well as automatically generated captions.
Manual captions. Prior works show that people are more likely to use racial descriptors when describing non-white individuals. We observe this pattern in human-annotated captions by conducting a keyword search of the captions in the COCO 2014 training set using a precompiled list of racial descriptors (details in Appendix B). For ambiguous terms (e.g. “white”, “black”) that can be used in a non-racial context, we manually inspect the captions. Assuming the training distribution mirrors that of the validation, for the manual captions, annotators used racial descriptors to describe individuals who appear to be white of the time versus of the time for individuals who appear to be Black. Furthermore, in of the instances when a racial descriptor for a white individual is used, the annotator is also mentioning an individual of a different race in the caption as well (e.g. “the white woman and Black woman”). We see this as a manifestation of the belief that “white” is the norm, and race is only salient when there is a deviation or explicit difference between multiple people.
In addition to looking for racial descriptors, we check for the presence of slurs and offensive language using a precompiled list of profane words . There are 1,691 instances of profane language, occurring in of the sentences in the COCO 2014 training set. We find alarming occurrences not only of racial slurs but also of homophobic and sexist language as well, similar to the NSFW discoveries by Prabhu and Birhane .
Automated captions. Racial descriptors are not found in the automated captions generated by FC, Att2In, DiscCap, AoANet, or Oscar. While this may be attributed to the fact that racial descriptors are uncommon in the training set, we disprove the idea that this is wholly the reason. To do so, we observe that other words which occur at similar rates (and are thus equally uncommon) are in fact still present in the model-generated captions. For example, the word “Japanese” occurs 69 times in the training set and 0 times in AoANet-generated captions while other descriptors, such as “uncooked” and “soaked”, which appear 88 and 61 times in the training set, occur 2 and 6 times in the generated captions respectively.
While it is rare, we find that racial and cultural descriptors as well as offensive language do propagate into the captions generated by the newer transformer-based models. For Transformer, AoANet, and Oscar, we find instances of offensive language. In addition, there are racial descriptors in 2 of the captions generated by Transformer and 12 cultural descriptors. Furthermore, for 10 of the 14 images, the model uses these descriptors when the human captions do not contain any racial or cultural descriptors (Fig. 3). This leads to the worry that models may replicate offensive language or exploit spurious correlations to assign descriptors in a stereotypical and harmful way.
2 Performance differs slightly between lighter and darker images
We next evaluate whether image captioning models produce captions of different qualities on images with lighter-skinned people than darker-skinned people. To do so, we first assess the differences in BLEU , CIDEr and SPICE scores between captions on lighter and darker images. Both BLEU and CIDEr rely on n-gram matching with BLEU measuring precision and CIDEr the similarity between the generated caption and the “consensus” of manual captions. SPICE, however, focuses more on semantics, capturing how accurately a generated caption describes the image’s scene graph (e.g. objects, attributes).
From these results (Tbl. 1), we make two key observations. First, according to both BLEU and CIDEr, the models Att2in, Transformer, AoANet, and Oscar perform somewhat better on lighter images than darker images: e.g., they achieve , , , and higher CIDEr scores respectively on lighter than darker images. We observe that these differences in BLEU and CIDEr are not significant for the FC and DiscCap — likely because their overall CIDEr scores are worse, at only and respectively, whereas the other four models attain CIDEr scores above (see Appendix C). This suggests that the way models are choosing to describe the images may be better-suited for the majority group. In fact, we see there is a slight positive correlation between the performance of the model (as measured by CIDEr) and the differences in performance between the two groups with an of (Fig. 4). Second, there are no noticeable differences with SPICE, indicating that the captions identify key visual concepts equally accurately across both groups. Nonetheless, it is important to note that negative results do not indicate something is bias-free, but merely that our particular experiment did not uncover strong biases.
3 Visual appearance differs between lighter and darker images
The analyses so far only consider issues in the captions themselves, irrespective of the image. We now explore how the visual depictions of people of different groups differ. We analyze simple image layout statistics, apply the REVISE tool for discovering bias in datasets, and consider differences in visual appearance of the image content.
We split our skin-tone-labeled image dataset of 10,969 images into 9,609 images for training and 1,360 for testing.These images belong to the COCO 2017 training and validation set respectively; recall that all belong to the COCO 2014 validation set. We use area under the ROC curve (AUC) as our metric on a balanced (through re-weighting) test set, so random guessing would have an AUC of . We bootstrap over 1,000 resamples and report a confidence interval.
Our two best performing models are trained on the distance from center and the distance plus the gender. Distance alone achieves an AUC of ; adding gender increases the AUC to . Distance is predictive because darker-skinned individuals tend to be further from the image center than lighter-skinned individuals; this is troubling since the “important” parts of an image tend to be more centered . Gender is a useful feature since from Sec. 3.3 we know that the gender distribution differs between the two groups.
REVISE bias discovery. We next apply the REvealing VIsual biaSEs (REVISE) tool.We additionally include the 813 images labeled both in both groups. Using REVISE we discovered that darker-skinned people appear more frequently with outdoor objects, and lighter-skinned people appear more frequently with indoor objects (Fig. 5). Specifically, objects like sink, potted plant, and toothbrush all appear with lighter-skinned people over 13x as much as with darker-skinned people, despite lighter-skinned people only appearing in 7x as many images as darker-skinned people. Although at the moment the differences in object co-occurrences do not appear to have noticeable downstream effects (Sec. 4.2), these differences may lead to discrepancies in performance as certain objects become more easily identifiable for different skin tone groups.
Visual appearance. Finally, we use image classification models for a detailed examination of how the content of the images differs between different skin tones. To ensure that the skin color of the pictured individual does not affect the model’s prediction, we use COCO’s object-level segmentations to mask all the people objects. We fill in these masks with the average color pixel in the image. Using the masked images, we fine-tune a pre-trained ResNet-101 over five epochs using the Adam optimizer and a batch size of 64. We oversample the darker images to account for the imbalanced class sizes. During training, the learning rate is initialized to be 0.01 and decays by a factor of 0.1 after three epochs. The model achieves an AUC of , indicating that there is a slight learnable difference between the scenes of lighter and darker images.
4 Captions describe people differently based on skin tone
Finally, we consider how both manual and automatic captions differ when describing lighter versus darker images. To do so, we first control for the visual differences, in order to disentangle the issues coming from the image content versus from the words used in the caption. We do so by finding images that are as similar as possible in content, and differ only by the skin color of the people pictured, i.e., constructing counterfactuals within the realm of our existing dataset. Concretely, for each darker image, we find the corresponding lighter image that minimizes the Euclidean distance between the extracted ResNet-34 features of the masked images using the Gale-Shapley algorithm for stable matching (Fig. 6). After examining the results, we select the top most similar image pairs.
The resulting dataset has 876 images. When needed, we use 700 for training (80%) and 176 for testing (20%); otherwise we compute statistics over the whole dataset. As expected, a visual classifier trained on these images (with the people masked) achieves an AUC of only , failing to differentiate between the two groups.
In the following analyses, we use the same six models and training setup as in previous experiments. However, we use the dataset, introduced above, which consists of 876 unmasked images for evaluation. This data thus allows us to examine whether human-annotated and model-generated captions diverge even when visual differences (except skin color) are controlled.
For our first line of inquiry, we use the Valence Aware Dictionary and Sentiment Reasoner (VADER) to perform sentiment analysis on the human-annotated captions. Limitations include that sentiment analysis tools have been shown to encode societal biases themselves , and may not generalize well to out-of-distribution machine-generated text. VADER returns a compound polarity score from (strongly negative) to (strongly positive). Scores less than are considered negative; scores greater than positive. We find that human-annotated captions describing lighter images have a mean compound score of whereas those describing darker images have a mean compound score of . The difference in compound scores is statistically significant (), with captions describing lighter images being more positive.
We find that automated captioning systems do not appear to amplify the difference in sentiment scores between the two groups (Tbl. 2). The lack of difference is largely due to the fact that automated captions tend to be more neutral than the human-annotated ones, thus removing most of the sentiment. In fact, the compound scores were all less than 0.03, excluding scores for captions generated by Transformer (0.046 for lighter and 0.042 for darker).
4.2 Sentence embedding differences
For our next analysis, we use sentence embeddings from the Universal Sentence Encoder to compare how the semantic content of captions differs between lighter and darker images. To note, racial descriptors in the captions are not removed for this experiment.
We find that the classifier can differentiate between the captions with an AUC of , indicating a learnable difference in the resulting caption content despite the visual content (with skin tone masked) being indistinguishable.
We see in Tbl. 2 that the ability to differentiate based on embeddings drops in the generated captions, especially for the more advanced Transformer model to , which is almost random. Although humans appear to be assigning different content to similar images with people of different skin tones, automated captioning models do not appear to uphold this trend, at least with respect to the particular sentence embeddings we use.
4.3 Vocabulary differences
Finally, we consider word choice in the captions. We use a logistic regression model and a vocabulary of the 100 most commonly used words (filtering out articles, prepositions, and racial descriptors, e.g. “white”) in the COCO 2014 training set. Our features are size 100 binary indicators of whether a particular word is present in a caption. The classifier achieves an AUC of on human captions. Beyond the differential use of racial descriptors we already observed in Sec. 4.1, this suggests annotators use different vocabularies to describe images even with similar visual content (other than skin tone).
The ability to distinguish between lighter and darker images further increases when automated captions are used. Particularly, in Tbl. 2 we see from Att2in to DiscCap and FC to Transformer, the AUCs slightly increases from to and to , respectively. From FC to AoANet, there is a greater increase in AUC from to . We do note that, for Oscar, the ability to differentiate based on vocabulary decreases compared to FC as the AUC drops from to . This may be due to the fact that Oscar is pre-trained on a larger corpus of data; the greater dataset diversity may help diminish the differences between the vocabularies used. Overall, this leads us to believe that more advanced models are more likely to employ different word choices when describing different groups of people.
Interpreting these results relative to that of the previous section in which we found that the semantic content of generated captions did not differ much between different groups, we consider whether different words are being used despite caption content being similar. As an example, the sentences “Apples are good.” and “Apples are great.” may map to similar sentence embeddings, but the specific word choice employed is different. In this vein, we find, for instance, that on AoANet’s captions, the average coefficent of the word “road” is higher than that of the word “street” (where higher coefficients are predictive of darker), even though upon manual inspection the images being described are similar (see Appendix D). While differences in the usage of words, such as “road” and “street,” are relatively innocuous, these subtle differences in vocabulary may become more problematic when we consider how certain words like “articulate” have developed a different meaning when applied to Black people . Thus, in future work, it is important to consider not only the semantic differences captured in the sentence embeddings but also the specific words being employed.
Discussion and Conclusion
In this work, we seek to understand not only what racial biases are present in the COCO image captioning dataset, but also how these biases propagates into models trained on them. We annotate skin color and gender expression of people in the images, and consider various forms of bias such as those in the form of differentiability between different groups. We find instances of bias in the dataset and the automated image captioning models. However, we are careful to note that cases in which we did not find bias do not mean there are not any, merely that our particular experiments did not uncover them. By looking at the models that seem to be most indicative of where the image captioning space is progressing, we can see that the bias appears to be increasing. For researchers, this serves as a reminder to be cognizant that these biases already exist and a warning to be careful about the increasing bias that is likely to come with advancements in image captioning technology.
Based on these analyses, we propose directions for mitigating the biases found in captioning systems. First, from our findings in Sec. 3.2 and 4.1, we see that human annotators make assumptions about the demographics of people pictured or use different language when describing people of different skin tone groups. To mitigate this, dataset collectors can provide more explicit instructions for annotators (e.g. do not label gender or include racial descriptors to people). In addition, we also find that ground-truth captions contain profane language (Sec. 4.1). In line with existing mitigation efforts , manual captions containing slurs or other offensive concepts should be removed from the dataset. Additionally, in Fig. 2 we see that only of the dataset contained images of people with darker skin tones, i.e., 1096 images. We need to collect more diverse datasets such that we can measure disaggregated statistics and compare metrics such as the difference in SPICE scores with the knowledge that our measurements do not suffer from a high sampling bias. Finally, from our analysis of generated captions (Sec. 4.4), we note that Oscar exhibits less bias compared to the other transformer-based models. This suggests the greater dataset diversity from pre-training the model may help reduce the amount of bias that propagates into the automated captions.
Acknowledgements. This work is supported by the National Science Foundation under Grant No. 1763642 and the Friedland Independent Work Fund from Princeton University’s School of Engineering and Applied Sciences. We thank Arvind Narayanan, Karthik Narasimhan, Sunnie S. Y. Kim, Vikram V. Ramaswamy, and Zeyu Wang for their helpful comments and suggestions, as well as the Amazon Mechanical Turk workers for the annotations.
References
Appendix A Comparing collected gender annotations with automatically derived ones
We explore extending the schema introduced in Sec. 3.2 for deriving gender labels from captions in three ways: 1) only labeling images where there is a person who has a bounding box greater than 5,500 pixels, 2) expanding the list of gendered words beyond “man” and “woman”, and 3) having different cutoffs for how many captions (of the 5 per image) need to mention a gender for the image to be labeled. We call the use of the gendered set man, woman “few,” and that of our expanded set “many.”
Our expanded set “many” consists of the following words: [“male”, “boy”, “man”, “gentleman”, “boys”, “men”, “males”, “gentlemen”] and [“female”, “girl”, “woman”, “lady”, “girls”, “women”, “females”, “ladies”].
Our results in Fig. 7 show that while these extensions significantly increase both the number of images correctly labeled and the accuracy of labeled images, all methods are inaccurate and/or incomplete. The gender labels derived from captions remain highly imperfect, as expected, cautioning against automated means of gender derivation .
Appendix B Racial descriptors
When searching for descriptors of race and ethnicity in Sec. 4.1, we first convert the captions to lowercase. We then use the following keywords [“white”, “Caucasian”, “Black”, “African”, “Asian”, “Latino”, “Latina”, “Latinx”, “Hispanic”, “Native”, and “Indigenous”] — also in lowercase — to query the captions.
Appendix C Caption performance
In Sec. 4.2 we assess the differences in caption performance for BLEU , CIDER , and SPICE when evaluated on the COCO 2014 validation set. We extend this analysis by providing the overall scores across the four image captioning models and looking at two additional automated image captioning metrics.
To start, we look at the performance for our four models. As seen in Tbl. 3, Oscar has the best performance across all metrics. Further, we see that the newer transformer-based models outperform older models (e.g. FC, Att2in, and DiscCap) across all metrics as well.
We also report the differences in performance between lighter and darker images for two commonly used image captioning metrics — METEOR and ROUGE (Tbl. 4). Similar to the results for BLEU and CIDER, the differences for METEOR and ROUGE are greater for Att2in, Transformer, and Oscar. AoANet also shows some slight differences in performance for METEOR and ROUGE. This supports our observation that the better performing captioning models also tend to show greater discrepancies in performance between lighter and darker images.
Appendix D Vocabulary differences coefficients
In Sec. 4.4 we explore the different word choices in the captions describing lighter and darker images. We provide the most predictive words for lighter and darker across the manual captions and the automatically generated captions in Tbl. 5.