Fair Diffusion: Instructing Text-to-Image Generation Models on Fairness

Felix Friedrich, Manuel Brack, Lukas Struppek, Dominik Hintersdorf, Patrick Schramowski, Sasha Luccioni, Kristian Kersting

Fair Diffusion

In the following, we present our novel strategy, Fair Diffusion, for approaching fairness in text-to-image generation models. Therefore, we propose to instruct, as motivated before, text-to-image DMs on fairness with textual guidance. Furthermore, we elaborate on fairness definitions in the scope of investigating DMs. In the Methods section, we further present in more detail the textual interface for DMs and the underlying technique to guide their image generation.

We propose Fair Diffusion to instruct a pre-trained model on fairness during the deployment stage (cf. Fig. 1). In line with this goal, previous work has proposed approaches to alter the image generation process . Fair Diffusion can facilitate these various approaches as instructing interface. Here, we evaluate Fair Diffusion with Semantic Guidance (Sega ) since it enables highly flexible instructions through text (more details in Methods). With this tool, Fair Diffusion can be applied on a DM as illustrated in Fig. 1. Based on detected biases—e.g., stored in a lookup table as instructions—the model is guided to reduce bias in its outcomes, eventually moving toward fairer and user-aligned behavior.

Intuitively and on a high level, Fair Diffusion extends classifier-free guidance with an additional fair guidance term γ\gamma (cf. Eq. 5) to generate an image xx:

This way, the image generation, η\eta, is a function of the text input prompt pp (e.g., “A photo of a firefighter”) with fair guidance, γ\gamma, during the generation. In turn, γ\gamma depends on additional textual descriptions of attribute expressions eie_{i}, scaled by seis_{e_{i}} with guidance direction. As a result, the image generation is guided towards the input prompt pp and fairness instructions eie_{i} at the same time. In order to realize the different expressions of an attribute with Fair Diffusion, we control the guidance direction by randomly sampling from a desired probability distribution PP. That means each eie_{i} is either increased or decreased depending on the expression that should be promoted/ suppressed. In Fig. 1, we illustrate a binary case, where concept e1e_{1} is promoted (++) and e2e_{2} suppressed (−-) during the image generation. Their direction can be changed based on PP. Fair Diffusion can easily facilitate arbitrary distributions by adjusting the guidance probability accordingly. This level of flexibility allows the implementation of different definitions of fairness using the same underlying model. For the evaluation, we again verify the resulting attribute proportion with a classifier. Ideally, the user-defined and measured proportions match.

Fig. 1 illustrates how Fair Diffusion can be applied after deployment. As shown in Fig. 2 and discussed later, the default image generation with SD of “firefighters” suffers from gender and ethnic biases. With Fair Diffusion, a user provides fair instructions to the model at inference without touching the input prompt. In turn, fair guidance aligns the output with users’ fairness notions, e.g., more diverse “firefighters”. Here, a lookup table serves as means to identify text prompts requiring fair guidance. It contains user-designed instructions to promote the fair generation of images. This way, Fair Diffusion could be automatized and integrated into APIs. We show further background and full details in the Methods section.

Fairness for Diffusion Models

As said before, fairness has always been a challenging concept to define . Definitions of fairness and bias are, like many ethical concepts, always controversial, and a variety of definitions exist . Roughly, fairness can be summarized as the absence of any tendency in favor of a person due to some attribute. However, fairness is inherently subjective and suffers from incompleteness. In general, only in very specific and constrained situations, it is possible to satisfy multiple of these fairness notions. In turn, a universal definition is not available . We define fairness for Fair Diffusion, in line with closely related work on fairness , as algorithmic fairness for a dataset and model.

Given a (synthetic) dataset D\mathcal{D}, fairness or statistical parity is defined as

Here, y∈Yy\in\mathcal{Y} is the label of a respective data point x∈Xx\in\mathcal{X}, aa is a protected attribute and PP is a probability. For example, xx can be an image with the label yy “firefighter” and aa the protected attribute “gender”. This fairness definition can be used to evaluate the fairness of a dataset but also of a generative model. Typically, datasets consist of real-world data (xx) with human labels (yy). For the evaluation of a generative model, a dataset can be synthetically generated to enable an empirical fairness evaluation. In that case, a data point is obtained through x=η(y)x=\eta(y), where the model η\eta is prompted by the user with the desired label yy. The model η\eta can represent any generative downstream task for any input modality (visual, textual, etc.), e.g., η\eta can be a generative diffusion model mapping from text (also called prompt pp) to images. In other words, we define a dataset to be fair if Def. 1 holds, i.e., there is no disproportionate weight in favor of attribute aa in the data, and we define a model to be fair if the same holds true for a model’s generated output. However, this fairness definition ensures fairness for one binary attribute. In the case of multiple non-binary attributes that may interfere with each other, it becomes more challenging to satisfy them at the same time. Furthermore, this fairness definition requires all attributes to be known, definable, measurable, and separable. We discuss the limitations of this definition in Discussion. And, while there exist multiple definitions of fairness, in this work, we focus on the introduced one to first identify sources of unfairness in DMs and to second evaluate the mitigation of unfairness.

In line with the presented fairness definition, in Fig. 1, the concepts e1e_{1} and e2e_{2} describe the expressions of the (binary) attribute aa. To ensure statistical parity (Def. 1), i.e., that both attribute expressions are equally represented in the model’s outcome, we randomly choose the direction of the guidance with equal probability. This results in a uniform probability distribution, assigning the same probability to each expression of an attribute, i.e. P(a)=1∣a∣P(a)=\frac{1}{|a|}.

Experiments

In this section, we investigate the components of DMs as potential sources of bias on the prominent example of gender occupation biases. Subsequently, we demonstrate their mitigation in the outcome using Fair Diffusion.

For assessing the bias of text-to-image DMs and its mitigation, we inspect the publicly-available diffusion model Stable Diffusion 1.5 (SD https://huggingface.co/runwayml/stable-diffusion-v1-5), its underlying large-scale dataset (LAION-5B More specifically, the subset LAION-2B(en) available at https://laion.ai/blog/laion-5b/) and pre-trained model (CLIP https://github.com/openai/CLIP with “ViT-L/14”). For the inspection and generation, we used p=“A photo of the face of a {occ}”p=\text{``A photo of the face of a \{{occ}\}''} as a text prompt, where occ∈{“firefighter”, “teacher”, “aide”, ...}\textit{occ}\in\{\text{``firefighter'', ``teacher'', ``aide'', ...}\}. In terms of fairness, we assume that an equal proportion of female- and male-appearing images is desired as derived from Def. 1. We present more details on our experimental setup, including measures to quantify DMs’ inherited biases and the full experimental protocol in the Methods section.

We start our empirical study by illustrating the presence of biases in SD and laying the foundation for the subsequent bias mitigation. To this end, we investigated SD’s training data, LAION-5B, its pre-trained model, CLIP, and SD’s outcome on gender occupation biases.

To begin with, we evaluated LAION-5B on four subsets of our occupation list (cf. Fig. 3a), where each subset contains images of occupations belonging to one field (science, arts, engineering, and caregiving).Selection is explained in App. D Tab. 3. Therefore, we classified the images of the subsets for gender to obtain insights into the rate of female-appearing persons, the share of images κ\kappa (Eq. 7) classified as “female”, as a measure of gender occupation bias. Fig. 3b shows that LAION-5B contains several occupation biases. One can observe that the rate of female-appearing persons is higher for occupation fields like arts or caregiving. On the other hand, the rate is lower for science or engineering. Both demonstrate stereotypical proportions in the dataset. For inspecting CLIP, we performed a bias association test, the iEAT. In this experiment, we tested the similarity between encoded images of different concepts. In the spirit of Steed et al. , we applied their setup to the CLIP encoder and uncovered similar gender occupation biases (cf. Tab. 1). For instance, encoded images of male-appearing persons are closer to engineering-related images than encoded images of female-appearing persons, which are, in turn, closer to caregiving-related images. More interestingly, we found a bias amplification when the association test on gender occupation bias is modified by ethnic attributes. In this case, images of male-appearing people are represented by European-American appearance and images of females by African-American appearance. Accordingly, an intersectionality bias is present in CLIP too, amplifying the gender occupation bias. We show more details on this experiment in App. Tab. 2.

Next, we examine the third model component, the downstream task and its outcome, for gender occupation biases. Here, we evaluated the generated images from SD. Since SD builds on LAION-5B, we also compared the respective rates in the following to investigate SD’s outcome for mirrored unfairness. This time, we extend the previous investigation to the complete occupation subset of LAION-5B (cf. Fig. 3a). Fig. 4 depicts the rate of female-appearing persons for six exemplary occupations and its per-group median. One can observe that the SD-generated images (blue lines) also contain clear gender biases for various occupations. For instance, a firefighter or a social worker are significantly affected by bias. The evaluated LAION-5B images (darkgray lines) contain similar biases, providing further evidence for our previous findings. However, one can observe a discrepancy in gender occupation biases, e.g., for “firefighter”, between LAION-5B and SD-generated images. The rate of female-appearing persons is higher for the generated images than for its training data, representing a stronger gender bias. On the other hand, if we look at “coach”, the gender bias in the SD outcome is on par with LAION-5B. Furthermore, for a designer, the gender bias in the generated images is smaller than in its training data. At the same time, one can also see that there are occupations like “teacher”, which have nearly no gender bias in LAION-5B (true for overall 5% of the evaluated occupations); nonetheless, there is a bias in the generated images. Lastly, one can also observe that the gender bias in the generated images for “aide” is smaller than in LAION-5B but beyond the fair boundary in the opposite direction.

Overall, we found that LAION-5B’s gender biases in the generated images are amplified for 56%, reflected for 22%, and mitigated for 22% of the evaluated occupationsWe denote a reflection if the SD outcome bias corresponds to the LAION-5B bias ±4%\pm 4\%, and an amplification (mitigation) if the bias is farther away from (closer to) the fair boundary.. As one can find various behaviors in the SD outcome (amplification, reflection, and mitigation) regarding fairness, we computed a per-group statistic (binary fe/male) to further insight into the overall gender bias in each component. One can observe that the median (orange line) of SD-generated images and the inspected LAION-5B images are distinctly outside the fair boundary. This means LAION-5B and the SD-generated images are unfair according to Def. 1. In total, 64 of 150 occupations are more female-biased, while 86 are more male-biased. More importantly, the median of SD-generated images is farther away from the fair boundary (middle) for both lists (f and m) compared to LAION-5B. This provides evidence that the SD output from this experiment is, on average, more biased than the images from LAION-5B. However, one can observe high variance, urging more research in this direction. Furthermore, attributing the discrepancy in bias to a specific component of the model or aspect of the training procedure is difficult. The shift in bias results from a complex interplay between training data and objective, and CLIP’s inherently biased representations, which are in turn influenced by a different training set.

During our inspection, we found biases and unfairness in each component of the SD pipeline: in the LAION-5B dataset, in the CLIP encoder, and in the generated images. At the same time, the biases are not simply mirrored between LAION-5B and SD’s outcome and do not show a clear tendency.

Instructing on Fairness with Fair Diffusion

After discovering several biases in SD’s components, we turn to mitigate them. Since the interplay between SD’s components is complex, debiasing them is a challenging task. Unlike related work on debiasing, we instruct a model on fairness at the deployment stage. Next, we evaluate the guidance toward fairness of text-to-image generations.

In contrast to the image generation with default SD, we further conditioned the image generation with Fair Diffusion. To this end, we steered the image generation toward “female person” (e1e_{1}) and away from “male person” (e2e_{2}) and randomly switched the direction toward male-appearing and away from female-appearing with a 50% chance.We would like to emphasize again the limitation of binary gender in this setup (cf. Discussion). This way, we utilize the concepts encoded in a DM to simultaneously suppress one and reinforce the other, with alternating directions. Due to this approach, 50% of the images should contain (fe)male-appearing generated persons.We applied the chosen guidance to the image generation regardless of its outcome without guidance or any biases present in LAION-5B. For an apt comparison, we applied Fair Diffusion to SD-generated images, i.e., we re-generated an image with the same seed and parameters and included the additional conditioning (fair guidance γ\gamma) for gender.

Fig. 4 demonstrates Fair Diffusion’s performance (green arrow) in mitigating gender occupation biases detected in Stable Diffusion (blue bar). Looking at our six exemplary cases, one can observe a shift of the gender proportion to the inside of the fair boundary, no matter in which direction the bias was previously in. Furthermore, the biases can be addressed with the Fair Diffusion strategy, whether they are present in LAION-5B or the SD outcome and whether the SD outcome amplifies, reflects, or mitigates LAION-5B biases. For example, though the proportion in SD-generated images for “designer” is less biased than in LAION-5B, it is still not within the fair boundary. In turn, Fair Diffusion mitigates the bias further and shifts the gender proportion within the fair boundary. Moreover, the per-group (m/f) median (cf. Methods) is within the fair boundary, so Fair Diffusion successfully reduced unfair occupation proportions. Hence, on average, Fair Diffusion achieves fairness according to Def. 1, i.e., for the model outcome. However, we can see that there remains variance in the generated images for some occupations. This can generally be due to the non-binary nature of gender, and gender is also not to be determined simply based on outward appearance. Moreover, we could identify some outliers, e.g., images of “dishwasher” were generally difficult to generate but also difficult to edit for gender, as it does describe not only an occupation but also a cleaning device. When searching LAION-5Bhttps://knn5.laion.ai/ for “a photo of the face of a dishwasher”, we also mainly found images of the cleaning device and no humans. So, we assume this to be an artifact due to the ambiguity of “face of a dishwasher”. For a broader evaluation, in particular, on the design of the editing prompts, we refer to App. C.

Apart from our quantitative analysis, we also show qualitative results on our bias mitigation approach in the following. In Fig. 5, one can observe a shift in outward appearance. The top row shows images generated with SD and the prompt pp for each of our six exemplary occupations. As our inspection experiments already demonstrated, there is mostly a strong bias toward one gender for certain occupations. In contrast, the images generated with Fair Diffusion (bottom row) shift the gender appearance toward the other gender appearance and by that ultimately toward a more diverse output. Notably, the overall image composition remains the same, with only minor changes to the rest of the image.

Another step toward a fairer model outcome is to go beyond editing gender appearance. In Fig. 2, we used Fair Diffusion to edit multiple features of firefighters’ outward appearance. The result is a more diverse output of facial features in terms of gender, skin tone, and ethnicity. This is due to the capability of Fair Diffusion to modify image features according to user preferences in isolation. This way, we can remove biases and increase outcome impartiality beyond gender. In App. A, we show further possible approaches to biases beyond gender, e.g., heteronormativity and age discrimination.

In summary, it is difficult to create a model with fair outcomes in all aspects. Here, we evaluated instructing text-to-image models to approach outcome impartiality. Our empirical results on Fair Diffusion demonstrated its potential as a reliable approach. However, we emphasize the interplay between processing training data and fairness in representations as well as the outcome. We envision a future with models that are able to generate a more diverse outcome—hand in hand with the user who steers and controls the generation process.

Discussion

We have found several severe biases in all components of the SD pipeline and introduced Fair Diffusion to mitigate these biases. Still, we need to delve deeper into some of the insights gained and use this section to discuss them in greater detail. In particular, we touch upon opportunities for society.

This work provides food for thought on current debiasing techniques, mostly focused on dataset curation and in-process bias prevention, shifting the focus to the deployment stage. As Nichol et al. showed, curating datasets by filtering has drawbacks, such as persisting biases and worse generalization capabilities. In turn, Fair Diffusion operates at the deployment stage enabling fair outcomes according to Def. 1. On the other hand, our inspection also showed that the evaluated images from LAION-5B are, on average, still remarkably affected by bias. Consequently, debiasing it might remain important. Ultimately, we believe debiasing all components might be necessary to increase fairness in generative models further. Nevertheless, as long as this is not available, especially for end-users, we demonstrated that the text interface of DMs enables instructions as an easy-to-use technique that can be immediately deployed to mitigate biases in the outcome of current image generation models. This does not introduce an entirely fair model but instead a way to control unfair models to increase fairness. In the presented version, Fair Diffusion is applied regardless of, e.g., existing biases in LAION-5B. While Fair Diffusion enables a user to take control of fairness, a future vision for even fairer models is automated detection of unfairness. If biased concepts are known beforehand, a user could supply Fair Diffusion with these in order to take action without a user actively intervening in the moment of generation. For example, the lookup table and instructions in Fig. 1 could be filled out beforehand.

As shown, Fair Diffusion promotes fairness according to Def. 1. However, as discussed before, fairness is inherently incomplete, such that the setup used in this work does not account for other fairness definitions.

For example, it realizes fairness according to Def. 1 in the model outcome but cannot in the dataset. To achieve such fairness, the dataset has to be filtered or modified so that no biases remain. Still, Fair Diffusion can be used to realize different notions of fairness. If the desired output proportion differs from an equal proportion (50/50 for binary attributes), a user can realize this easily by setting a different edit probability PP that is non-uniform (e.g., to 70/30 for binary attributes). This may be used for a fairness definition that simply reflects the occupation proportions of current society, i.e., utilizing a country’s occupation statistic a user lives in. This further illustrates that fairness is versatile, and Fair Diffusion is easily adaptable to these different notions.

We acknowledge the limited representation of gender in this study. Current automated measures treat gender as a binary-valued attribute, which it is not . Due to the lack of tools that classify beyond binary gender, an empirical evaluation at a large scale remains limited to binary-valued gender evaluation. Similarly, we assume encoded concepts in Stable Diffusion to be limited to binary gender terms. Consequently, we observed non-binary fairness instructions to result in fragile behavior (cf. App. C). Although current diffusion models have inherent limitations, Fair Diffusion builds on them to make a first step toward fairness by mitigating, e.g., gender-biased image generation. We advocate for research on encoding gender as non-binary in generative and predictive models.

Furthermore, we also want to touch upon some technical limitations of this work. First, when evaluating the LAION-5B dataset, we observed that several images are stock photos. This is a reminder that LAION-5B is a web-crawled dataset that does not represent reality, nor does it reflect the internet in its entirety. Second, we use the CLIP encoder for searching LAION-5B. However, as demonstrated, CLIP is inherently biased , which may affect the search results. Therefore, the resulting images are not entirely disentangled from this confounding factor. However, there are barely any alternatives, as manual labeling is biased too and infeasible due to a large amount of data, and other automated approaches will suffer from imprecision too. We empirically chose the threshold ρ\rho to be low enough to counteract this behavior, so images of disadvantaged genders will be included in the search result. Third, the gender classification results rely on a pre-trained classifier, FairFace . As said, the classifier is an inherent limitation for classifying (binary) gender. Furthermore, we cannot guarantee that this classifier is bias-free and hence promote investigating its function and alternative ways. Yet, it seems to be the best available choice for an automated evaluation. Moreover, FairFace and the other limitations are only relevant to evaluate Fair Diffusion, while the strategy itself is independent of, e.g., the used classifier for evaluating. Lastly, Fair Diffusion builds on Sega and inherits its constraints. The concepts (biases) to be removed must be describable, i.e., either with natural language or with images. However, this is inherent to all currently available approaches that enable editing generative DMs. Fair Diffusion is agnostic to that, and Sega can, in principle, be replaced with other editing techniques.

Fair Diffusion currently operates on the textual interface to steer image generation. However, it is not limited to the textual modality. Fair Diffusion adds fair guidance to the image generation according to edit encodings ceic_{e_{i}}. Currently, CLIP’s text encoder embedded the edit prompt eie_{i} to ceic_{e_{i}}. The way the encoding is obtained can go beyond (English) natural language. Fair Diffusion could be extended with other approaches like AltCLIP for multilingual encodings, Textual Inversion for visual encodings, or MultiFusion for multimodal (text and image) encodings. This variety of approaches offers versatile interfaces for fair guidance with Fair Diffusion.

While human interaction has generally proven to be helpful , at the same time, certain dangers can arise. For example, a user in control with malicious intentions could target the model to misuse it. Like many other pieces of research, Fair Diffusion faces the dual-use dilemma. The strategy can be used in an adversarial manner as well, such that biased outcomes of generative models can be further amplified, and its diversity decreased. Hence, further detection mechanisms for malicious interaction are required. This is an active research topic that needs consideration when using human interaction.

Our work is of greater relevance as it offers the opportunity to immediately promote fairness in many real-world applications. As image generation models become increasingly popular and integrated into our lives, fairness must be kept in mind. DMs come into play even in high-stakes applications such as medicine and drug development . These models are also used in other areas, such as advertisement or designhttps://www.cosmopolitan.com/lifestyle/a40314356/dall-e-2-artificial-intelligence-cover/. Imagine a firefighter advertisementFor example, stock images are already being generated using diffusion models (Shutterstock, https://www.shutterstock.com/press/20435) containing people from Fig. 2 top or bottom row only. This way, generative models can have a crucial impact on societies and how we include and value diversity in them. Furthermore, Fair Diffusion can make another step toward fairness in society. This work focuses on a specific definition of fairness for evaluation purposes. However, the way such a tool is used also has a political dimension beyond research. Sometimes, the goal is not to achieve an equal outcome for each attribute. Temporarily over-representing a certain attribute, e.g., in advertisements, can be desired as it can promote awareness and transparency for bias and discrimination concerning this attribute. Or, current over-representations can be gradually reduced, to slowly habituate new proportions in a society. The pathway toward an ideal discrimination-free world may take measures that might contradict the fairness definition used in this work (Def. 1) but align with other fairness definitions . Hence, our approach facilitates flexible outcome proportions, which, in turn, enables over-representation or any other proportion. However, it remains an open question, what the new world (à la Aldous Huxley) should look like that an AI system reflects. We do not argue for any specific proportion or promote any specific political direction. Instead, we provide a strategy that can be used by society and politics immediately with ease for purposes that can ultimately promote a fairer depiction of society. Therefore, the overall goal might not be a fair tool itself, as it is rather a means to an end, but to use it in a way that promotes a fairer society without discrimination.

Conclusion

We introduced Fair Diffusion and demonstrated that it can instruct generative text-to-image models in terms of fairness. To measure fairness, we explored the publicly available large-scale training dataset of Stable Diffusion (LAION-5B) and further applied iEAT on its underlying pre-trained representation encoder (CLIP). Both show severe gender and racial biases mirrored in the downstream diffusion model. However, its textual interface and advanced steering approaches provide the necessary control to instruct the generative model on fairness, as our extensive evaluation demonstrates. Specifically, we show how to shift the bias in generated images in any direction yielding arbitrary proportions for, e.g., gender and race. In this way, our method prevents diffusion models from implicitly and unintentionally reflecting or even amplifying biases. Based on our findings, we strongly advise careful usage of such models. However, we also envision easily accessible generative models as a tool to amplify fairness, i.e., itself introducing syntactic biases—compared to real-world distributions—into realistic images. This enables media to display various genders in (fe)male over-represented occupations motivating younger people to follow their interests despite societal biases .

An exciting avenue for future work is disentangling the components to pinpoint the sources of bias in the model. Furthermore, it is interesting to compare different generative diffusion models for fairness. In addition, this work can be extended to image-to-image diffusion, facilitating the editing of real-world images rather than just generated ones. Lastly, Fair Diffusion can be easily integrated into any real-world diffusion application, mitigating unfair image generation or even amplifying fairness.

Methods

In this section, we introduce the underlying components of Fair Diffusion, measures to evaluate fairness in diffusion models and a detailed experimental protocol.

As visualized in Fig. 6, many recent models, like DMs, rely on large-scale training datasets as well as large-scale pre-trained models. These requirements are necessary to perform well in text-to-image generation tasks and generalize over multiple domains. The information transfer from the pre-trained model and the downstream training data helps these models achieve remarkable performance. However, combining both also increases the risk of introducing biases into the models, as we demonstrated in our experiments.

Bias Mitigation

Recently, many approaches have been proposed to create models with fairness in mind. For large-scale models, these methods can be categorized with respect to three paradigms: (i) pre-processing the training data to remove bias before learning , (ii) enforcing fairness during training by introducing constraints on the learning objective , and (iii) post-processing approaches to modify the model outcome at the deployment stage .

For DMs, Nichol et al. filtered the data prior to training in order to mitigate bias through the removal of data related to certain concepts. However, they observed the filtered model to continue exhibiting bias while encountering adverse effects such as a loss in generalization ability. These results highlight again that creating a completely bias-free dataset is generally infeasible. Additionally, different definitions of fairness would each require a dedicated dataset and respective model tailored to the targeted fairness characteristics. This contradicts a major principle of large-scale pre-training, i.e., training one model on only one (large) dataset and subsequently using it for downstream tasks. Hence, data pre-processing alone does not provide an apt solution for mitigating biases.

In contrast, this work targets the (post-process) deployment stage of DMs. Fortunately, large-scale models are not at the mercy of under-curated data. Schramowski et al. demonstrated that representations learned during pre-training can be exploited and leveraged to suppress unwanted and inappropriate behavior in the downstream task. Whereas their work focused on suppressing inappropriate (e.g., pornographic) content, we used a similar approach for fair outcomes, giving power to the individual user. This allows them to instruct the model on their individual definition of fairness. Other works have already shown that user instructions are an essential component for machine learning models to enable user alignment , trust , and overall model performance .

As discussed before, Fair Diffusion can be implemented with any image editing approach. In this work, we chose Sega as it offers a powerful interface to conduct image editing. Sega extends on the principles of classifier-free guidance, presented in previous section about text-guided image generation. In addition to the text prompt pp (and its guidance scale sgs_{g}), a user can provide additional textual concept descriptions eie_{i} (with their own scale and direction). This extends the previous noise estimation for image generation from ϵˉθ(zt,cp)\mathbf{\bar{\epsilon}}_{\theta}(\mathbf{z}_{t},\mathbf{c}_{p}) (Eq. 3) to ϵˉθ(zt,cp,ce)\mathbf{\bar{\epsilon}}_{\theta}(\mathbf{z}_{t},\mathbf{c}_{p},\mathbf{c}_{e}) (Eq. 5). In this way, the initially unconditioned image generation is additionally conditioned on the input prompt (classifier-free guidance) and the user instructions, e.g. on fairness (fair guidance γ\gamma). In particular, multiple concepts can be combined arbitrarily and either increased or decreased during the image generation. This way, we can also realize more complex changes by e.g. guiding toward the encoding ce1c_{e_{1}} of one concept e1e_{1} and simultaneously away from ce2c_{e_{2}} with text e2e_{2}, e.g., its alleged opposite. Overall, the resulting ϵ\epsilon-estimate, ϵθˉ\bar{\epsilon_{\theta}}, can be written as:

More technical details on this technique can be found by Ho et al. and Brack et al. .

In contrast to this implementation of fair guidance, other approaches such as simple prompt engineering are limited in their abilities and often change the entire image composition. On the other hand, Sega enables editing of dedicated features in isolation. Furthermore, changing gender appearance by inserting gender-related words into the input prompt is a non-trivial task for machines and requires language understanding. Determining the right position for the edit word(s) is challenging, especially if biases in the prompt are present only implicitly. Sega implements the challenging integration, e.g. of negations and conditions, in a sentence in a technical way. This way, it needs no such language understanding while enabling the same capacity as the edit instruction, i.e. the fair guidance) is independent of the input prompt. We illustrate this issue in App. Fig. 15.

Measuring Fairness in the Components of Diffusion Models

In previous sections, we showed that DMs are built around various components (cf. Fig. 6) and each can be affected by bias. Namely, bias in the datasets, the learned representations, and its reflection in the outcome. Next, we describe the measures used to quantify and track biases across all three components.

The first potential source of bias, according to the general model setup (cf. Fig. 6), is the dataset. Given a potentially biased attribute, e.g., gender, we investigate its co-occurrence with a target attribute such as occupation. If, for example, the proportion of genders within all samples of an occupation is not in line with the fairness definition (e.g. Def. 1), we have identified a source of bias already emanating from the dataset. This proportion also serves as a reference to investigate whether the model outcome reflects, amplifies, or mitigates such a bias. If the investigated dataset has no pre-existing labels for the attribute(s) of interest, they have to be derived first. However, this is generally a non-trivial task and cannot easily be transferred between different domains. For vision-language tasks, a sensible approach is to employ a multimodal model capable of computing text-image similarity. We identified relevant images R\mathcal{R} in the dataset by computing their similarity to a textual description pp of the target concept . Along these lines, we obtain the label at the same time, as textual description pp corresponds to label yy.

In this work, we selected images aligned with description pp by filtering the entire dataset with an empirically-determined threshold ρ\rho:

where I\mathcal{I} denotes the set of images from the dataset. Next, we used a pre-trained classifier, κ\kappa, to determine the (missing) label for the protected attribute under investigation. Consequently, we obtained each label ara_{r} for image r∈Rr\in\mathcal{R} with:

Second, we investigated the bias of learned representations using the image Embedding Association Test (iEAT) . Intuitively, iEAT tests for statistically significant associations between sets of representations, e.g., encoded images. These consist of two attribute sets AA and BB and two target sets KK and LL. A common example is target images of female-appearing people LL and male-appearing people KK compared against images related to career AA and family BB. This way, a biased model may associate the images of male-appearing people closer to “career” than to “family” and vice versa. Formally, the test statistic can be computed as:

This way, s(w,A,B)s(w,A,B) computes the association of an encoded image ww with the attributes (aa and bb) and eventually the differential association of the encoded target images with the attributes. We assess the statistical significance by computing the one-sided pp-value along with the effect size dd as:

where σ\sigma denotes the standard deviation.

The third source of bias we inspected is the downstream task approximated by its outcome. The modality of the outcome generally depends on the type of model and task. Here, we evaluated images generated by Stable Diffusion. The procedure to inspect these images for bias is similar to the dataset inspection: a synthetic image dataset is created, attribute correlations in it are calculated which are in turn evaluated for fairness, e.g. according to Def. 1. To investigate potential bias transfer between the training data and outcome (mitigation, reflection or amplification), we generated images using the same text prompt pp used for searching the dataset. Similarly, we again used the same classifier κ\kappa (Eq. 7) to determine label aga_{g} for the protected attribute in the generated images (g∈Gg\in\mathcal{G}) with ag=κ(g)a_{g}=\kappa(g).

Experimental Protocol

For the inspection, we created a new subset of LAION-5B with over 1.8 million images displaying humans with recognizable faces and in recognizable occupations (cf. Fig. 4a) and generated over 37,000 images each with SD and Fair Diffusion, respectively. In total, we evaluated more than two million images for gender occupation biases. Our instruction tool is built around SegaCode publicly available at https://github.com/ml-research/semantic-image-editing to edit images and guide the image generation toward fairer outcomes and employed FairFaceCode publicly available at https://github.com/joojs/fairface as κ\kappa to classify the protected attribute, i.e. facial (gender) attributes. Yet, Fair Diffusion can facilitate other image editing and classifying tools, too.

We employed CLIP to identify relevant images in LAION-5B—i.e., depicting people in recognizable occupations—and computed text-image similarities between LAION-5B images and a text prompt representing an occupation. To this end, we used p=“A photo of the face of a {occ}”p=\text{``A photo of the face of a \{{occ}\}''} as a text prompt and empirically determined a similarity threshold ρ=0.27\rho=0.27. We also used this prompt to generate images with SD, where occ∈{“firefighter”, “teacher”, “aide”, ...}\textit{occ}\in\{\text{``firefighter'', ``teacher'', ``aide'', ...}\}. The whole list consists of over 150 different occupationsTaken from https://huggingface.co/spaces/society-ethics/DiffusionBiasExplorer and we generated 250 images for each occupation prompt.

We made the assumption that an equal proportion of female- and male-appearing images is desired as derived from Def. 1. However, our evaluation is limited by current classifiers (like FairFace) facilitating only binary-valued gender classification, whereas gender is clearly non-binary (extensively examined in Discussion). Interestingly, Fair Diffusion can, in principle, be applied to non-binary gender identities and may also be used to realize different target distributions. Moreover, we employed a fair boundary (i.e., allow for a deviation of ±4%\pm 4\%) to soften this theoretical assumption. This way, we try to account for natural non-perfect binarity, i.e. the continuous spectrum of gender with its diversity and non-equal birth rate and world population.

We computed a per-group statistic (binary fe/male appearing) to further insight into the overall gender occupation bias in each component. Therefore, we divided the list of occupations into f and m, where the f-group denotes more female-biased occupations and the m-group otherwise. If the rate of female-appearing persons in LAION is > ⁣0.5>\!0.5 we use f and otherwise m, respectively. Subsequently, we evaluate these lists for each component and generate respective box plots. Without this group distinction, the average bias lies within the fair boundary (although the box plot shows high variance) as there are strong biases in both directions, which cancel each other out in an overall mean computation.

Acknowledgments

This work benefited from the ICT-48 Network of AI Research Excellence Center “TAILOR” (EU Horizon 2020, GA No 952215), the Hessian research priority program LOEWE within the project WhiteBox, and the Hessian Ministry of Higher Education, Research and the Arts (HMWK) cluster projects “The Adaptive Mind” and “The Third Wave of AI”, and from the German Center for Artificial Intelligence (DFKI) project “SAINT”.

References

Appendix

Apart from the results shown in the main text, we also generated more images, to provide insights into Fair Diffusion. Fig. 7 shows again that generated images of “firefighters” by default SD (top row) are strongly male-biased. In contrast, Fair Diffusion changes the outward appearance towards female-appearing “firefighters”. More interestingly, the changes in gender appearance do not change the overall image composition and the occupation remains identifiable. We regard this as a very powerful property of our approach.

Furthermore, in Fig. 8, we show results for the prompt “A photo of a woman”. Usually, the generated images by default SD represent people with Caucasian appearance. Here, we generated images with Fair Diffusion by instructing with +“Asian”, +“Indian”, +“African”, +“European”, +“Middle Eastern”, and + “Latino (Hispanic)”. These instructions are taken from FairFace’s “race” class. One can observe that the outward appearance changes according to the instruction given. Interestingly, one can observe that the instruction refers to multiple features, like hair color and style, lip color, and shapes of nose, cheek, and chin. This figure illustrates the potential capabilities of Fair Diffusion and should motivate further research in more diverse image generation beyond the presented gender biases.

We found further discriminating behavior in default SD, shown in Fig. 9. As one can observe, SD tends to generate images in line with heteronormativity and images of younger persons. Instead, Fair Diffusion can be employed again to generate homosexual couples and people of different ages. Please note that these images are only illustrative to show the potential of Fair Diffusion. As elaborated in the Discussion section, certain dangers also come along. For example, “homosexual” and “gay” often generated male-homosexual couples. This might be due to the fact that there is a specific word for female homosexuality, i.e., lesbian, and with +“lesbian”, it is also possible to generate female couples. However, these results are preliminary, showing avenues for future research.We down-scaled all images to a smaller and easier-to-handle size. Higher-resolution images can be generated with our code.

B A more detailed Inspection of CLIP Biases

In this experiment, we tested the similarity between images of different concepts. In line with , we applied their setup to the CLIP encoder and extended their occupation experiments with engineering and caregiving. Therefore, we added the first 12 images from Google search, that did not contain people, to the set of evaluation images of . All images can be found in our code baseanonymous link to reproduce the results.

Besides the results shown in Tab. 1 we show further results in Tab. 2. One can observe further cultural and racial biases for people looking Arab-Muslim or African-American as CLIP relates them to unpleasant. Accordingly, an intersectionality bias is present in CLIP too, amplifying the gender occupation bias. In other words, in CLIP, some people are confronted with multiple factors of advantage (e.g. White male) or disadvantage (e.g. Black female). Overall, we find that CLIP is inherently affected by bias too, and can be attributed as a source for the bias shift between LAION-5B and SD.

C Different Setups for Fair Diffusion

In addition to the results shown in Fig. 4, we investigated different setups for Fair Diffusion. The edit instructions used in the main text (-“female person” +“male person”, and with switched signs) represent a limitation, as examined in Discussion. Hence, we also examined other setups shown in Figs. 11, 10, 12 and 13. The applied edit instructions are given below each figure. In general, one can observe that the overall performance of Fair Diffusion for the different edit instructions presented here is worse than for the instructions used in the main text.

However, we cannot clearly attribute the loss in performance to Fair Diffusion. For one, the measured rate of female-appearing persons depends on the FairFace classifier. In case the edit instructions led to less clearly identifiable generated persons, FairFace struggled to classify them correctly and had high uncertainty. Furthermore, the success of the edit instruction depends on the parameters used. We used rather low parameters to not alter the image too strongly. If one increases the parameters, the changes are enforced stronger, at the expense of more substantial changes to the re-generated image. On the other hand, the instructions given to Fair Diffusion have to be known by the model. SD seems to have only a little understanding of e.g. the words “non-binary” (Fig. 13) and “gender” (Fig. 12). Thus, it is difficult to appropriately use these concepts for steering the image toward fairer outcomes. Furthermore, Figs. 10 and 13 show that fair instructions should be distinctive. If they contain similar concepts, they might interfere with each other. Lastly, (Fig. 11) shows that positive guidance alone is insufficient compared to positive and negative but still achieves remarkable performance. More research is needed here to investigate these findings further.

In Fig. 15, we compare three different image editing techniques for diffusion models. Fair Diffusion is agnostic to the underlying method and able to integrate various approaches –thus also the three shown. However, the capabilities of the methods differ. One can observe in this exemplary comparison, that only Sega is able to preserve the image composition and edit the gender appearance in isolation. The other methods in contrast change the overall image composition and struggle with addressing the gender appearance.

D Further Experimental Details

As described in our experimental evaluation, we also stumbled upon some challenges. For example, it was difficult to generate images for the occupation “dishwasher”. In Fig. 14, one can observe that “face of a dishwasher” is very ambiguous and mainly yielded results of the front side of dishwashing machines. Hence, further prompts beyond “A photo of the face of a occocc” should be evaluated in future research.

We hand selected subgroups “science”, “arts”, “engineering”, and “caregiving”. The selection can be found in Tab. 3.