Debiasing Vision-Language Models via Biased Prompts

Ching-Yao Chuang, Varun Jampani, Yuanzhen Li, Antonio Torralba, Stefanie Jegelka

Introduction

Foundation vision-language models, such as CLIP , DALLE-2 , Imagen , and Stable Diffusion , which are trained on extensive multimodal data at a massive scale, have led to a significant shift in the landscape of machine learning systems. Specifically, contrastive vision-language encoders like CLIP have the ability to perform zero-shot inferences without fine-tuning, and language embeddings can be used to train high-quality text-to-image models .

While vision-language models demonstrate impressive capabilities, it is important to recognize that they may also exacerbate biases . Recent studies have shown that the datasets these models are trained on can contain inappropriate image-text pairs with stereotypes, racist content, and ethnic slurs. The biases are then propagated to downstream applications , resulting in biased predictions. In addition to social biases, zero-shot models derived from vision-language models can also suffer from more general forms of spurious correlations such as image background, leading to poor group robustness . Biases also exist in generative models, where generated images may exhibit bias towards certain genders and races . Substantial progress has been made recently toward mitigating biases in vision-language models . However, many current approaches for addressing bias in models require training or fine-tuning the models using resampled datasets or modified objectives, which can be computationally intensive for foundation models.

In this work, we propose a general approach for self-debiasing foundation vision-language models by projecting out biased directions in the text embedding. Given a vision-language encoder such as CLIP, we define a set of biased directions in the embedding using prompts that describe the biases. For instance, prompts like “a photo of a male/female” define a biased subspace in the latent space. One approach to mitigating these biases is to construct a projection matrix, a linear transformation of the text embedding that projects out the biased directions . However, solely relying on prompts to define biased directions may be unstable and noisy . To address this issue, we propose a calibration loss that minimizes the discrepancy of a pair of prompt embeddings. For example, given a projection matrix that removes gender information, the projected vectors of prompts “a photo of a male doctor” and “a photo of a female doctor” should be similar. Based on this principle, we design an objective to calibrate the projection matrix, which has an easily solvable closed-form solution. This allows for the construction of the projection matrix to be training-free and requires no downstream dataset or labels, making it suitable for large-scale models. Empirically, we find that debiasing only the text embedding with a calibrated projection matrix suffices to improve the group robustness of zero-shot models on well-established benchmarks.

We then extend our approach to generative models such as Stable Diffusion , a widely adopted text-to-image model conditioned on text embeddings from CLIP . The inherent challenge lies in the fact that generative models are distinctly dissimilar from zero-shot classification, where the target classes are explicitly defined. With generative models, our objective is to develop a debiasing matrix that is universally applicable to every prompt. This matrix can then be employed as a standard preprocessing stage prior to feeding the text embedding into the generative model. To accomplish this, we solve the calibration matrix with a set of positive pairs which comprise various prompts from the training dataset, and debias the unseen prompts with the obtained matrix. Similar to debiasing zero-shot models, the projection matrix improves the diversity of generated images from text-to-image models without altering the model parameters.

In short, this work makes the following contributions:

We present a simple and general approach for debiasing vision-language models;

The proposed approach does not require training, data, or labels, making it computationally efficient for use with foundation models;

We evaluate our approach through experiments on both discriminative (zero-shot, text-image retrieval) and generative (diffusion) vision-language models.

Related Works

Vision-Language models have become increasingly widespread in recent years. However, these models are known to suffer from spurious correlations and can be biased towards certain races and gender. Birhane et al. study the datasets these models are trained on and show that their biases can be inherited by the models. Various methods have been proposed to address biases, but many of them only address single-modality models.

Large-scale language models have been shown to contain harmful or misrepresentative biases . Previous research has demonstrated the presence of gender bias in natural language processing systems as well as racial bias . Bolukbasi et al. first proposed the use of orthogonal projections to remove gender biases in word embeddings. This approach was later extended to debiasing sentence embeddings . Alternative methods include regularizing the models with constraints on training data or directly modifying the dataset . However, scaling these approaches to large foundation models can be challenging as they often require retraining the backbone encoders.

Biases in Vision Models

Gender and racial biases have also been widely explored in computer vision , in terms of dicsriminative models and generative models . Many debiasing approaches aim to learn good representations via adversarial training , or augmenting the biased dataset . Beyond social bias, many works study spurious correlations, a more general form of bias that can include features such as image background or other non-target attributes that are correlated with labels. This problem of spurious correlations is often studied and tackled as a group robustness problem . Kirichenko et al. show that last layer re-training is sufficient for robustness to spurious correlations, which aligns with our finding that debiasing the zero-shot weights suffices to yield robust classifiers.

Biases in Vision-Language Models

Recently, biases in multimodal settings have gained significant attention . Wang et al. propose to remove dimensions in the CLIP embedding that are highly correlated with gender attributes. Berg et al. debias the CLIP models with prompt learning via an adversarial approach. Seth et al. learn additive residual image representations to offset the biased representations. Recently, Zhang and Ré address the group robustness of vision-language models with contrastive learning. These previous works are data-oriented, where models are trained or finetuned on labeled datasets. In contrast, our approach is fully zero-shot, which does not require any downstream dataset and model training. To debias generative models, a recent work pre-defines a look-up table to provide fair guidance for text-to-image diffusion models. Nevertheless, this method encounters limitations when faced with previously unseen classes that are absent from the look-up table, while our approach generalizes well to new concepts.

Biases and Spurious Correlations

We consider a dataset in which each input x∈Xx\in{\mathcal{X}} is associated with multiple attributes, including the target class y∈Yy\in{\mathcal{Y}} and a spurious attribute a∈Aa\in{\mathcal{A}}. We focus on the case where biases are present and the attribute aa is spuriously correlated with the label yy. For instance, the class “doctor” could be correlated with the spurious attribute “gender” in the datasets foundation models are trained on . Importantly, these biases can be transferred to downstream tasks, both discriminative and generative.

The definition of metrics for text-image retrievals, such as maximal skewness , will be deferred to the experiment section.

Generative Models

A text-to-image model learns a conditional distribution P^(X∣Z=z)\hat{P}(X|Z=z), where zz is the embedding of the prompt. However, the biased nature of the dataset used to train the generative model can affect the distribution P^\hat{P}. To measure the bias present in generative models, recent works propose using statistical parity. Specifically, given a classifier h:X→Ah:{\mathcal{X}}\rightarrow{\mathcal{A}} for the spurious attribute, the discrepancy of the generative distribution P^\hat{P} is defined as the L2 norm between empirical and uniform distributions :

In practice, the expectation is estimated with empirical samples. A fair generative model minimizes the discrepancy by ensuring that each attribute a∈Aa\in{\mathcal{A}} has an equal probability (uniformly distributed).

Debiasing Discriminative Models

It is essential for a robust classifier to evade dependence on irrelevant features present in images. This necessitates the classifier to be invariant to image backgrounds and insensitive to attributes such as race or gender. Prior research has employed datasets with target labels and spurious attributes to quantify and eliminate biases . However, this approach is not feasible in a zero-shot setting, where data and training are prohibitive.

In contrast to previous approaches, our proposed method for measuring biases utilizes prompts, drawing inspiration from studies on debiasing word embeddings . The use of vision-language contrastive training allows for the description of irrelevant features through natural language. As such, embeddings of prompts such as “a photo of a [irrelevant attribute]" can capture these spurious features in the visual embedding. Consequently, the bias of a classifier can be quantified by computing the cosine similarity between its weights and the corresponding spurious feature. Table 1 illustrates the cosine similarity between the embeddings of prompts that describe the target classes and irrelevant attributes, using two popular group robustness benchmarks: Waterbird and CelebA . The details of datasets and the specific prompts can be found in section 5 and appendix C.1. The results demonstrate that the classifier weights are inclined towards certain irrelevant attributes (gender or image background), implicitly implying that the classifiers are using these spurious directions to make predictions.

2 Debiasing via Orthogonal Projection

We can use the projection matrix to eliminate spurious directions in a text embedding zz as P0zP_{0}z.

3 Calibrating the Projection Matrix

It is essential to acknowledge that the estimation of the irrelevant feature directions may introduce an approximation error in the projection matrix . Additionally, in certain scenarios, it may be challenging to thoroughly describe the irrelevant attribute using a limited number of prompts, resulting in increased uncertainty in the projection matrix estimation. This issue is also evident in our empirical results (Table 2 and 4), where the use of orthogonal projection fails to enhance performance.

To improve the estimation of the projection matrix, we leverage positive pairs of prompts that are expected to have the same semantic meaning after projection. In particular, the embedding of prompts such as “a photo of a [class name] with [spurious attribute]” should only contain information about “[class name]” after projecting out the spurious information, as Figure 1 illustrates. Motivated by this intuition, we propose to regularize the difference between the projected embeddings using a set of positive pairs SS:

where (zi,zj)(z_{i},z_{j}) is the embedding of pair (i,j)(i,j) in S{\mathcal{S}} and (i,j)(i,j) are prompts that describe the same class but different spurious attributes. The loss encourages the linear projection PP to be invariant to the difference between (i,j)(i,j), i.e., the spurious attributes. The optimization problem has a convenient closed-form solution, as demonstrated in Lemma 4.1.

We can see that U(I+λ′Σ2)−1UTU(I+\lambda^{\prime}\Sigma^{2})^{-1}U^{T} acts as a calibration term. Before multiplying the text embedding with the projection matrix P0P_{0}, variation due to the change of the spurious feature, namely, the eigenvectors with large squared singular value in ZdiffZ_{\textnormal{diff}} (spurious direction) will be down-weighted due to the inverse (I+λ′Σ2)−1(I+\lambda^{\prime}\Sigma^{2})^{-1}. Therefore, varying the spurious attributes should result in similar embeddings after multiplying the calibration matrix.

4 Relation to an Equalization Loss

The loss encourages the embedding zz to have similar cosine similarity to embeddings in positive pairs while maintaining proximity to the initialization z0z_{0}. Objective (4) has the same optimal solution as the calibration loss (3).

In particular, we have P0z∗=P∗z0P_{0}z^{\ast}={P^{\ast}}z_{0} where P∗P^{\ast} is the minimizer of the calibration loss (3).

Lemma 4.2 shows that the optimal solution of (4) is equivalent to multiplying the original embedding zz with the calibration matrix defined before. Applying the projection P0P_{0} to z∗z^{\ast} leads to the same weight in Lemma 4.1. This interpretation is particularly useful in cases where the ideal solution does not lie in the middle of ziz_{i} and zjz_{j}, as will be shown in section 6 where we address biases in generative models.

The equalization objective has a similar motivation as the equalization step proposed by Bolukbasi et al. in their work on removing gender bias from word embeddings. Similar to the idea of positive pairs, given a set of word embeddings that has the same semantic meaning except for gender, their approach centers these embeddings by setting them to the average embedding of the set. After centering, any word in the dictionary will be equidistant to all words in the set. However, our approach differs in that we modify the embedding of the target prompt zz, rather than the embedding of positive pairs, making it more suitable for debiasing zero-shot classifiers as we are primarily concerned with the embedding of zz.

Experiments: Discriminative Models

By following the setting of Zhang and Ré , we evaluate our approach on two popular benchmarks for evaluating spurious correlations, Waterbird and CelebA . On Waterbird, a water/land background is a confounding factor for the waterbirds/landbirds class, while on CelebA the binary gender is the spurious feature for blond/dark hair. Therefore, both datasets contains four groups defined by the labels and the spurious attributes. As such, both datasets contain four groups defined by the labels and the spurious attributes.

We evaluate our approach against several baselines, including zero-shot classification , empirical risk minimization (ERM) with linear probing , and ERM with non-linear adapter . Additionally, we also consider three recent methods designed to improve the group robustness of vision-language foundation classifiers:

Weight Space Ensembling (WiSE-FT) , which trains a linear classifier first using ERM and then combines the classifier outputs with the initial zero-shot predictions;

Deep Feature Reweighting (DFR) , which trains a linear probe on embeddings obtained from a pre-trained model using group-balanced data. Following Zhang and Ré , the group labels are replaced with zero-shot predictions;

Contrastive Adapter (CA) , which trains adapters using contrastive learning to bring embeddings in the same class closer.

It is important to note that all of the baselines except the zero-shot classifier require at least training data and class labels, while our debiasing approach does not require access to any input data, labels, or group labels, which follows the principles of zero-shot learning.

We evaluate the performance of our proposed approach using two CLIP backbones: ResNet-50 and ViT-L/14 . The results are presented in Table 2. The results indicate that a simple application of the orthogonal projection (Orth-Proj) by itself only yields limited improvement of the worst group accuracy, whereas the calibration loss (Orth-Cali) significantly improves robustness across datasets and base models. The proposed Orth-Cali method achieves comparable or even smaller gaps between average and worst group accuracy compared to the state-of-the-art contrastive adapter , without the need for any data or labels. Note that the baselines generally achieve better average accuracy as they require fine-tuning on the target datasets.

Empirically, we found that gradually increasing the parameter λ\lambda improves the worst group accuracy and leads to a stable solution as shown in Table 3. Therefore, for all the experiments on discriminative models, we set λ\lambda to 10001000 by default. To investigate the importance of orthogonal projection and calibration, we present an ablation study in Table 4. The results indicate that the calibration loss alone (P0=IP_{0}=I) performs well on the CelebA dataset, as the spurious feature (gender) is relatively easy to describe with prompts. However, performance drops on the Waterbird dataset without a good initialization from the orthogonal projection. More ablation studies can also be found in Appendix D, where we demonstrate the importance of class names in positive pairs.

2 Debiased Information Retrieval

Fairness in text-image retrieval has gained increasing attention in recent years. Building on the work of Berg et al. , we propose to utilize the MaxSkew metric, introduced by Geyik et al. , to evaluate the level of fairness in the retrieval results. Specifically, we conduct our analysis on the FairFace dataset , which is specifically designed to address issues of fairness in facial recognition systems. Given a ranked list of images in response to a text query, let ra,kr_{a,k} be the ratio of the top k images that are labeled with attribute aa. Then MaxSkew@k is defined as max⁡a∈Alog⁡ra,k1/∣A∣\max_{a\in{\mathcal{A}}}\log\frac{r_{a,k}}{1/|{\mathcal{A}}|}. It quantifies the maximal discrepancy between the ratio of top k images labeled with a specific sensitive attribute, denoted as ra,kr_{a,k}, and the uniform weight 1/∣A∣1/|{\mathcal{A}}|, where A{\mathcal{A}} represents the set of sensitive attributes. The MaxSkew metric provides a useful measure of fairness in text-image retrieval systems, by assessing the degree to which the retrieval results are evenly distributed across sensitive attributes. A small MaxSkew value indicates that the distribution of retrieved images across different sensitive attributes is close to being uniform.

To measure the bias, we query the validation set of FairFace based on 10 prompts that are uncorrelated with facial expressions or sensitive attributes, e.g., “a photo of a [concept] person”, where the [concept] is a neutral concept such as evil or smart. The detailed prompts are described in Appendix C. We measure the MaxSkew based on three labeled attributes of FairFace: gender, race, and age. Table 5 shows the average MaxSkew@1000 over concepts, demonstrating that our approach significantly reduces the MaxSkew across different attributes and backbones.

Debiasing Generative Models

We now explore the possibility of extending the methodology developed for discriminative models to generative models. Our primary focus is on addressing social group biases, specifically gender and race discrepancy, as measured by metric (1). In particular, the main experiment is to query the generative model using profession-related prompts, specifically “a photo of a [profession]". Empirically, the generated images were found to exhibit a strong bias towards certain gender and race, and we attempt to improve the diversity of generated images with the proposed equalization loss in this section. We also demonstrate that our approach can also address spurious correlations beyond social biases.

Unlike the well-defined targets prevalent in zero-shot classification, the nature of generative models requires a more universal solution. Specifically, we seek to derive a debiasing matrix capable of accommodating any prompt. This matrix could subsequently be treated as a standardized preprocessing step, applied prior to the introduction of the embedding into the generator.

To achieve this, we optimize the equalization loss with positive pairs consisting of an enumeration of “a photo of a [attribute] [profession]” where the [attribute] is a member of the set of gender or races and the [profession] is a job title sampled from a training set. For instance, to mitigate gender bias, we adopt S={\mathcal{S}}= {\{(“a photo of a male doctor”, “a photo of a female doctor”), ⋯\cdots, (“a photo of a male engineer”, “a photo of a female engineer”) }\}. By solving the calibration matrix with professions in the training set, we expect the obtained matrix can also mitigate the biases in unseen professions. Note that we optimize the equalization loss (4) without applying the initial orthogonal projection matrix P0P_{0}. This is because our goal is to balance rather than completely eliminate biased information in the generated images.

Experiments: Generative Models

To evaluate the effectiveness of our approach in the context of generative models, we conducted experiments using the Stable Diffusion (SD) v2.1 framework . We construct a list of professions that consists of 100 job titles with GPT-4 and randomly separate them into 80 training and 20 testing professions. The complete list can be found in appendix C. In alignment with the framework proposed by Kärkkäinen and Joo , we consider the gender attributes of male and female, and racial attributes of White, Asian, Black, Indian, and LatinoIt is essential to recognize gender and race are complex social constructs that cannot be simply reduced to binary or discrete categories. The choice of using binary gender and discrete race attributes in our work was primarily based on the existing literature and benchmark datasets that have commonly adopted this setting for evaluation purposes..

Evaluating generative models can be challenging without the use of human labels. Inspired by Cho et al. , we used sensitive attribute classifiers to predict the sensitive attributes of the generated images. The discrepancy, as defined in equation (2), was then calculated. In particular, we generate 100 images for each train / test profession for evaluation, resulting in 10000 images for each model. We then leverage the CLIP classifier to predict the sensitive attributes to calculate the discrepancy. An alternative to CLIP is the FairFace classifier ; however, we found that the domain shift between the FairFace dataset and the generated images significantly impairs its performance. The debiased and biased models share the same random seed for fair comparison. We set λ=500\lambda=500 for all the experiments in this section.

We first examine whether minimizing the calibration loss with training prompts can yield a calibration matrix that also works for unseen (testing) professions. In particular, we measure the average L2 difference between the projected embedding ∑(i,j)∈Stest∥Pzi−Pzj∥/∣Stest∣\sum_{(i,j)\in{\mathcal{S}}_{\textnormal{test}}}\left\|Pz_{i}-Pz_{j}\right\|/|{\mathcal{S}}_{\textnormal{test}}| for testing prompts and show the results in Table 6. We can see that the calibration matrix successfully minimizes the difference after projection, even for unseen professions.

2 Quantitative and Qualitative Results

The results presented in Table 7 demonstrate a significant reduction in both gender and race discrepancy after debiasing. Importantly, the improvements are observed for both training and testing professions, implying that the obtained debiasing matrix can generalize beyond training prompts. To further illustrate the effectiveness of our approach, we present quantitative results for mitigating gender bias in Figure 2. By applying the calibration matrix to balance the male and female directions, the gender diversity of the generated images significantly improved. Additional examples can be found in appendix D.

Compared to gender bias, we found that addressing racial bias is a more challenging task. One source of complexity is the ambiguity of ethnicity, as individuals may identify with multiple races. Nevertheless, as Figure 3 and Table 7 demonstrate, the diversity in the output images is improved by simply debiasing the prompt embedding with the calibration matrix.

3 Human Evaluation

Despite the scalability, the prediction from a trained classifier could be erroneous. Therefore, we also evaluate our approach with human evaluation, where we invite annotators of different genders, races, and nationalities to label the sensitive attributes of the generated images. Details and the interface are included in appendix C.2. For human evaluation, we generate 25 images for each test profession, resulting in 500 images for each model. As Table 8 shows, our approach greatly improves the diversity of the generated images, corroborating the previous results.

4 Beyond Social Biases

Our approach can also be applied to address general spurious attributes beyond social biases. As an example, we draw inspiration from the WaterBird dataset and debias the prompt “a photo of a waterbird” by using {\{“a photo of a [animal] with water background” and “a photo of a [animal] with land background” }\} as positive pairs, where we construct a list of 100 names of animals with GPT-4 .

As Figure 4 illustrates, our approach successfully generates images of water birds in both land and water backgrounds, whereas the original models only generated images with water background.

Conclusion

In this work, we present a new approach to debiasing vision-language foundation models by utilizing prompts to mitigate biases. The proposed calibrated projection effectively mitigates biases in both discriminative and generative vision-language models without any additional training or data.

Thanks to Arjun Akula, Susanna Rico, Joshua Robinson, Lucy Chai, Kabir Swain, Manel Baradad, Joanna Materzynska, Shobhita Sundaram, Pei-Ling Chiang, and Yi-Yi Chu for their helpful comments and suggestions. This work was in part supported by NSF BIGDATA IIS- 1741341, NSF CAREER 1553284, and NSF AI Institute TILOS. CYC is supported by an IBM PhD Fellowship.

References

Appendix A Broader Impact

The development and implementation of debiasing techniques in vision and language models has the potential to significantly impact a wide range of industries and applications. By reducing the biases in the models, they will be better able to accurately recognize and understand diverse individuals and groups, leading to more fair and equitable decision-making in fields such as education, employment, and law enforcement. Nevertheless, our approach also has limitations. For instance, the proposed debiased technique for generative models does not work for certain classes or biases. Despite the limitations, our work on debiasing vision and language models is a crucial step towards creating more inclusive and fair technology for all.

Appendix B Proof

We will leverage the first order optimality criteria to derive the solution.

Setting the derivate w.r.t. P to zero yields:

Note that two optimums are equivalent, where the second one is simply the matrix form of the first. ∎

B.2 Lemma 4.2

Similarly, the objective can be rewritten as

Appendix C Experiment Details

In this section, we provide the exact prompt we use for all the experiments in the paper in Table 9, 10, 11, 12.

C.2 Human Evaluation

We generate 100 images for each profession for evaluation. Therefore, there are 1500 images for each model in total. The random seed is fixed for the original and debiased Stable Diffusion models. In particular, automatic evaluation and human evaluation adopt the same set of images for fair comparison. The interface for human evaluation is shown in Figure 5. Note that some generated images might be corrupted, or do not even contain humans. In this case, the annotators can click 3 or 6 to indicate that the current image is not identifiable. We remove these images while calculating the discrepancy.

Appendix D More Experiment Results

In this section, we study the importance of class names in the positive pairs. In particular, instead of using “a photo of a [class name] with [spurious attribute]”, we instead use “a photo of a [spurious attribute]” to estimate the calibration matrix. The results are shown in Table 13. We can see that the performance significantly drops after removing the class name from the prompt, emphasizing the importance of class-conditioned prompts.

D.2 More Samples from Biased and Debiased Generative Models

In this section, we show more generated images from Stable Diffusion 2.1 to provide a qualitative experiment. We can see that the proposed debiasing approach significantly improves the diverisity across training and testing professions as as Figure 6, 7, 9 and 10 show. Nevertheless, there are also failure cases, where both our approach and the original model fail. For instance, biased and debiased models fail to generate females for many engineer-related professions such as carpenter as Figure 8 shows.