Is Fairness Only Metric Deep? Evaluating and Addressing Subgroup Gaps in Deep Metric Learning
Natalie Dullerud, Karsten Roth, Kimia Hamidieh, Nicolas Papernot, Marzyeh Ghassemi
Introduction
Deep metric learning (DML) extends standard metric learning to deep neural networks, where the goal is to learn metric spaces such that embedded data sample distance is connected to actual semantic similarities (Globerson & Roweis, 2006; Weinberger et al., 2006; Hoffer & Ailon, 2018; Wang et al., 2014). The explicit optimization of similarity makes deep metric spaces well suited for usage in unseen classes, such as zero-shot image or video retrieval or facial re-identification (Milbich et al., 2021; Roth et al., 2020c; Musgrave et al., 2020; Hoffer & Ailon, 2018; Wang et al., 2014; Schroff et al., 2015; Wu et al., 2018; Roth et al., 2020c; Brattoli et al., 2020; Hu et al., 2014; Deng et al., 2019; Liu et al., 2017). However, while DML is effective in establishing notions of similarity, work describing potential fairness issues is limited to individual fairness in standard metric learning (Ilvento, 2020), disregarding embedding models.
Indeed, the impacts and metrics of fairness are well studied in machine learning (ML) generally, and representation learning specifically (Dwork et al., 2012; Mehrabi et al., 2019; Locatello et al., 2019b). This is especially true on high-risk tasks such as facial recognition and judicial decision-making (Chouldechova, 2017; Berk, 2017), where there are known risks to minoritized subgroups (Samadi et al., 2018). Yet, relatively little work has been done in the domain of DML (Rosenberg et al., 2021). It is crucial to address this knowledge gap – if DML embeddings are used to create upstream embeddings that facilitate downstream transfer tasks, biases may propagate unknowingly.
To tackle this issue, this work first proposes a benchmark to characterize fairness in non-balanced DML - finDML. finDML introduces three subgroup fairness definitions based on feature space performance metrics – recall@k, alignment and group uniformity. These metrics measure clustering ability and generalization performance via feature space uniformity. Thus, we select the metrics for our definitions to enforce independence between inclusion in a particular cluster or class, and a protected attribute (given the ground-truth label). We leverage existing datasets with fairness limitations (CelebA (Liu et al., 2015) and LFW (Huang et al., 2007)) and induce imbalance in training data of standard DML benchmarks, CARS196 (Krause et al., 2013) and CUB200 (Wah et al., 2011), in order to create an effective benchmark for fairness analysis in DML.
Making use of finDML, we then perform an evaluation of 11 state-of-the-art (SOTA) DML methods representing frequently used losses and sampling strategies, including: ranking-based losses (Wang et al., 2014; Hoffer & Ailon, 2018), proxy-based (Kim et al., 2020) losses, semi-hard sampling (Schroff et al., 2015) and distance-weighted sampling (Wu et al., 2018). Our experiments suggest that imbalanced data during upstream embedding impacts the fairness of all benchmarks methods in both upstream embeddings (subgroup gaps up to 21%) as well as downstream classifications (subgroup gaps up to 45.9%). This imbalance is significant even when downstream classifiers are given access to balanced training data, indicating that data cannot naively be used to de-bias downstream classifiers from imbalanced embeddings.
Finally, inspired by prior work in DML on multi-feature learning (Milbich et al., 2020), we introduce PARtial Attribute DE-correlation (PARADE). PARADE addresses imbalance by de-correlating two learned embeddings: one learnt to represent similarity in class labels, and one learnt to represent similarity in the values of a sensitive attribute, which is discarded at test-time. This creates a model in which the ultimate target class embeddings have been de-correlated from the sensitive attributes of the input. We note that as opposed to previous work on variational latent spaces, PARADE de-correlates a learned similarity metric. We find that PARADE reduces gaps of SOTA DML methods by up to 2% downstream in finDML.
In total, our contributions can be summarized as follows:
We define finDML; introducing three definitions of fairness in DML to capture multi-faceted minoritized subgroup performance in upstream embeddings through focus on feature representation characteristics across subgroups, and five datasets for benchmarking.
We analyze SOTA DML methods using finDML, and find that common DML approaches are significantly impacted by imbalanced data. We show empirically that learned embedding bias cannot be overcome by naive inclusion of balanced data in downstream classifiers.
We present PARADE, a novel adaptation of previous zero-shot generalization techniques to enhance fairness guarantees through de-correlation of class discriminative features with sensitive attributes.
Background
DML extends standard metric learning by fusing feature extraction and learning a parametrized metric space into one end-to-end learnable setup. In this setting, a large convolutional network provides the mapping to a feature space , while a small network , usually a single linear layer, generates the final mapping to the metric or embedding space . The overall mapping from the image space is thus given by . Generally, the embedding space is projected on the unit hypersphere through normalization (Weisstein, 2002; Wu et al., 2018; Roth et al., 2020c; Wang & Isola, 2020) to limit the volume of the representation space with increasing embedding dimensionality. The embedding network is then trained to provide a metric space that operates well under some predefined, usually non-parametric metric such as the Euclidean or cosine distance defined over .
Typical objectives used to learn such metric spaces range from contrastive ranking-based training using tuples of data, such as pairwise (Hadsell et al., 2006), triplet- (Schroff et al., 2015; Wu et al., 2018) or higher-order tuple-based training (Sohn, 2016; Wang et al., 2020a), procedures to bring down the effective complexity of the tuple space (Schroff et al., 2015; Harwood et al., 2017; Wu et al., 2018) or the introduction of learnable tuple constituents (Movshovitz-Attias et al., 2017; Qian et al., 2019; Kim et al., 2020).
More recent work (Milbich et al., 2020; Roth et al., 2020c; Jacob et al., 2019) extends standard DML training through incorporation of objectives going beyond just sole class label discrimination: e.g., through the introduction of artificial samples (Lin et al., 2018; Duan et al., 2018), regularization of higher-order moments (Jacob et al., 2019), curriculum learning (Zheng et al., 2019; Harwood et al., 2017; Roth et al., 2020a), knowledge distillation (Roth et al., 2020b) or the inclusion of additional features (DiVA) to produce diverse and de-correlated representations (Milbich et al., 2020).
DML Evaluation
Standard performance measures reflect the goal of DML: namely, optimizing an embedding space for best transfer to new test classes via learning semantic similarities. As immediate applications are commonly found in zero-shot clustering or image retrieval, respective retrieval and clustering metrics are predominantly utilized for evaluation. Recall@k (Jegou et al., 2011) or mean average precision measured on recall (Roth et al., 2020c; Musgrave et al., 2020) typically estimate retrieval performance. Normalized mutual information (NMI) on clustered embeddings (Manning et al., 2010) is used as a proxy for clustering quality (see Supplemental for detailed definitions). We leverage these performance metrics to inform finDML and our experiments.
Fairness in Classification
Formalizing fairness in ML continues to be an open problem (Mehrabi et al., 2019; Chen et al., 2018a; Chouldechova, 2017; Berk, 2017; Locatello et al., 2019b; Chouldechova & Roth, 2018; Dwork et al., 2012; Hardt et al., 2016; Zafar et al., 2017). In classification, definitions for fairness such as demographic parity, equalized odds, and equality of opportunity, rely on model outputs across the random variables of protected attribute and ground-truth label (Dwork et al., 2012; Hardt et al., 2016).
Fairness in Representations
A more relevant family of fairness definitions for DML would be those explored in fairness for general representation learning (Edwards & Storkey, 2015; Beutel et al., 2017; Louizos et al., 2015; Madras et al., 2018). Here, the goal is to learn a fair mapping from an original domain to a latent domain so that classifiers trained on these representations are more likely to be agnostic to the sensitive attribute in unknown downstream tasks. This assumption distinguishes our setting from previous fairness work in which the downstream tasks are known at train time (Madras et al., 2018; Edwards & Storkey, 2015; Moyer et al., 2018; Song et al., 2019; Jaiswal et al., 2019). DML differs from this form of representation learning as it aims to learn a mapping capturing semantic similarity, as opposed to latent space representation.
Earlier works in fair representation learning intended to obfuscate any information about sensitive attributes to approximately satisfy demographic parity (Zemel et al., 2013) while a wealth of more recent works focus on using adversarial methods or feature disentanglement in latent spaces of VAEs (Locatello et al., 2019a; Kingma & Welling, 2013; Gretton et al., 2006; Louizos et al., 2015; Amini et al., 2019; Alemi et al., 2018; Burgess et al., 2018; Chen et al., 2018b; Kim & Mnih, 2018; Esmaeili et al., 2019; Song et al., 2019; Gitiaux & Rangwala, 2021; Rodríguez-Gálvez et al., 2020; Sarhan et al., 2020; Paul & Burlina, 2021; Chakraborty et al., 2020). In this setting, the literature has focused on optimizing on approximations of the mutual information between representations and sensitive attributes: maximum mean discrepancy (Gretton et al., 2006) for deterministic or variational (Li et al., 2014; Louizos et al., 2015) autoencoders (VAEs); cross-entropy of an adversarial network that predicts sensitive attributes from the representations (Edwards & Storkey, 2015; Xie et al., 2017; Beutel et al., 2017; Zhang et al., 2018; Madras et al., 2018; Adel et al., 2019; Zhao & Gordon, 2019; Xu et al., 2018); balanced error rate on both target loss and adversary loss (Zhao et al., 2019); Weak-Conditional InfoNCE for conditional contrastive learning (Tsai et al., 2021).
PARADE shares aspects of these previous methods in its choice of de-correlation or disentanglement. However, PARADE de-correlates the learned similarity metric as opposed to the latent space. In addition, with DML-specific criteria, PARADE learns similarities over the sensitive attribute while not directly removing all information about the sensitive attribute, as the sensitive attribute and target class embeddings share a base network.
Extending Fairness to DML - finDML Benchmark
To characterize fairness with finDML, this section introduces the key constituents – definitions to characterize fairness in embedding spaces and respective benchmark datasets.
Our embedding space fairness definitions rely on embedding space metrics adapted from (Wang & Isola, 2020) and (Roth et al., 2020c), namely alignment and uniformity. Both metrics we use to characterize embeddings for our definitions in the next section (intra- as well as inter-class alignment and uniformity) have been successfully linked to generalization performance in contrastive self-supervised and metric learning models (Wang & Isola, 2020; Roth et al., 2020c; Sinha et al., 2020). Alignment succinctly captures the similarity structure learned by the representation space with respect to the target labels through measuring distances between pairs of samples. On the other hand, notions of uniformity can differ. Uniformity of the sample distribution over the hypersphere has been studied through the radial basis function (RBF) over pairs of samples. Alternatively, uniformity of the feature space has been studied through the KL-divergence between the discrete uniform distribution and the sorted singular value distribution of the representation space on dataset .
Here, lower scores indicate more significant directions of variance in learned representations. Both introduced notions of uniformity represent important aspects of the embedding space, but the computational overhead in computing RBF over all pairs of samples in large datasets makes it impractical for our uses and is less interpretable than . Therefore, we leave the uniformity metric utilized in finDML general, but utilize for our experiments.
2 Defining Fairness
Building on the aforementioned performance metrics, we introduce three definitions for fairness in the embedding spaces of DML models. As the recall@k and alignment metrics inform inclusion in an embedded cluster (or class), we follow fair classification literature in the motivation for our first fairness definition: inclusion in a class should be independent of a protected attribute given the ground-truth label. Thus, we examine the probability of encountering a data instance of the same class in a data point’s -nearest neighbors to form the first definition. The second definition relies on equal expectation of alignment across sensitive attribute values. Departing from classification literature, our third definition encapsulates fairness in a task-agnostic sense (as DML is often applied in such settings): fairness across the “goodness" of the learned features via a uniformity metric.
Define as a function that receives a point and returns a set in the powerset of , , containing points in that map to the nearest neighbors of in . Thus, is -close fair with respect to attribute if:
Note: the criteria weakens as increases, similar to recall@k.
is fair, according to alignment with respect to attribute , if:
i.e. the expectation of the alignment is equal across domain of .
is fair, according to uniformity, and with respect to attribute , if the expectation of the uniformity is equal across domain of :
where denotes some measure of uniformity over a set .
3 Constructed finDML benchmark datasets
finDML encompasses existing DML benchmark datasets, CUB200 and CARS196, and facial recognition datasets, CelebA and LFW (Wah et al., 2011; Krause et al., 2013; Liu et al., 2015; Huang et al., 2007). For fairness analysis, we investigate bird color in CUB200While bird color in CUB200 does not represent a real-world fairness setting, CUB200 is widely used as a DML benchmark. Thus, a fairness angle allows fairness analysis of previous methods benchmarked on CUB200., Race in LFW and Skintone in CelebA (Kumar et al., 2009). A detailed description of dataset and attribute labeling is included in the Supplemental. To create additional fairness benchmarks, we induce class imbalance in CUB200 and CARS196, as both datasets are naturally balanced w.r.t. class.
Manually Introduced Class Imbalance We introduce imbalance by reducing the number of training data samples of randomly selected classes by (Imbalanced). We run an experiment with the original datasets as a balanced control (Balanced) for comparison. In the imbalanced setting, we adjust (increase) the number of training samples of the majoritized groups to match the number of datapoints in the balanced control experiments. We average metrics over sets of randomly selected classes for imbalanced experiments. We use the standard ratio of for train-test split of these datasets, but split over number of data points per class, as opposed to splitting over the classes themselves. The manually imbalanced datasets are used to benchmark standard DML methods, validate our framework, and analyze downstream effects.
Although dataset imbalance does not constitute the sole source of bias in machine learning applications, unfairness as a result of imbalance is the most well-understood in the literature (Chen et al., 2018a). Additionally, we do not assume for our naturally imbalanced datasets, particularly the facial datasets, that attribute imbalance is the only source of bias we observe.
PARtial Attribute DE-correlation (PARADE)
In this section, we present Partial Attribute De-correlation, or PARADE, in which we incorporate adversarial separation (Milbich et al., 2020) during training to de-correlate separate embeddings. We enumerate several significant changes: 1) only target embedding released at test-time; 2) triplet formation and loss term w.r.t. sensitive attribute; 3) de-correlation with sensitive attribute as opposed to de-correlation to reduce redundancy in concatenated feature space. These two representations branch off from the deep metric embedding model at the last layer. The two representations encode the similarity metrics learned over the sensitive attribute and target class, respectively. The sensitive attribute embedding layer is discarded at test time. The resulting network expresses a similarity metric with respect to the target class, de-correlated from the sensitive attribute (Figure 1). Therefore, PARADE figuratively optimizes the first two fairness definitions proposed in Section 3.2 via an objective that maximizes independence between the sensitive attribute and target class.
To achieve efficient training and de-correlation of the target class and the sensitive attribute embedding layers, we simultaneously train both layers that branch from the penultimate layer of the model and de-correlate at each iteration. Because PARADE must learn one embedding w.r.t. target class () and one embedding w.r.t. the sensitive attribute (), we introduce separate objectives for each embedding:
where is the number of training triplet samples, and represents a generic loss function, such as triplet loss (Hoffer & Ailon, 2018). We use to illustrate sampling over triplets of the form where and are of the same target class and and are of differing target classes. Similarly, indicates sampling over triplets of the form where and are of the same sensitive attribute subgroup and and are of differing sensitive attribute subgroups. See Figure 2 for a t-SNE visualization of the distinct embeddings of PARADE.
Partial De-correlation
In order to minimize the correlation between and , we use the adversarial separation (de-correlation) method from (Milbich et al., 2020), which minimizes the mutual information between a pair of embeddings. The task of mutual information minimization is accomplished through learning an MLP to maximize the pair’s correlation, , and consequently performing a gradient reversal , which inverts the gradients during backpropagation. The MLP, , is trained to maximize , s.t. denotes element-wise multiplication. Combining the loss terms results in total loss:
where weights the sensitive attribute loss and weights the degree of de-correlation. modulates the de-correlation term to allow to retain some attribute information (i.e. partial de-correlation). Thus, the deployed model can retain information about the sensitive attribute in its feature representations, as appears in the loss function back-propagated through the full model . The extent to which the sensitive attribute affects the output features is controlled by ; we suggest optimizing and through maximization of worst-group performance (Lahoti et al., 2020) (See Supplemental C.5 for further analysis of PARADE hyperparameters).
Experiments
Baseline DML Methods For all datasets, we use a ResNet-50 (He et al., 2016) architecture with best performing parameters on a validation set (for further implementation details, see Supplemental). To investigate a sweeping set of frequently used DML methods, we benchmark across a diverse, representative set of 11 techniques, including: three standard ranking-based losses (margin, triplet, n-pair, and contrastive) three batch mining strategies (random, semi-hard and distance-weighted sampling) and three common loss functions (multisimilarity loss, ArcFace loss for handling facial datasets, and proxy-based loss, ProxyNCA) (Hoffer & Ailon, 2018; Hadsell et al., 2006; Wu et al., 2018; Sohn, 2016; Hadsell et al., 2006; Kim et al., 2020; Wang et al., 2020a; Deng et al., 2019; Wu et al., 2018; Schroff et al., 2015). See Supplementary for more details.
Fairness Evaluation In the embedding space, we analyze fairness via performance gaps between minoritized groups and majoritized groups, or worst-group performance gaps (Lahoti et al., 2020). For fairness of the feature representations, we compute gaps in three metrics: recall@k and NMI for intra- and inter-class distance (Section 3.2), and the uniformity measure corresponding to Definition 5 (defined in Section 3.1).
Training and Evaluation on Downstream Classifiers To link fairness performance in the embedding space to downstream classification (in which more extensive prior work has been completed), we train downstream classifiers and evaluate classification bias. After training the DML model with the aforementioned criteria, the network is fixed. The output embeddings from the image training datasets, in addition to the class labels, are used to train four downstream classification models: logistic regression (LR), support vector machine (SVM), K-Means (KM), and random forest (RF) (Pedregosa et al., 2011). In the manually imbalanced upstream setting, we train downstream classifiers on the original balanced image datasets to ascertain if bias incurred in the embedding can propagate downstream even if the downstream classifier is trained with real balanced data.
We execute class imbalanced experiments for CARS196 and CUB200 and vary the level of imbalance between minoritized and majoritized classes in the upstream training set.
PARADE Configuration We test Partial Attribute De-correlation, PARADE, by training models in the listed settings: manually color imbalanced dataset for CUB200, CelebA and LFW. The attribute used to train the sensitive attribute embedding for each dataset, and the attribute used for fairness evaluation. We compare PARADE with margin loss and distance-weighted sampling (Wu et al., 2018) to standard margin loss and distance-weighted sampling.
Results
Our experiments indicate that current DML methods encounter crucial fairness limitations in the presence of imbalanced training data. Table 1 (along with a corresponding table for CARS196 in the Supplemental) demonstrate that gaps in the manually class imbalanced setting are greater than the balanced control setting. In four combinations of loss functions and sampling strategies, we do not observe a scenario in which the class imbalanced setting achieves a smaller gap than the control in the embedding space, nor the downstream classification. This is particularly significant due to the nature of sampling strategies studied (Wu et al., 2018; Schroff et al., 2015), which batch samples to force the model to correct “hard" examples. The results validate finDML as a benchmark and framework for fairness through the lens of well-studied fairness characterization in classification.
Interestingly, Table 1 displays non-negligible gaps in downstream performance metrics recall and precision even in the balanced control case. This could represent stenography of underlying structures in the data, such as car color or bird size. More likely, however, these gaps are due to use of macro-averaging in recall and precision calculations. Nonetheless, the manually class imbalanced settings consistently produce larger gaps.
2 Propagation of bias to downstream tasks
The tabular results emphasize a significant result: naive re-balancing with real data downstream cannot overcome bias incurred in the upstream embedding in any setting studied. Indeed, Table 1 exhibits propagation of bias from upstream embeddings (trained on imbalanced data) to downstream tasks (trained on fixed upstream embeddings with a re-balanced dataset). To provide additional context for the result, we direct to increasing use of DML models as components of larger classification models. This trend is arising in literature such as supervised contrastive learning, and recent developments in pre-training and lifting DML models for classification (Khosla et al., 2020). This necessitates tackling bias in the representation space of DML as opposed to patches downstream, and emphasizes the importance of defining fairness in this setting as done in our work.
Figure 3 shows that gaps in downstream classification mimic those upstream, even as we vary the level of imbalance introduced when training the upstream embedding. Here, the random forest classifier sees greater gaps in downstream metrics than the control, even when manual imbalance is set at upstream, and the downstream training dataset is balanced. For results with additional downstream classifiers, see Supplemental. This experiment demonstrates that the propagation of bias to downstream will occur even with lower levels of imbalance, and does not appear to depend on the downstream classifier chosen.
3 Reduced subgroup gaps through partial de-correlation with sensitive attribute
Table 2(a) shows results for performance gaps between relevant subgroups in both facial recognition datasets. PARADE shows strong results for CUB200 bird color dataset, primarily reducing gaps downstream and accordingly to recall@1 (Definition 1). PARADE can reliably reduce gaps for both the representation space and downstream classifiers on LFW. Interestingly, we observe that the majoritized subgroup (“White") had worst performance of all “Race" subgroups (see Supplemental), contrary to previous results (Samadi et al., 2018).Note: minoritized subgroups can still encounter notable bias across other axes more difficult to measure (Radford & Espenshade, 2014). As such, we measure gaps between the worst-performing subgroup and others.
For CelebA, we find for standard methods the minoritized subgroups to generally perform worst. PARADE excels at gap reduction upstream but encounters larger subgroup gaps downstream compared to standard methods. While PARADE does reduce downstream gaps between light skintones (I, II, and III), and the two lighter dark skintones (IV, V), gaps increase between lighter skintones and the darkest skintone (VI) (see Supplemental). Because skintone VI constitutes % of the CelebA dataset, PARADE is likely not able to learn similarity between faces over attributes besides skintone. And PARADE is prevented from learning similarities based on skintone due to de-correlation. In such settings, PARADE could be combined with oversampling minoritized subgroups to ensure better performance.
In general, the results show promising benefits of PARADE to adequately address and improve on the challenge of subgroup gaps for DML models used in facial recognition; and in the standard DML dataset CUB200, for recall@1 upstream (Definition 1) and across metrics downstream (Table 2(b)).
Discussion
In this work, we introduce the finDML benchmark, a framework for fairness in deep metric learning (§3.2). We demonstrate the fairness limitations of established DML techniques, and the surprising propagation of embedding space bias to downstream classifiers. Importantly, we find that this bias cannot be addressed at the level of downstream classifiers but instead needs to be addressed at the DML stage. We investigate the limit of this propagation in manually introduced imbalance, and finally show that PARADE can reduce subgroup gaps in several settings.
PARADE suffers from pitfalls similar to other “fairness with awareness" methods: PARADE uses information only on pre-defined sensitive attributes and therefore can be unfair w.r.t. other sensitive attributes. PARADE does have an advantage in addressing the combinatorial number of attributes considered in multi-attribute fairness through DML, which will scale sub-combinatorially in time/space complexity. We also note that subgroup gaps are not sufficient to capturing societal understandings of fairness, and there is no consensus as to how to remedy such gaps (Chouldechova & Roth, 2018; Dwork et al., 2012; Hardt et al., 2016; Zemel et al., 2013; Zafar et al., 2017). Additionally, while PARADE intentionally optimizes Definitions 1 and 2, we provide no explicit guarantee and optimization of uniformity, Definition 5, remains an open problem. Finally, PARADE does incur slight decrease in overall performance, similar to other methods (Wick et al., 2019) (see Supplemental for per-subgroup performance and additional fairness-utility trade-off analysis for PARADE).
Code of Ethics Statement
The work presented here deals with fairness in deep metric learning. A portion of our studies in the paper focus on CARS196 and CUB200-2011 datasets, which have consistently been used in benchmarking novel DML frameworks (Krause et al., 2013; Wah et al., 2011). The fairness analysis considered for CUB200-2011 deals with bird color, which does not, to our knowledge, correspond with any societal problems relating to fairness. Nonetheless, as CUB200-2011 is used in a litany of papers for SOTA performance comparison, finDML includes CUB200 so that DML methods can be analyzed w.r.t. fairness on a dataset used in their original paper.
We do include facial recognition datasets and tasks and analyze fairness with respect to facial attributes. Facial recognition does raise ethical concerns in practice. We note that our paper attempts to address primary social concerns in facial identity recognition. We do not encourage the task of facial attribute recognition, and solely use labeled attributes that correspond to known axes of bias for fairness analysis (e.g. Race and Skintone). As PARADE has solely been tested in two widely used public facial recognition datasets, we cannot guarantee fairness nor privacy in practical settings with private facial datasets.
Reproducibility Statement
Additional experimental results discussed in the main paper and others are contained in Supplemental C. Implementation details including attribute information, generation of attributes, training parameters, metric calculation and gap computation are listed in Supplemental D. Code available here: https://github.com/ndullerud/dml-fairness.
Acknowledgements
We would like to acknowledge and thank our sponsors, who support our research with financial and in-kind contributions: CIFAR through the Canada CIFAR AI Chair, Intel, NFRF through an Exploration grant, NSERC through the COHESA Strategic Alliance, International Max Planck Research School for Intelligent Systems (IMPRS-IS), European Laboratory for Learning and Intelligent Systems (ELLIS) PhD program, Mitacs Accelerate through internship program, and Microsoft Research. Resources used in preparing this research were provided, in part, by the Province of Ontario, the Government of Canada through CIFAR, and companies sponsoring the Vector Institute. We would like to thank members of the CleverHans Lab and HealthyML Lab for their invaluable feedback.
References
Appendix A Additional Background
The batch sampling procedure in deep metric learning methods differ from that of generic deep classifiers in that canonical loss functions require tuples or pairs of samples in order to utilise ranking objectives as training surrogates to learn an appropriate similarity metric. To ensure that tuples with positive and negative examples can be extracted from the batch, the Samples-Per-Class- (SPC-) heuristic (see e.g. Roth et al. (2020c)) is generally used, where commonly . Given a batch size , the SPC- technique randomly selects classes from which training samples are then drawn randomly to be included in each batch .
After feeding the batch through the network, tuples are mined from the batch to use in the loss function. We refer to mining in this paper either as batch mining or overload the both as batch sampling terminology. The naive solution to tuple mining is random mining, in which all possible tuples of the form are considered and are randomly chosen from the batch. However, this method lacks the capacity to utilize valuable information about the current embedding space, and is prone to significant redundancy in the training signal Schroff et al. (2015); Wu et al. (2018).
Hu et al. (2014) For each , we randomly draw a positive example from and a negative example from to form the triplet .
Intuitively, this could be mitigated by hard mining heuristics searching for negative samples that are closer to the anchor sample in the embedding space than positive samples, thereby always ensuring a significant training signal. Unfortunately, such approaches are prone to heavy overfitting, training instability and large gradient variance, thereby commonly resulting in less-than-optimal solutions (see e.g. Schroff et al. (2015); Harwood et al. (2017); Wu et al. (2018)). Recent approaches thus establish more lenient heuristics, such as through the introduction of slack parameters to the hard mining objective (e.g. semi-hard mining Schroff et al. (2015) or softhard mining Roth & Brattoli (2019)).
For each , we randomly draw a positive example from , and a negative example from the set
While other adaptive means (e.g. Harwood et al. (2017); Roth et al. (2020a)) have shown strong performance improvements, modern predefined heuristics such as distance-weighted tuple mining Wu et al. (2018) offer a better cost-to-performance tradeoff Roth et al. (2020a). Here, the heuristic leverages the fact that embeddings are commonly normalized to have unit norm for regularization purposes Wu et al. (2018). This ensures a distribution over a unit hypersphere, in which explicit pairwise distributions can be established Weisstein (2002); Wu et al. (2018). By inverting this distribution, distance-weighted mining can thus encourage a much more diverse coverage of tuple difficulties, improving generalization performance and reducing gradient variance Wu et al. (2018).
For embedding spaces normalized to the -dimensional hypersphere , we haveWeisstein (2002); Wu et al. (2018) the following pairwise sampling distribution :
for embedding pairs and Euclidean distance . For each , we randomly draw a positive example from , and sample a negative example based on an inverse distance distribution w.r.t. :
Examined Objectives
Hoffer & Ailon (2018) The triplet loss extends the contrastive objective with sample triplets and can be defined as:
Margin loss extends the triplet objective through the inclusion of a learnable boundary between positive and negative pairs Wu et al. (2018). In our experiments, we utilise . These criteria are widely used (see e.g. Roth et al. (2020c); Musgrave et al. (2020)) and require mining to make use of the batch information.
Wu et al. (2018) The margin objective integrates the learnable distance boundary between positive and negative pairs of samples for a relative ordering of pairs with respect to as
Going beyond pairs and triplets, one can also consider the case of more general n-tuples, which was investigated e.g. in the N-Pair objective Sohn (2016) and the Multisimilarity loss Wang et al. (2020a).
Sohn (2016) N-Pair loss is a simple augmentation of the triplet framework in which all negatives in the batch are incorporated in the objective function as:
where denotes an embedding regularization parameter due to slow convergence for normalized embeddings stated in Sohn (2016)
Wang et al. (2020a) Multisimilarity loss fits into the ranking loss category, but in addition to evaluation of cosine similarity between positive-anchor pairs and negative-anchor pairs, the objective evaluates positive-positive and negative-negative pairs with respect to the anchor:
where cosine similarity for two normalized vectors .
Notably, the Multisimilarity loss employs a masking process as a stand-in for the lack of batch-mining heuristic. While this proves to be similarly successfull in addressing the tuple sampling complexity issue, this can also be addressed through the usage of proxy-samples. These are dummy variables that represent various contextual properties (such as mean class representations) to serve as standing for actual samples, which is found e.g. in the ArcFace Deng et al. (2019) or ProxyNCA loss Movshovitz-Attias et al. (2017).
where the angular component is encoded in additive angular margin penalty , and is a scaling parameter, which denotes the radius of the effective utilized hypersphere .
Standard Performance Metrics
Performance metrics in deep metric learning aim to capture the quality of the similarity metric learned by the deep embedding model. Therefore, standard performance metrics in DML reflect the closeness between samples of the same class, the separability of samples of different classes, the clustering quality of embedding, and the uniformity over the hypersphere embedding space, which has been linked to zero-shot generalization capability Wang & Isola (2020), as discussed in Section 2. In our experiments, we utilize recall@1 Jegou et al. (2011), normalized mutual information Manning et al. (2010) between cluster labels assigned by the well-known K-Means Lloyd (1982) algorithm and ground-truth class labels, and to measure the closeness between samples of same class, cluster quality of the embedding (and hence, the separability of distinct classes) and uniformity, respectively. Here, we define these metrics formally, but we note that there exist multitudinous performance metrics for DML that we do not define here or use explicitly for our results, including f1 score, mean average precision (mAP), and recall@k for Jegou et al. (2011).
Jegou et al. (2011) Given , denote as defined in Definition 1. Then, Recall@k is measured as:
Manning et al. (2010) Let a clustering algorithm, such as -Means Lloyd (1982) with the number of clusters set to , such that indicates the cluster label for data point . The normalized mutual information score between the target labels and the cluster labels is measured as:
where for random variables , denotes the mutual information function:
and denotes the entropy function:
The performance metric , used to measure feature uniformity for our empirical evaluations, is defined in Section 3.1.
A.2 Classification Fairness Definitions
Fairness definitions and criteria in classification are briefly mentioned in Section 2 of the main paper. Here, we provide explicit formulas for the most common fairness definitions, including demographic parity, equalized odds, and equality of opportunity Hardt et al. (2016), and provide some additional context on fairness definition evolution.
The predictor satisfies demographic parity with respect to attribute and class if the predictor is independent of :
Specifically, demographic parity has largely been used over the years as a simple and intuitive definition of fairness, in which a classifier is said to satisfy demographic parity if the sensitive attribute is independent of the output of the classifier. While demographic parity provides a simple fairness definition, the measure cannot capture fairness in classification tasks where the ground-truth label is inherently related to a certain attribute value Li et al. (2017).
The predictor satisfies demographic parity with respect to attribute and class if the predictor is independent of conditional on :
The predictor satisfies demographic parity with respect to attribute and class if the predictor is independent of conditional on positively labelled :
This lead to the introduction of other fairness definitions that capture such nuances, the most well-known of which are probably equalized odds and equality of opportunity Hardt et al. (2016). However, fairness metrics overall have been criticized due to the choice of protected attribute over which to measure, and the inability of these metrics to capture bias with respect to certain attributes which are not known at test-time. We discuss this to a limited extent in Section 7.
Appendix B Dataset Summary Statistics
Appendix C Additional Results
Additional results for all loss and batch mining strategies for the manually class imbalanced experiments and balanced controls for CARS196 are located in Tables 6 and 7. K-Means was also tested as a downstream classifier but showed poor performance. The impact of varying imbalance in the manually class imbalanced CARS196 experiments with all tested downstream classifiers is displayed in Table 8. Additional results for benchmarking of further fairness improvement methods in downstream classification of "imbalanced" embeddings (aside from naive use of balanced datasets) are shown in Table 8(b).
C.2 CUB200
Additional results for all loss and batch mining strategies for the manually class imbalanced experiments and balanced controls for CUB200 are located in Tables 9 and 10. K-Means was also tested as a downstream classifier but showed poor performance. The impact of varying imbalance in the manually class imbalanced experiments in the upstream embedding, and all tested downstream classifiers is displayed in Table 9. Additional results for benchmarking of further fairness improvement methods in downstream classification of "imbalanced" embeddings (aside from naive use of balanced datasets) are shown in Table 11(b). Benchmarking of fairness improvement methods in downstream classification for bird color are shown in Table 12(b). Per-subgroup and overall results for CUB200 color experiments with standard margin-distance and PARADE are displayed in Table 13.
C.3 CelebA
Additional results for all loss and batch mining strategies for the CelebA dataset are located in Tables 14 and 15. Additional PARADE results for subgroup gaps excluding Fitzpatrick Skintone VI (as mentioned in Section 6.3) are located in Table 16.
C.4 LFW
Additional results for all loss and batch mining strategies for the LFW dataset are located in Tables 18 and 19. Per-subgroup results for LFW to demonstrate worst-group performance for the “White" subgroup (as mentioned in Section 6.3) are located in Table 20. Benchmarking of fairness improvement methods in downstream classification for bird color are shown in Table 21(b).
C.5 Exploration of Fairness - Utility Tradeoff and Varying Hyperparameters in PARADE
We vary and in the PARADE objective to explore the relationship between the overall performance, subgroup gap, and worst-group performance in PARADE. As stated in the main paper, we optimize and via worst-group performance. Results of this analysis are displayed in Figure 12. We use our exploration to expound on how to optimize for and . As seen in Figure 12, a clear trend that inversely relates overall performance, and fairness as measured by subgroup gap and worst-group performance is seen for the uniformity metric, over the grid of and values (Note that higher values of correspond to worse performance). Recall@1 and NMI demonstrate noisier relationships between overall performance and fairness; and several , choices appear to select an optimal tradeoff. In Figure 12, for Recall@1, we observe that at the location , in the optimization grid, PARADE reaches peak overall performance and fairness (measured by low subgroup gap and high performance for the worst-performing subgroup) simultaneously. Thus, we could conclude that this choice of and represents an optimal tradeoff for utility and fairness in PARADE as measured by Recall@1. By the other displayed metrics, we see that , demonstrates a reasonable utility-fairness tradeoff. Therefore, the choice of , would be optimal for PARADE in CUB200 bird color setting. Note that the choice of where to operate within this trade-off should depend on the application that is being targeted. For example, here we use Recall@1 to determine the optimal choice of hyperparameters and validate with the other two considered metrics. However, for LFW, which has a high population of singleton classes (see Figure 7), NMI would be a better metric to use for selecting optimal point.
Appendix D Implementation Details
Dataset manipulation for the CARS196 and CUB200 manually class imbalanced experimentsis explained in Section 3.3.
For CUB200 bird color experiments, we utilized the labeled bird color attributes from Wah et al. (2011). Each image can have multiple “primary color" labels. Therefore, we take the mode over all bird colors labeled for each image in order to determine a single bird color associated with the image. For CelebA, we calculate the Fitzpatrick skintone based on the image pixel information for each image. The calculation is described in Section D.2. For LFW, we construct the “Race" attribute from labels of “White", “Black", “Asian," and “Indian" as labelled by Kumar et al. (2009). For each of these attributes, the labelling provided by Kumar et al. (2009) has a float value, which we map to binary values: the image is considered to have the attribute if the value is greater than , and the image is considered to not have the attribute if the value is less than . Naturally, the labelling is not necessarily correct for each image, as the confidence about the “Race" labelling can be quite low for some images. We remove all images without at least one of these attributes, though we note that these attributes do not encompass all races. Therefore, our analysis may not be relevant for other races not labelled by Kumar et al. (2009).
D.2 Fitzpatrick Skintone Calculation
We follow the methods from Cheng et al. (2021) for calculation of Fitzpatrick Skintone based on image pixel information. However, we calculate these values for CelebA, as opposed to CelebA-HQ. As CelebA-HQ incorporates higher resolution images, but has fewer images, our process of Fitzpatrick Skintone calculation on CelebA is slightly modified to account for lower resolution, and differing image size.
In Cheng et al. (2021), two sample skin patches are selected from each image of CelebA-HQ to determine the skintone. We select three sample skin patches, as we are forced to reduce the dimensions of the patches to account for the smaller image size of CelebA. Additionally, we leverage facial landmark attributes provided by CelebA Liu et al. (2015) in order to choose our sample patches. Specifically, given the landmarks for the left eye, right eye, and nose for each image, we choose to sample square patches of size (all color channels are selected) with the following center points:
The first two center points are intended to capture the likely location of the left and right cheeks, respectively, as these are likely located below each eye and adjacent to the nose. The last center point is the nose. We note that this protected attribute generation is not perfect. In some cases, such label generation can accidentally use aspects of the background, if, the individual’s face position in the image is not facing forward. Also, extreme lighting can lead to misclassification of skintone. Nonetheless, we believe the procedure provides a good approximation of Fitzpatrick skintone category, but do not recommend these attribute labels for use outside of fairness analysis.
The selected sample patches are converted to CIELab-space to retrieve the (luminance) and (yellow) values. We then calculate the Mean Individual Typology Angle (ITA) value:
Based on the Mean ITA calculation, we classify each image into one of the Fitzpatrick skintone categories, as listed in Table 23. To calculate subgroup gaps, we calculate gaps between the mean value over the lightest Fitzpatrick skintones and the mean value over the darkest Fitzpatrick skintones.
D.3 Training Parameters
For CUB200 and CARS196, we did not perform hyperparameter search but followed reported hyperparameters from Roth et al. (2020c) for best performance with an ImageNet Deng et al. (2009) pretrained ResNet50 He et al. (2016) and frozen batch normalization layers. As detailed in Roth et al. (2020c), we train for epochs with embedding dimension , learning rate with no scheduler, and weight decay . We train with a batch size of , with the Adam optimizer Kingma & Ba (2015) over five seeds inclusive for the balance control datasets, and for CUB200 color experiments; and seeds for the manually class imbalanced experiments. For training transforms, we normalize each image using color channel means and standard deviations , randomly crop the image and re-size to and horizontally flip with probability . For testing transforms, we normalize each image with the aforementioned color channel means and standard deviations, resize to , and center crop to .
For CelebA and LFW, we performed hyperparameter search over the following hyperparameters: architectures: ResNet50 He et al. (2016), and SE-Net50 (both with and without frozen batch normalization layers); number of training epochs; learning rates; last linear layer learning rate (differ from other layer learning rates); learning rate schedulers; embedding dimensions: , , ; pre-training; image augmentations. We evaluated hyperparameter sets on a validation set we cut from the typical training set (20% of training set), and chose the set of hyperparameters with best recall@k score for CelebA and best NMI score for LFW. NMI is used for LFW due to the high number of singleton classes present in the dataset (recall@1 is meaningless for singleton classes).
For CelebA, we train on the ResNet50 He et al. (2016) architecture with frozen batch normalization layers, for epochs with learning rate , and no scheduler, weight decay , Adam Kingma & Ba (2015) optimizer, and batch size of . For training transforms, we normalize each image using color channel means and standard deviations , resize to , center crop to and horizontally flip with probability . For testing transforms, we normalize each image with the aforementioned color channel means and standard deviations, resize to , and center crop to . We average over runs with seeds , inclusive.
For LFW, we train on the ResNet50 He et al. (2016) architecture with frozen batch normalization layers, for epochs with initial learning rate for all model parameters except the last linear layer, which has initial learning rate , and a multi-step learning rate scheduler which reduces the learning rate by a factor of at epochs and , weight decay , Adam Kingma & Ba (2015) optimizer, and batch size of . For training transforms, we normalize each image using color channel means and standard deviations , resize to , center crop to and horizontally flip with probability . For testing transforms, we normalize each image with the aforementioned color channel means and standard deviations, resize to , and center crop to . We average over runs with seeds , inclusive.
For each dataset we chose a set of loss and batch mining strategies that have historically been used for the relevant task, encompassing a broad range of methods, and / or achieved good performance. However, for n-pair loss and sampling, good performance was not achieved for the facial datasets despite use in the past for facial recognition Sohn (2016). For manually class imbalanced experiments with CARS196 and CUB200 and the associated balanced controls, we used: margin loss / distance-weighted sampling, margin loss / semi-hard sampling, triplet loss / distance-weighted sampling, triplet loss / semi-hard sampling, contrastive loss / distance-weighted sampling, multisimilarity loss, and proxy-NCA loss. For the color experiments with CUB200, we used: margin loss / distance-weighted sampling. For CelebA and LFW, we used: margin loss / distance-weighted sampling, arcface loss, and n-pair loss and sampling. For all testing and evaluation experiments with PARADE, we used margin loss and distance-weighted sampling, but PARADE can be used with any loss and mining strategy.
Here we list the hyperparameters that we use for each evaluated loss function and batch mining strategy, if applicable. Refer to A for explicit formulas associated with the parameters here. We set in semi-hard mining. For distance-weighted mining, we set and clip the maximum distance to . In the triplet objective, we use for triplet loss. For margin loss, the learning rate of the boundary is set to , with initial value and triplet margin . For N-Pair uses embedding regularization parameter . In Multisimilarity loss, we use , , and . Finally, for ArcFace, additive angular margin penalty is set to , while scaling parameter and class centers are optimized with learning rate .
The two PARADE parameters, and , as described in Section 4, were optimized via worst-group performance over a grid search. For CUB200, we set , . For CelebA, we set , . For LFW, we set , .
D.4 Fairness Evaluation
For each dataset, we calculate subgroup gaps between the majoritized and minoritized subgroup (CARS196, CUB200 class, CelebA) or between the worst-performing subgroup and other subgroups (LFW). In CUB200 color experiments, due to the large number of subgroups, we calculate the gap between the top performing subgroups and the bottom performing subgroups (there are total subgroups).
In the upstream embedding tasks, in which we denote as the embedding function for the learned model, and use to denote the value of the attribute for data point , we calculate recall@1 for data samples in with associated class label and attribute as:
Note here that the nearest neighbors function is computed with respect to all , not exclusively with attribute , but the input to the nearest neighbors function is exclusively . To calculate NMI, let be the output of a clustering algorithm on the entire dataset , i.e. and let denote the output of clustering algorithm restricted to some subset . The important note here is that the clustering algorithm is run over the entire dataset, but expresses the cluster labels only for . Then, we measure NMI for data samples in with associated class label and attribute as:
We calculate for data samples in with attribute as:
where denotes the singular values over .
Downstream
In the downstream tasks, for data samples in with class label , and predictor , let express the value of the ground-truth label for data sample and let express the value of the predicted label. Then, we denote the number of true positives with attribute :
the number of false positives with attribute
and the number of false negatives with attribute :
We calculate macro-averaged recall for data samples in with associated class label and attribute as:
where is the number of possible class labels, i.e. the size of the set of all possible values of . We calculate macro-averaged precision for data samples in with associated class label and attribute as:
We calculate accuracy for data samples in with associated class label and attribute as:
The subgroup gaps are then considered to be the difference between the metric value for the majoritized subgroup and the metric value for the minoritized subgroup (CARS196, CUB200, CelebA); or between the metric value for each subgroup with better performance than the worst-performing subgroup and the metric value for the worst-performing subgroup (LFW). As stated in Section D.4, for CUB200 bird color experiments, the subgroup gaps were calculated between the top performing % of subgroups and bottom performing of subgroups.