Examining CNN Representations with respect to Dataset Bias

Quanshi Zhang, Wenguan Wang, Song-Chun Zhu

Introduction

Given a convolutional neural network (CNN) that is pre-trained to estimate image attributes (or labels), how to diagnose black-box knowledge representations inside the CNN and discover potential representation flaws is a crucial issue for deep learning. In fact, there is no theoretical solution to identifying good and problematic representations in the CNN. Instead, people usually just evaluate a CNN based on the accuracy obtained using testing samples.

In this study, we focus on representation flaws caused by potential bias in the collection of training samples (?). As shown in Fig. 1, if an attribute usually co-appears with certain visual features in training samples, then the CNN may be learned to use the co-appearing features to represent this attribute. When the used co-appearing features are not semantically related to the target attribute, we consider these features as biased representations. This idea is related to the disentanglement of the local, bottom-up, and top-down information components for prediction (?; ?; ?). We need to clarify correct and problematic contexts for prediction. CNN representations may be biased even when the CNN achieves a high accuracy on testing samples, because testing samples may have a similar bias.

In this paper, we propose a simple yet effective method that automatically diagnoses representations of a pre-trained CNN without given any testing samples. I.e., we only use training samples to determine the attributes whose representations are not well learned. We discover blind spots and failure modes of the representations, which can guide the collection of new training samples.

Intuition, self-compatibility of network representations: Given a pre-trained CNN and an image II, we use the CNN to estimate attribute AA for II. We also mine inference patternsWe regard a neural pattern as a group of units in a channel of a conv-layer’s feature map, which are activated and play a crucial role in the estimation of the attribute AA. of the estimation result, which are hidden in conv-layers of the CNN. We can regard the mined inference patterns as exact representations of the attribute AA in the CNN. Then, based on inference patterns, we compute the relationship between each pair of attributes (Ai,Aj)(A_{i},A_{j}), i.e. identifying whether AiA_{i} is positively/negatively/not related to AjA_{j}.

The intuition is simple, i.e. according to human’s common sense, we set up several ground-truth relationships between some pairs of attributes as rules to diagnose CNN representations. The mined attribute relationships should well fit the ground truth; otherwise, the representation is probably not well learned. Let us take a CNN that is learned to estimate face attributes for example. As shown in Fig. 2, the smiling attribute is supposed to be represented by features (patterns), which appear on the mouth region in conv-layers. Whereas, the black hair attribute should be inferred by features extracted from hairs. Therefore, the attribute relationship “smiling is not related to the black hair” is trustworthy enough to become a ground truth. However, the CNN may use eye/nose features to represent the two attribute, because these attributes always co-appear with specific eye/nose appearances in a biased dataset. Thus, we will mine a specific relationship between the two attributes, which conflicts with the ground truth.

Our method: Given a pre-trained CNN, we mine relationships between each pair of attributes according to their inference patterns. Then, we annotate some ground-truth attribute relationships. For example, the heavy makeup attribute is positively related to the attractive attribute; black hair and smiling are not related to each other. We compute the Kullback-Leibler (KL) divergence between the mined relationships and ground-truth relationships to discover attributes that are not well learned, including both blind spots and failure modes of attribute representations.

In fact, how to define ground-truth relationships is still an open problem. We can ask different people to label attribute relationships in their personal opinions to approach the ground truth. More importantly, our method is compatible with various types of ground-truth distributions. People can define their ground truth w.r.t. their tasks as constraints to examine the network. Thus, our method is a flexible and convincing way to discover representation bias at the level of human cognition.

The annotation cost of our method is usually much lower than end-to-end learning of CNNs. Our annotation cost is O(n2)O(n^{2}), where nn denotes the number of attribute outputs. In contrast, it usually requires thousands or millions of samples to learn a new CNN in real applications.

Why is the proposed method important? As a complement to using testing samples for evaluation, our zero-shot diagnosis of a CNN is of significant values in applications:

A high accuracy on potentially biased testing samples cannot prove correct representations of a CNN.

Potential bias cannot be fully avoided in most datasets. Especially, some attributes (e.g. smiling) mainly describe specific parts of images, but the dataset (?; ?) only provides image-level annotations of attributes for supervision without specifying regions of interests, which makes the CNN more sensitive to dataset bias.

More crucially, the level of representation bias is not necessary to be proportional to the dataset bias level. We need to diagnose the actual CNN representations.

In conventional studies, correcting representation flaws caused by either dataset bias or the over-fitting problem is a typical long-tail problem. If we blindly collect new training samples without being aware of failure modes of the representation, it would require massive new samples to overcome the bias problem. Our method provides a new perspective to solve the long-tail problem.

Unlike methods of CNN visualization/analysis (?; ?; ?; ?) that require people to one-by-one check the representation of each image, our method discovers all biased representations in a batch.

Contribution: In this study, to the best of our knowledge, we, for the first time, propose a method to discover potentially biased representations hidden in a pre-trained CNN without testing samples. Our method mines blind spots and failure modes of a CNN in a batch manner, which can guide the collection of new samples. Experiments have proved the effectiveness of the proposed method.

Related work

Visualization of CNNs: In order to open the black box of a CNN, many methods (?; ?; ?; ?; ?; ?) have been developed to visualize and analyze patterns of response units in a CNN. Some methods (?; ?; ?) back-propagate gradients w.r.t. a given unit to pixel values of an image, in order to obtain an image that maximizes the score of the unit. These techniques mainly visualize simple patterns. As mentioned in (?), attributes are an important perspective to model images, but it is difficult to visualize a complex attribute (e.g. the attractive attribute).

Given a feature map produced by a CNN, Dosovitskiy et al. (?) trained a new up-convolutional network to invert the feature map to the original image. Similarly, this approach was not designed for the visualization of a single attribute output.

Interpreting semantic meanings of CNNs: Going beyond the “passive” visualization of neural patterns, some studies “actively” retrieve mid-level patterns from conv-layers, which potentially corresponds to a certain object/image part. Zhou et al. (?; ?) mined patterns for “scene” semantics from feature maps of a CNN. Simon et al. discovered objects (?) from CNN feature maps in an unsupervised manner, and retrieved patterns for object parts in a supervised fashion (?). Zhang et al. (?) used a graphical model to organize implicit mid-level patterns mined from a CNN, in order to explain the pattern hierarchy inside conv-layers in a weakly-supervised manner. (?) used a gradient-based method to interpret visual question-answering models. Zhang et al. (?) transformed CNN representations to an explanatory graph, which represents the semantic hierarchy hidden inside a pre-trained CNN.

Model diagnosis: Many methods have been developed to diagnose representations of a black-box model. (?) extracted key features for model outputs. The LIME method proposed by Ribeiro et al. (?) and gradient-based visualization methods (?; ?) extracted image regions that were responsible for each network output, in order to interpret the network representation.

Unlike above studies diagnosing representations for each image one by one, many approaches aim to evaluate all potential attribute/label representations for all images in a batch. Lakkaraju et al. (?) and Zhang et al. (?; ?) explored unknown knowledge hidden in CNNs via active annotations and active question-answering. Methods of (?; ?) computed the distributions of a CNN’s prediction errors among testing samples, in order to summarize failure modes of the CNN. However, we believe that compared to (?; ?), it is of larger value to explore evidence of failure cases from mid-layer representations of a CNN. (?) required people to label dimensions of input features that were related to each output according to common sense, in order to learn a better model. Hu et al. (?) designed some logic rules for network outputs, and used these rules to regularize the learning of neural networks. In our research, we are inspired by Deng et al. (?), which used label graph for object classification. We use ground-truth attribute relationships as logic rules to harness mid-layer representations of attributes. (?; ?; ?) tried to isolate and diagnose information from local, bottom-up, or top-down inference processes. More specially, (?) proposed to separate implicit local representations and explicit contextual information used for prediction. Following this direction, this is the first study to diagnose unreliable contextual information from CNN representations w.r.t. dataset bias.

Active learning: Active learning is a well-known strategy for detecting “unknown unknowns” of a pre-trained model. Given a large number of unlabeled samples, existing methods mainly select samples on the decision boundary (?) or samples that cannot well fit the model (?; ?), and require human users to label these samples.

Compared to active-learning approaches, our method does not require any additional unlabeled samples to test the model. More crucially, our method looks deep inside the representation of each attribute to mine attribute relationships; whereas active learning is closer to black-box testing of model performance. As discussed in (?), unless the initial training set contains at least one sample in each possible mode of sample features, active learning may not exhibit high efficiency in model refinement.

Algorithm

We are given a CNN that is trained using a set of images I{\bf I} with attribute annotations. The CNN is designed to estimate nn attributes of an image, denoted by A1,A2,…,AnA_{1},A_{2},\ldots,A_{n}. Meanwhile, we also have a certain number of ground-truth relationships between different attributes, denoted by a relationship graph G∗=({Ai},E∗)G^{*}=(\{A_{i}\},{\bf E}^{*}). Each edge (Ai,Aj)∈E∗(A_{i},A_{j})\in{\bf E}^{*} represents the relationship between AiA_{i} and AjA_{j}. Note that it is not necessary for G∗G^{*} to be a complete graph. We only select trustworthy relationships as ground truth. The goal is to identify attributes that are not well learned and to discover blind spots and failure modes in attribute representation.

In different applications, people use multiple ways to define attributes (or labels), including binary attributes (Yi∈{−1,+1}Y_{i}\in\{-1,+1\}) and continuous attributes (e.g. Yi∈[−1,+1]Y_{i}\in[-1,+1] and Yi∈(−∞,+∞)Y_{i}\in(-\infty,+\infty)). We can normalize all these attributes to the range of Yi∈(−∞,+∞)Y_{i}\in(-\infty,+\infty) for simplificationGiven annotations of continuous attributes Yi∗∈(−∞,+∞)Y^{*}_{i}\in(-\infty,+\infty), we can define L-2 norm loss L(Yi,Yi∗)=(Yi∗−Yi)2L(Y_{i},Y^{*}_{i})=(Y^{*}_{i}-Y_{i})^{2} to train the CNN. Given annotations of binary attributes for training Yi∗∈{−1,+1}Y^{*}_{i}\in\{-1,+1\}, we can use the logistic log loss L(Yi,Yi∗)=log⁡(1+exp⁡(−Yi⋅Yi∗))L(Y_{i},Y^{*}_{i})=\log(1+\exp(-Y_{i}\cdot Y^{*}_{i})) to train the CNN. In this way, YiY_{i} can be considered as an attribute estimation whose range is (−∞,+∞)(-\infty,+\infty).. To simplify the introduction, without loss of generality, we consider Yi∗>0Y^{*}_{i}>0 as the existence of a certain attribute AiA_{i}; otherwise not. Consequently, we flip the signs of some ground-truth annotations to ensure that we use positive values, rather than negative values, to represent the activation of AiA_{i}.

Mining attribute relationships

Attribute representation: Given an image I∈II\in{\bf I} and a target attribute AiA_{i}, we select the feature map xI{\bf x}^{I} of a certain conv-layer of the CNN to represent AiA_{i} and compute YiIY_{i}^{I}. Since the CNN conducts a series of convolution and ReLu operations on xI{\bf x}^{I} to compute YiIY_{i}^{I}, we can approximate YiIY_{i}^{I} as a linear combination of neural activations in xI{\bf x}^{I}.

where viI{\bf v}_{i}^{I} denotes a weight vector, and βiI\beta_{i}^{I} is a scalar for bias.

In the above equation, parameters νiI{\boldsymbol{\nu}}_{i}^{I} and βiI\beta_{i}^{I} reflect inherent piecewise linear representations of YiIY_{i}^{I} inside the CNN, whose values have been fixed when the CNN and the image are given. We will introduce the estimation of νiI{\boldsymbol{\nu}}_{i}^{I} and βiI\beta_{i}^{I} later. The target parameter here is ρi∈{0,1}N{\boldsymbol{\rho}}_{i}\in\{0,1\}^{N}, which is a sparse mask vector. It means that we select a relatively small number of reliable neural activations from xI{\bf x}^{I} as inference patterns of AiA_{i} and filters out noises. ∘\circ denotes element-wise multiplication between vectors. We can regard ρi{\boldsymbol{\rho}}_{i} as a prior spatial distribution of neural activations that are related to attribute AiA_{i}. For example, if AiA_{i} represents an attribute for noses, then we expect ρi{\boldsymbol{\rho}}_{i} to mainly represent nose regions. Note that except ρi{\boldsymbol{\rho}}_{i}, parameters νiI{\boldsymbol{\nu}}_{i}^{I} and βiI\beta_{i}^{I} are only oriented to image II due to ReLu operations in the CNN.

We can compute the inherent piecewise linear gradient w.r.t. xI{\bf x}^{I}, i.e. νiI{\boldsymbol{\nu}}_{i}^{I} via gradient back propagation.

where the CNN contains MM conv-layres (including fully-connected layers), and xk{\bf x}_{k} denotes the output of the kk-th conv-layer (x=defxm{\bf x}\overset{\text{def}}{=}{\bf x}_{m} corresponds to the mm-th conv-layer). We can further compute the value of βiI\beta_{i}^{I} based on the full representation without pattern selection YiI=(νiI)TxI+βiIY_{i}^{I}=({\boldsymbol{\nu}}_{i}^{I})^{T}{\bf x}^{I}+\beta_{i}^{I}.

Inspired by the LIME method (?), the loss of mining inference patterns is similar to a Lasso selection:

where L(YiI,ρi){\mathcal{L}}(Y_{i}^{I},{\boldsymbol{\rho}}_{i}) measures the fidelity of the representation on image II, and L(ρi){\bf L}({\boldsymbol{\rho}}_{i}) denotes the representation complexity. We can simply formulate L(YiI,ρi)=[(viI)TxI+βiI−YiI]2{\mathcal{L}}(Y_{i}^{I},{\boldsymbol{\rho}}_{i})=[({\bf v}_{i}^{I})^{T}{\bf x}^{I}+\beta_{i}^{I}-Y_{i}^{I}]^{2}, and L(ρi)=λ∥ρi∥1{\bf L}({\boldsymbol{\rho}}_{i})=\lambda\|{\boldsymbol{\rho}}_{i}\|_{1}, where ∥⋅∥1\|\cdot\|_{1} denotes L-1 norm, and λ\lambda is a constant. Based on the above equation, ρ^i\hat{\boldsymbol{\rho}}_{i} can be directly estimated using a greedy strategy.

Attribute relationships: For each pair of attributes AiA_{i} and AjA_{j}, we define a cosine distance ϖijI=def(vjI)TviI∥vjI∥∥viI∥\varpi_{ij}^{I}\overset{\text{def}}{=}\frac{({\bf v}^{I}_{j})^{T}{\bf v}^{I}_{i}}{\|{\bf v}^{I}_{j}\|\|{\bf v}^{I}_{i}\|} to represent their attribute relationship. If AiA_{i} and AjA_{j} are positively related, vi{\bf v}_{i} will approximate to vj{\bf v}_{j}, i.e. ϖijI\varpi_{ij}^{I} will be close to 1. Similarly, if AiA_{i} and AjA_{j} are negatively related, then ϖijI\varpi_{ij}^{I} will be close to -1. If AiA_{i} and AjA_{j} are not closely related, then vj{\bf v}_{j} and vi{\bf v}_{i} will be almost orthogonal, thus ϖjiI≈0\varpi_{ji}^{I}\approx 0.

The actual representation of an attribute in a CNN is highly non-linear, and the linear representation in Eq. (1) is just a local mode oriented to a specific image II. When we compute the gradient νiI=∂Yi∂xI{\boldsymbol{\nu}}_{i}^{I}=\frac{\partial Y_{i}}{\partial{\bf x}^{I}}, the ReLu operation blocks irrelevant information in gradient back-propagation, thereby obtaining a local linear representation. It is possible to cluster νiI{\boldsymbol{\nu}}_{i}^{I} of different images into several local modes of the representation. Expect extreme cases mentioned in (?), these local modes are robust to most small perturbations in the image II.

Diagnosis of CNN representations

Given each image I∈II\in{\bf I}, we compute ϖijI\varpi_{ij}^{I} to represent the relationship between AiA_{i} and AjA_{j} w.r.t. the image II. In this way, we use the distribution of ϖijI\varpi_{ij}^{I} among all training images in I{\bf I}, denoted by Q(ϖij∣Ai,Aj){\bf Q}(\varpi_{ij}|A_{i},A_{j}), to represent the overall attribute relationshipWithout loss of generality, we modify attribute annotations to ensure Yi∗=+1Y_{i}^{*}=+1 rather than Yi∗=−1Y_{i}^{*}=-1 to indicate the existence of a certain attribute. We find that the CNN mainly extracts common patterns from positive samples as inference patterns to represent each attribute. Thus, we compute distributions P{\bf P} and Q{\bf Q} for (Ai,Aj)(A_{i},A_{j}) among the samples in which either Yi∗=+1Y_{i}^{*}=+1 or Yj∗=+1Y_{j}^{*}=+1. Similarly, in Experiment 3, we also ignored samples with Yi∗=Yj∗=−1Y_{i}^{*}=Y_{j}^{*}=-1 to compute the entropy for the competing method.. Fig. 3 shows the mined distributions of Q(ϖij∣Ai,Aj){\bf Q}(\varpi_{ij}|A_{i},A_{j}) for different pairs of attributes.

Besides the observation distribution Q{\bf Q}, we also manually annotate a number of ground-truth attribute relationships G∗G^{*}, and define a distribution for each ground-truth attribute relationship P(ϖij∣Ai,Aj){\bf P}(\varpi_{ij}|A_{i},A_{j}). People can label several types of ground-truth relationships for (Ai,Aj)∈E∗(A_{i},A_{j})\in{\bf E}^{*}, lij∈L={L1,L2,…}l_{ij}\in{\bf L}=\{L_{1},L_{2},\ldots\}, to supervise the diagnosis of CNN representations. Let (Ai,Aj)∈E∗(A_{i},A_{j})\in{\bf E}^{*} be labeled with lij=L∗l_{ij}=L^{*}. We assume the ground-truth distribution P(ϖij∣Ai,Aj)∼N(μL∗,σL∗2){\bf P}(\varpi_{ij}|A_{i},A_{j})\sim{\mathcal{N}}(\mu_{L^{*}},\sigma_{L^{*}}^{2}) follows a Gaussian distribution. We assume most pairs of attributes are well learned, so we can compute μL∗\mu_{L^{*}} and σL∗2\sigma_{L^{*}}^{2} as the mean and the variation of ϖij\varpi_{ij}, respectively, among all pairs of attributes that are labeled with L∗L^{*}. In this way, biased representations correspond to outliers of ϖij\varpi_{ij} w.r.t the ground-truth distribution.

We can compute the KL-divergence between P{\bf P} and Q{\bf Q}, KL(P∥Q){\bf KL}({\bf P}\|{\bf Q}), to discover biased representations.

where P(Aj∣Ai)=1/deg(Ai)P(A_{j}|A_{i})=1/\textrm{deg}(A_{i}) is a constant given the degree of AiA_{i}. We approximately set Ω ⁣= ⁣\Omega\!=\!, because P(ϖij∣Aj,Ai)≈0{\bf P}(\varpi_{ij}|A_{j},A_{i})\approx 0 when ∣ϖij∣>1|\varpi_{ij}|>1 in real applications. We believe that if KLAi{\bf KL}_{A_{i}} is high, AiA_{i} is probably not well learned.

Blind spots & failure modes: Each pair of attributes (Ai,Aj)∈E∗(A_{i},A_{j})\in{\bf E}^{*} with a high KLAiAj{\bf KL}_{A_{i}A_{j}} may have two alternative explanations. The first explanation is that (Ai,Aj)(A_{i},A_{j}) represents a blind spot of the CNN. I.e. (Ai,Aj)(A_{i},A_{j}) should be positively/negatively related to each other according to the ground-truth, but the CNN has not learned many inference patterns that are shared by both AiA_{i} and AjA_{j}. In this case, the CNN does not encode the inference relationship between AiA_{i} and AjA_{j}.

The alternative explanation is that (Ai,Aj)(A_{i},A_{j}) represents a failure mode. If the mined relationship is that AiA_{i} is strongly positively related to AjA_{j}, which conflicts with the ground-truth relationship. Then, samples with opposite ground-truth annotations for AiA_{i} and AjA_{j}, Yi∗⋅Yj∗ ⁣< ⁣0Y_{i}^{*}\cdot Y_{j}^{*}\!<\!0 may correspond to a failure mode in attribute estimation. Note that these samples belong to two modes, i.e. the modes of Yi∗>0,Yj∗ ⁣< ⁣0Y_{i}^{*}>0,Y_{j}^{*}\!<\!0 and Yi∗<0,Yj∗ ⁣> ⁣0Y_{i}^{*}<0,Y_{j}^{*}\!>\!0. We simply select the mode with fewer samples as a failure mode. Similarly, if the CNN incorrectly encodes a negative relationship between (Ai,Aj)(A_{i},A_{j}), then we select a failure mode from candidates of (Yi∗ ⁣> ⁣0,Yj∗ ⁣> ⁣0)(Y_{i}^{*}\!>\!0,Y_{j}^{*}\!>\!0) and (Yi∗ ⁣< ⁣0,Yj∗ ⁣< ⁣0)(Y_{i}^{*}\!<\!0,Y_{j}^{*}\!<\!0).

In practise, we determine blind spots and failure modes as follows. Given a pair of attributes (Ai,Aj)(A_{i},A_{j}) with a high KLAiAj{\bf KL}_{A_{i}A_{j}}, if ∣EI[ϖijI]∣ ⁣< ⁣0.2|{\bf E}_{I}[\varpi_{ij}^{I}]|\!<\!0.2 and ∣EI[ϖijI]−μlij∣>0.2|{\bf E}_{I}[\varpi_{ij}^{I}]-\mu_{l_{ij}}|>0.2, then (Ai,Aj)(A_{i},A_{j}) correspond to a blind spot. If ∣EI[ϖijI]∣ ⁣> ⁣0.2|{\bf E}_{I}[\varpi_{ij}^{I}]|\!>\!0.2 and ∣EI[ϖijI]−μlij∣ ⁣> ⁣0.2|{\bf E}_{I}[\varpi_{ij}^{I}]-\mu_{l_{ij}}|\!>\!0.2, we extract a failure mode from (Ai,Aj)(A_{i},A_{j}).

Experiments

Dataset: We tested the proposed method on the Large-scale CelebFaces Attributes (CelebA) dataset (?) and the SUN Attribute database (?). The CelebA dataset contains more than 200K celebrity images, each with 40 attribute annotations. In order to simplify the story, we first used annotations of face bounding boxes provided in the dataset to crop face regions from original images, and then used the cropped faces as input to learn a CNN. The SUN database contains 14K scene images with 102 attributes, but most attributes only appear in very few images. Thus, we selected 24 attributes with minimum scores of max⁡(#(Yi∗>0),#(Yi∗<0))\max(\#(Y_{i}^{*}>0),\#(Y_{i}^{*}<0)) as target attributes for experiments, where #(Yi∗>0)\#(Y_{i}^{*}>0) denotes the number of positive annotations of AiA_{i} among all images. Furthermore, in the SUN dataset, value ranges for ground-truth attribute annotations are Yi∗∈Y_{i}^{*}\in. We modified the ground-truth to binary annotations Yi∗,new=sign(Yi∗−0.5)Y_{i}^{*,new}=sign(Y_{i}^{*}-0.5) for simplicity.

Implementation details: In this study, we used the AlexNet (?) as the target CNN, which contains five conv-layers and three fully-connected layers. We tracked inference patterns of an attribute through different conv-layers, and we used inference patterns in the first conv-layer for CNN diagnosis. It is because that feature maps in lower conv-layers have higher resolutions and that inference patterns in lower conv-layers are better localized than higher conv-layers. Although low-layer patterns mainly represent simple shapes (e.g. edges), edges on black hairs and edges describing smiling should be localized at different positions.

For the CelebA dataset, we defined five types of attribute relationships“Definitely negative” is referred to as exclusive attributes, e.g. black hair and blond hair. Whereas, “probably” means a high probability. For example, a heavy makeup person is probably attractive., i.e. lij∈{definitely negative,l_{ij}\in\{\textrm{definitely negative}, probably negative,\textrm{probably negative}, not related,\textrm{not related}, probably positive,\textrm{probably positive}, definitely positive}\textrm{definitely positive}\}4. We obtained μdefinitely positive>μprobably positive>…>μdefinitely negative\mu_{\textrm{definitely positive}}>\mu_{\textrm{probably positive}}>\ldots>\mu_{\textrm{definitely negative}}. For the SUN dataset, we defined two types of attribute relationships, i.e. lij∈{negative,positive}l_{ij}\in\{\textrm{negative},\textrm{positive}\}. We manually annotated 18 probably positive relationships, 549 not-related relationships, 21 probably negative relationships, and 9 definitely negatively relationships in the CelebA dataset. In the SUN dataset, we labeled a total of 83 positive relationships and 63 negative relationships.

Experiment 1, mining potentially biased attribute representations: We trained two CNNs using images from the CelebA dataset and those from the SUN dataset, respectively. Then, we diagnosed attribute representations of the CNNs. Fig. 4 compares the mined and the ground-truth attribute relationships.

Our method is not sensitive to a small number of errors in ground-truth relationship annotations. It is because as shown in Fig. 4, for each attribute, we compute multiple pairwise relationships between this attribute and other attributes. A single inaccurate relationship will not significantly affect the result. Similarly, because we calculate the KL divergence among all training images, KL divergence results in Fig. 5 are robust to noise and appearance variations in specific images. The low error rate of attribute estimation is not necessarily equivalent to good representations. Error rates for the “wearing lipstick” and “double chin” attributes are only 8.1% and 4.5%, which are lower than the average error 11.1%. However, these two attributes have the top-4 representation biases.

As shown in Fig. 7, when the CNN uses patterns in incorrect positions to represent the attribute, we will probably obtain a significant KL divergence.

Experiment 2, testing the proposed method on manually biased datasets: In this experiment, we manually biased training sets to learn CNNs. We used our method to explore the relationship between the dataset bias and the representation bias in the CNNs.

From each of the CelebA and the SUN datasets, we randomly selected 10 pairs of attributes. Then, for each pair of attributes, (Ai,Aj)(A_{i},A_{j}), we biased the distribution of AiA_{i} and AjA_{j}’s ground-truth annotations to produce a new training set, as follows. Given a parameter τ\tau (0≤τ≤10\leq\tau\leq 1) that denotes the bias level, we randomly removed τ⋅NYi∗Yj∗ ⁣< ⁣0\tau\cdot N_{Y_{i}^{*}Y_{j}^{*}\!<\!0} samples from all samples whose ground-truth annotations Yi∗Y_{i}^{*} and Yj∗Y_{j}^{*} were opposite, where NYi∗Yj∗ ⁣< ⁣0N_{Y_{i}^{*}Y_{j}^{*}\!<\!0} denotes the number of samples that satisfied Yi∗Yj∗ ⁣< ⁣0Y_{i}^{*}Y_{j}^{*}\!<\!0.

Initially, for each pair of attributes (Ai,Aj)(A_{i},A_{j}), we generated a fully biased training set with τ=1\tau=1, and our method mined a significant KL divergence of KLAiAj{\bf KL}_{A_{i}A_{j}}. We then gradually added samples with Yi∗Yj∗ ⁣< ⁣0Y_{i}^{*}Y_{j}^{*}\!<\!0 to reduce the dataset bias τ\tau, and learned new CNNs based on the new training sets. Fig. 6 shows the decrease of the KL divergence when the dataset bias τ\tau was reduced. Given each of the 10 pairs of attributes, we generated four biased datasets by applying four values of τ∈{0.25,0.5,0.75,1.0}\tau\in\{0.25,0.5,0.75,1.0\}. In this way, we obtained 40 biased CelebA datasets (τ∈{0.25,0.5,0.75,1.0}\tau\in\{0.25,0.5,0.75,1.0\}) and another 30 biased SUN Attribute datasets (τ∈{0.5,0.75,1.0}\tau\in\{0.5,0.75,1.0\}) to learn 70 CNNs. Fig. 6 shows KL divergences mined from these CNNs.

The experiment demonstrates that large KL divergences successfully reflected potentially biased representations, but the level of annotation bias was not proportional to the level of representation bias. When we reduced the annotation bias τ\tau, the corresponding KL divergence usually decreased. At the meanwhile, CNN representations had different sensitiveness to different types of annotation bias. Small bias w.r.t. some pairs of attributes (e.g. heavy makeup and pointy nose) led to huge representation bias. Whereas, the CNN was robust to annotation biases of other pairs of attributes. For example, it was easy for the CNN to extract correct inference patterns for male and oval face, so small annotation bias of these two attributes did not cause a significant representation bias (see the lowest gray curve in Fig. 6(left)).

Experiment 3, the discovery of blind spots and failure modes: In this experiment, we mined blind spots and failure modes. We obtained five blind spots from the CNN for the CelebA dataset, i.e. the CNN did not encode positive relationships between attractive and each of earrings, necktie, and necklace, the negative relationship between chubby and oval face, and the negative relationship between bangs and wearing hat. We list blind spots with top-10 KL divergences that were mined from the CNN for the SUN Attribute database in Table 2. For example, the CNN did not learn a strong negative relationship between man-made and vegetation as expected, because man-made and vegetation co-exist in many training images. This is a typical dataset bias, and the dataset should contain more images that have only one of the two attributes.

We mined failure modes with top-NN KL divergences. We compared our method with an entropy-based method for the discovery of failure modes. The competing method only used distributions of ground-truth annotations to predict potential failure modes of the CNN. For each pair of attributes (Ai,Aj)(A_{i},A_{j}), its failure mode was defined as the mode corresponding to the least training samples among all the four mode candidates (Yi∗ ⁣= ⁣+1,Yj∗ ⁣= ⁣+1)(Y_{i}^{*}\!=\!+1,Y_{j}^{*}\!=\!+1), (Yi∗ ⁣= ⁣+1,Yj∗ ⁣= ⁣−1)(Y_{i}^{*}\!=\!+1,Y_{j}^{*}\!=\!-1), (Yi∗ ⁣= ⁣−1,Yj∗ ⁣= ⁣+1)(Y_{i}^{*}\!=\!-1,Y_{j}^{*}\!=\!+1), (Yi∗ ⁣= ⁣−1,Yj∗ ⁣= ⁣−1)(Y_{i}^{*}\!=\!-1,Y_{j}^{*}\!=\!-1). The significance of the failure mode was computed as the entropy of the joint distribution of (Yi∗,Yj∗)(Y_{i}^{*},Y_{j}^{*}) among training samples4. Then, we selected failure modes with top-NN entropies as results. To evaluate the effectiveness of failure modes, we tested the CNN using testing images. Let (Yu∗ ⁣= ⁣a,Yv∗ ⁣= ⁣b)(Y_{u}^{*}\!=\!a,Y_{v}^{*}\!=\!b) (a,b∈{−1,+1}a,b\in\{-1,+1\}) be a failure mode and Acc(Yu∣IYu∗ ⁣= ⁣a)Acc(Y_{u}|{\bf I}_{Y_{u}^{*}\!=\!a}) denote the accuracy for estimating AuA_{u} on testing images with Yu∗ ⁣= ⁣aY_{u}^{*}\!=\!a. Then, [Acc(Yu∣IYu∗=a)+Acc(Yv∣IYv∗=b)]/2[Acc(Y_{u}|{\bf I}_{Y_{u}^{*}=a})+Acc(Y_{v}|{\bf I}_{Y_{v}^{*}=b})]/2 measures the accuracy on ordinary images, and [Acc(Yu∣IYu∗=a,Yv∗=b)+Acc(Yv∣IYu∗=a,Yv∗=b)]/2[Acc(Y_{u}|{\bf I}_{Y_{u}^{*}=a,Y_{v}^{*}=b})+Acc(Y_{v}|{\bf I}_{Y_{u}^{*}=a,Y_{v}^{*}=b})]/2 measures the accuracy on images with the failure modes. Fig. 8 compares the accuracy decrease caused by the top-10 failure modes mined by our method and the accuracy decrease caused by the top-10 failure modes produced by the competing method. Table 1 shows accuracy decreases caused by different numbers of failure modes. It showed that our method extracted more reliable failure modes. Fig. 7 further visualizes biased representations corresponding to some failure modes. Intuitively, failure modes in Fig. 8 also fit biased co-appearance of two attributes among training images.

Justification of the methodology: Incorrect representations are usually caused by dataset bias and the over-fitting problem. For example, if A1A_{1} often has a positive annotation when A2A_{2} is labeled positive (or negative), the CNN may use A2A_{2}’s features as a contextual information to describe A1A_{1}. However, in real applications, it is difficult to predict whether the algorithm will suffer from such dataset bias before the learning process. For example, when the conditional distribution P(Y1∗∣Y2∗>0)P(Y_{1}^{*}|Y_{2}^{*}>0) is biased (e.g. P(Y1∗>0∣Y2∗>0)>P(Y1∗<0∣Y2∗>0)P(Y_{1}^{*}>0|Y_{2}^{*}>0)>P(Y_{1}^{*}<0|Y_{2}^{*}>0)) but P(Y1∗∣Y2∗<0)P(Y_{1}^{*}|Y_{2}^{*}<0) has a balance distribution, it is difficult to predict whether the CNN will consider A1A_{1} and A2A_{2} are positively related to each other.

Let us discuss two toy cases of this problem for simplification. Let us assume that the CNN mainly extracts common features from positive samples with Y2∗>0Y_{2}^{*}>0 to represent A2A_{2}, and regards negative samples with Y2∗<0Y_{2}^{*}<0 as random samples without sharing common features. In this case, the conditional distribution P(Y1∗∣Y2∗>0)P(Y_{1}^{*}|Y_{2}^{*}>0) will probably control the relationships between A1A_{1} and A2A_{2}. Whereas, if the CNN mainly extracts features from negative samples with Y2∗<0Y_{2}^{*}<0 to represent A2A_{2}, then the attribute relationship will not be sensitive to the conditional distribution P(Y1∗∣Y2∗>0)P(Y_{1}^{*}|Y_{2}^{*}>0).

Therefore, as shown in Fig. 8 and Table 1, our method is more effective in the discovery of failure modes than the method based on the entropy of annotation distributions.

Summary and discussion

In this paper, we have designed a method to explore inner conflicts inside representations of a pre-trained CNN without given any additional testing samples. This study focuses on an essential yet commonly ignored issue in artificial intelligence, i.e. how can we ensure the CNN learns what we expect it to learn. When there is a dataset bias, the CNN may use unreliable contexts to represent an attribute. Our method mines failure modes of a CNN, which can potentially guide the collection of new training samples. Experiments have demonstrated the high correlations between the mined KL divergences and dataset bias and shown the effectiveness in the discovery of failure modes.

In this paper, we used Gaussian distributions to approximate ground-truth distributions of attribute relationships to simplify the story. However, our method can be extended and use more complex distributions according to each specific application. In addition, it is difficult to say all discovered representation biases are “definitely” incorrect representations. For example, the CNN may use rosy cheeks to identify the wearing lipstick attribute, but these two attributes are “indirectly” related to each other. It is problematic to annotate the two attributes are either positively related or not related to each other. The wearing necktie attribute is directly related to the male attribute, but is indirectly related to the mustache attribute, because the necktie and the mustache describe different parts of the face. If we label wearing necktie is not related to mustache, then our method will examine whether the CNN uses mustache as contexts to describe the necktie. Similarly, if we consider such an indirect relationship as reliable contexts, we can simply annotate a positive relationship between necktie and mustache. Moreover, if neither the “not-related” relationship nor the positive relationship between the two attributes is trustworthy, we can simply ignore such relationships to avoid the risk of incorrect ground truth. In the future work, we would encode ground-truth attribute relationships as a prior into the end-to-end learning of CNNs, in order to achieve more reasonable representations.

Acknowledgement

This work is supported by ONR MURI project N00014-16-1-2007 and DARPA XAI Award N66001-17-2-4029, and NSF IIS 1423305.

References