Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional Entropies
Itai Gat, Idan Schwartz, Alexander Schwing, Tamir Hazan
Introduction
Multi-modal data is ubiquitous and commonly used in many real-world applications. For instance, discriminative visual question answering systems take into account the question, the image and a variety of answers. In general, we treat data as multi-modal if it can be partitioned into semantic features, e.g., color and shape can be treated as multi-modal data.
To address this issue, we develop a novel regularization term based on the functional entropy. Intuitively, this term encourages to balance the contribution of each modality to classification. To address the computational challenges of computing the functional entropy we develop a method based on the log-Sobolev inequality which bounds the functional entropy with the functional Fisher information.
We illustrate the efficacy of the proposed approach on the three challenging multi-modal datasets Colored MNIST, VQA-CPv2, and SocialIQ. We find that our regularization maximizes the utilization of essential information. We verify this empirically on the synthetic dataset Colored MNIST. We also evaluate on popular benchmarks, finding that our method permits a state-of-the-art performance on two datasets: SocialIQ (68.53% vs. 64.82%) and VQA-CPv2 (54.55% vs. 52.05%).
Related Work
Multi-modal datasets. Over the years, the amount and variety of data that has been used across tasks has grown significantly. Unsurprisingly, present-day tasks are increasingly sophisticated and combine multiple data modalities like vision, text, and audio. In particular, in the past few years, many large-scale multi-modal datasets have been proposed . Subsequently, multiple works developed strong models to address these datasets . However, recent work also suggests that many of these advanced models predict by leveraging one of the modalities more than the others, e.g., utilizing question type to determine the answer in VQA problems . This property is undesirable since multi-modal tasks consider all data essential to solve the challenge without overfitting to the dataset.
Bias in datasets. Recently, datasets were proposed to study whether a model can generalize and solve the task or whether it uses a single modalities’ features. Usually, this evaluation is performed by partitioning data into train and test sets using different distributions. For example, VQA-CP is a reshuffle of the VQA dataset ensuring that question-type distributions differ between train and test splits. Another well-known dataset is Colored MNIST . In this dataset, each digit class is colored differently in the train set, while samples in the test set remain gray-scale. Different approaches were proposed to deal with such problems: Arjovsky et al. propose to improve generalization by ensuring that the optimal classifier equals all training distributions. Wang et al. suggest to regularize the overfitting behavior to different modalities. Methods like REPAIR prevent a model from exploiting dataset biases by re-sampling the training data. Kim et al. use an adversarial approach to learn unbiased feature representations. Clark et al. and Cadene et al. suggest methods to overcome language priors using a bias-only model in VQA tasks.
Entropy and information in deep nets. Entropy plays a pivotal role in machine learning and has been extensively used in losses and for regularization . However, its use is confined to probability distributions while we use functional entropy, which has a different form and is defined for any non-negative function. More broadly, other components of information theory have been studied in deep nets, for example, the information bottleneck criteria . Other works use information theory to overcome generalization . For instance, Krishna et al. propose to maximize the mutual information of the modalities by regularizing differences between modality representations. Fisher information is also used in various machine learning and deep learning settings, e.g., monitoring of the learning process . In contrast, our work considers the functional Fisher information of a non-negative function that represents a multi-modal learner, while Fisher information is defined over probability density functions. Also, we use the log-Sobolev inequality between the functional entropy and the functional Fisher information, which does not hold for entropy and Fisher information.
Background
Beyond the loss, a typical learning process employs a regularization term which encourages use of the ‘simplest’ function. Various regularization terms that favor ‘simple’ functions pose a considerable difficulty for multi-modal problems: deep learners easily find simple functions that ignore one of the modalities. For example, a simple discriminator for Colored MNIST, which consists of monochromatic images whose colors correlate with their labels, focuses almost exclusively on the color vector to predict the label rather than also assessing the shape of the image. Formally, if the monochromatic images are represented by their color and shape modalities then the simplest discriminator will only consider the -dimensional color . In this setting, the learned function avoids all important information within the shape modality .
In the following we describe the notion of functional entropy in Section 3.1. In Section 3.2 we present the log-Sobolev inequality, which bounds the functional entropy of a non-negative function by the functional Fisher information. We conclude with the notion of tensorization, which decomposes these components according to their multi-modal spaces.
2 Functional Fisher information
3 Tensorization and multi-modal data
The tensorization of the functional entropy amounts to
Regularization by Maximizing Functional Entropies
Functional entropy requires a probability measure. In the following we differentiate between multi-modal training points and general multi-modal points in the probability measure space, which we denote by . We use the training points to determine the measure and we denote by the variable of the integrands. In our work we consider a Gaussian product. Given a training point that resides in the multi-modal space we define the measure for the -th modality to be the Gaussian distribution with mean and variance , where is the -th modality of the training point and is the variance of the coordinate of :
The measure is the product measure over the different modalities . For example, given a monochromatic image in the training data, the distribution employed by the functional entropy in Eq. (2) is .
For each training data point , we define the functional entropy over the deep net softmax function as
This function measures the sensitivity of the softmax prediction to Gaussian perturbations of the input, since the random perturbation is sampled from a Gaussian with an expected value , as described in Eq. (6).
The cross-entropy function is a non-negative function, therefore, it is natural to apply the log-Sobolev inequality for Gaussian measures to bound the functional entropy using the functional Fisher information, in Eq. (3):
We use the functional Fisher information bound in Eq. (3) to regularize the training process, in order to implicitly encourage to maximize the information of each modality, while minimizing the training loss. In order to account for both the loss minimization and the information maximization, we take the inverse information. Given multi-modal training data , our learning objective is
The hyperparameter balances between the training loss and the inverse information.
We combine this approximation with the log-Sobolev inequality to measure the amount of the functional Fisher information added by each modality, for a given multi-modal training point :
Similarly to Eq. (9), we may use the tensorized functional Fisher information bound in Eq. (11) to regularize the training process. Given multi-modal training data , our tensorized learning objective is
Connection Between Functional Entropy and Variance
Rothaus has shown a connection between the functional entropy of a non-negative function and its variance.
Particularly, when the values of the non-negative function are small, one can expand the Taylor series of to show that
where the residual function is non-negative and approaches zero faster than approaches zero, i.e., . Interestingly, a similar bound to the log-Sobolev inequality (Eq. (3)) exists for the variance of continuous random variables , which is widely known as the Poincaré inequality:
The relation between the functional entropy and the variance, expressed in Eq. (14), suggests that these bounds should behave similarly in practice. To fit the variance into multi-modal settings we need to show tensorization (as in Sec. 3.3). In the variance case, this property is called the Efron-Stein theorem (cf. , Proposition 2.2),
Next, we present a similar regularization term to the one described in Sec. 4. This time we use variance and Poincaré inequality.
In Sec. 4, we were interested in bounding the functional entropy (Eq. (8)), for each training point , of . Similarly, we want to bound the variance of . For this purpose, we can use the Poincaré inequality, described in Eq. (15),
We use the above inequality to regularize the training process. To consider both the loss minimization and the regularization term we formulate the learning objective,
To fit our multi-modal settings, we need to follow the tensorization process as illustrated in Sec. 4.1. The same tensorization process can be applied to the variance using the Poincaré bound given in Eq. (15). For tensorized Poincaré bound leads to the learning objective
Experiments
In the following, we evaluate our proposed regularization on four different datasets. One of the datasets is a synthetic dataset (Colored MNIST), which permits to study whether a classifier leverages the wrong features. We show that adding the discussed regularization improves the generalization of a given classifier. We briefly describe each dataset and discuss the results of the proposed method.
Dataset: Colored MNIST is a synthetic dataset based on MNIST . The train and validation set consist of 60,000 and 10,000 samples, respectively. Each sample is biased with a color that correlates with its digit. The biasing process assigns to each digit an RGB vector which represents a mean color. Then, each sample receives its color, sampled from a normal distribution with a fixed variance around the digit’s mean color. This process results in a monochromatic image and high correlation between the digit’s color and its label. To introduce a bias in a multi-modal approach, we split each sample into a color modality and a shape (gray-scale representation of the image) modality . For humans it is evident that a digit should be classified based on its shape and not its color. For a learner this fact is not as clear. To minimize the loss, it is much easier for a classifier to leverage the color modality, which correlates very well with the label. In its nature, Colored MNIST evaluates the generalization of a model since it has a test set that assesses whether a classifier relies solely on color or both the color and the shape.
Baseline: A simple deep net achieves high accuracy on both colored train and colored validation set. However, on the gray-scale validation set, the network fails drastically, achieving only a 41.11% accuracy when using the model from the last training epoch. We note that the more we train the more the baseline relies on color rather than shape. We also compute an upper-bound by training the deep net on a gray-scale version. The upper-bound accuracy on the gray-scale validation set is 98.47%.
Results: Adding our proposed regularization encourages to exploit information from both shape and color modalities. We provide results in Tab. 1(c). Fig. 2 shows that without entropy regularization, the Fisher information value of the shape is almost zero while adding the regularization results in a higher shape information value than the color. This fact complements the classifier’s performance on the gray-scale validation set shown in Fig. 3. Using functional Fisher information based regularization outperforms the same classifier trained without regularization by almost 55%.
2 VQA-CPv2
Dataset: VQA-CPv2 is a re-shuffle of the VQAv2 dataset. Visual question answering (VQA) requires to answer a given question-image pair. observed that the original split of the VQAv2 dataset permits to leverage language priors. To challenge models to not use these priors, the question type distributions of the train and validation set were changed to differ from one another. VQA-CPv2 consist of 438,183 samples in the train set and 219,928 samples in the test set.
Results: We evaluated our method by adding functional Fisher information regularization to the current state-of-the-art . In doing so, the result improves by 2.5%, achieving 54.55% accuracy. We provide a comparison with recent state-of-the-art methods in Tab. 2.
The authors of raise the concern that new regularization methods mainly boost the performance of yes/no questions. Investigating the improvements due to our result shows that this is not the case. The accuracy difference to the previous state-of-the-art on the different answer types is: yes/no +1.5%, number +18%, and other -1%.
3 SocialIQ
Dataset: The SocialIQ dataset is designed to develop models for understanding of social situations in videos. Each sample consists of a video clip, a question, and an answer. The task is to predict whether the answer is correct or not given this tuple. The dataset is split into 37,191 training samples, and 5,320 validation set samples. Note that an inherent bias exists in this dataset: specifically the sentiment of the answer provides a good cue.
Baseline: A simple classifier based on only the answer modality performs significantly better than chance level accuracy (using our settings ~6% more). Such biases in the train set lead to a classic case of overfitting.
Results: As seen in Fig. 3, training without functional Fisher information regularization leads to ~80% accuracy on the train set and ~64% accuracy on the validation set. Although, functional Fisher information regularization results in 70% accuracy on the train set, it improves validation set accuracy to 67.93% accuracy.
We further investigate the information values during the training phase with and without functional Fisher information regularization. In Fig. 2 we observe that without our regularization, the answer modality has the highest information value while the question modality is almost entirely ignored. Adding the proposed regularization balances the information between modalities, the desired behavior in multi-modal learning.
4 Dogs and Cats
Dataset: Following the settings of Kim et al. , we evaluate our models on the biased “Dogs and Cats” dataset. This dataset comes in two splits: The TB1 set consists of bright dogs and dark cats and contains 10,047 samples. The TB2 set consist of dark dogs and bright cats and contains 6,738 samples. We use the image as a single-modality.
Baseline: The authors show that training of ResNet-18 on TB1 and testing on TB2 results in a poor performance of 74.98%. The authors also show that using TB2 as the train set and TB1 as the test set results in even worse accuracy of 66.45%.
Functional Fisher information regularization training on TB1 and testing on TB2 with (see Eq. (12)) set to equal 3e-10 results in 94.71% accuracy, exceeding by 3.5%. Training on TB2 while testing on TB1 achieves an accuracy of 88.11%, 1% higher than .
Conclusion
Classical regularizers applied on multi-modal datasets lead to models which may ignore one or more of the modalities. This is sub-optimal as we expect all modalities to contribute to classification. To alleviate this concern we study regularization via the functional entropy. It encourages the model to more uniformly exploit the available modalities.
Acknowledgements: This work is supported in part by NSF under Grant 1718221, 2008387, NIFA award 2020-67021-32799 and, BSF under Grant 2019783.
Broader Impact
We study functional entropy based regularizers which enable classifiers to more uniformly benefit from available dataset modalities in multi-modal tasks. We think the proposed method will help to reduce biases that present-day classifiers exploit when being trained on data which contains modalities, some of which are easier to leverage than others.
We think this research will have positive societal implications. With machine learning being used more widely, bias from various modalities has become ubiquitous. Minority groups are disadvantaged by present-day AI algorithms, which work very well for the average person but are not suitable for other groups. We provide two examples next:
It is widely believed that criminal risk scores are biased against minoritieshttps://www.propublica.org/article/bias-in-criminal-risk-scores-is-mathematically-inevitable-researchers-say, and mathematical methods that reduce the bias in machine learning are desperately needed. In our work we show how our regularization allows to reduce the color modality effect in colored MNIST, which hopefully facilitates to reduce bias in deep nets.
Consider virtual assistants as another example: if pronunciation is not mainstream, replies of AI systems are less helpful. Consequently, current AI ignores parts of society.
To conclude, we think the proposed research is a first step towards machine learning becoming more inclusive.