Robust Conditional GAN from Uncertainty-Aware Pairwise Comparisons
Ligong Han, Ruijiang Gao, Mun Kim, Xin Tao, Bo Liu, Dimitris Metaxas
Introduction
Generative adversarial networks (GAN) have shown great success in producing high-quality realistic imagery by training a set of networks to generate images of a target distribution via an adversarial setting between a generator and a discriminator. New architectures have also been developed for adversarial learning such as conditional GAN (CGAN) which feeds a class or an attribute label for a model to learn to generate images conditioned on that label. The superior performance of CGAN makes it favorable for many problems in artificial intelligence (AI) such as image attribute editing.
However, this task faces a major challenge from the lack of massive labeled images with varying attributes. Many recent works attempt to alleviate such problems using semi-supervised or unsupervised conditional image synthesis . These methods mainly focus on conditioning the model on categorical pseudo-labels using self-supervised image feature clustering. However, attributes are often continuous-valued, for example, the stroke thickness of MNIST digits. In such cases, applying unsupervised clustering would be difficult since features are most likely to be grouped by salient attributes (like identities) rather than any other attributes of interest. In this work, to disentangle the target attribute from the rest, we focus on learning from weak supervisions in the form of pairwise comparisons.
Pairwise comparisons. Collecting human preferences on pairs of alternatives, rather than evaluating absolute individual intensities, is intuitively appealing, and more importantly, supported by evidence from cognitive psychology . As pointed out by ? (?), we consider relative attribute annotation because they are (1) easier to obtain than total orders, (2) more accurate than absolute attribute intensities, and (3) more reliable in application like crowd-sourcing. For example, it would be hard for an annotator to accurately quantify the attractiveness of a person’s look, but much easier to decide which one is preferred given two candidates. Moreover, attributes in images are often subjective. Different annotators have different criteria in their mind, which leads to noisy annotations .
Thus, instead of assigning an absolute attribute value to an image, we allow the model to learn to rank and assign a relative order between two images . This method alleviates the aforementioned problem of lacking continuously valued annotations by learning to rank using pairwise comparisons.
Weakly supervised GANs. Our main idea is to substitute the full supervision with the attribute ratings learned from weak supervisions, as illustrated in Figure 1. To do so, we draw inspiration from the Elo rating system and design a Bayesian Siamese network to learn a rating function with uncertainty estimations. Then, for image synthesis, motivated by we use “corrupted” labels for adversarial training. The proposed framework can (1) learn from pairwise comparisons, (2) estimate the uncertainty of predicted attribute ratings, and (3) offer quantitative controls in the presence of a small portion of absolute annotations. Our contributions can be summarized as follows.
We propose a weakly supervised generative adversarial network, PC-GAN, from pairwise comparisons for image attribute manipulation. To the best of our knowledge, this is the first GAN framework considering relative attribute orders.
We use a novel attribute rating network motivated from the Elo rating system, which models the latent score underlying each item and tracks the uncertainty of the predicted ratings.
We extend the robust conditional GAN to continuous-value setting, and show that the performance can be boosted by incorporating the predicted uncertainties from the rating network.
We analyze the sample complexity which shows that this weakly supervised approach can save annotation effort. Experimental results show that PC-GAN is competitive with fully-supervised models, while surpassing unsupervised methods by a large margin.
Related Work
Learning to rank. Our work focuses on finding “scores” for each item (e.g. player’s rating) in addition to obtaining a ranking. The popular Bradley-Terry-Luce (BTL) model postulates a set of latent scores underlying all items, and the Elo system corresponds to the logistic variant of the BTL model. Numerous algorithms have been proposed since then. To name a few, TrueSkill considers a generalized Elo system in the Bayesian view. Rank Centrality builds on spectral ranking and interprets the scores as the stationary probability under the random walk over comparison graphs. However, these methods are not designed for amortized inference, i.e. the model should be able to score (or extrapolate) an unseen item for which no comparisons are given. Apart from TrueSkill and Rank Centrality, the most relevant work is the RankNet . Despite being amortized, RankNet is homoscedastic and falls short of a principled justification as well as providing uncertainty estimations.
Weakly supervised learning. Weakly-supervised learning focuses on learning from coarse annotations. It is useful because acquiring annotations can be very costly. A close weakly supervised setting to our problem is which learns the spatial extent of relative attributes using pairwise comparisons and gives an attribute intensity estimation. However, most facial attributes like attractiveness and age are not localized features thus cannot be exploited by local regions. In contrast, our work uses this relative attribute intensity for attribute transfer and manipulation.
Uncertainty. There are two uncertainty measures one can model: aleatoric uncertainty and epistemic uncertainty. The epistemic uncertainty captures the variance of model predictions caused by lack of sufficient data; the aleatoric uncertainty represents the inherent noise underlying the data . In this work, we leverage Bayesian neural networks as a powerful tool to model uncertainties in the Elo rating network.
Robust conditional GAN (RCGAN). Conditioning on the estimated ratings, a normal conditional generative model can be vulnerable under bad estimations. To this end, recent research introduces noise robustness to GANs. ? (?) apply a differentiable corruption to the output of the generator before feeding it into the discriminator. Similarly, RCGAN proposes to corrupt the categorical label for conditional GANs and provides theoretical guarantees. Both methods have shown great denoising performance when noisy observations are present. To address our problem, we extend RCGAN to the continuous-value setting and incorporate uncertainties to guide the image generation.
Image attribute editing. There are many recent GAN-style architectures focusing on image attribute editing. IPCGAN [2018b] proposes an identity preserving loss for facial attribute editing. ? (?) propose cycle consistency loss that can learn the unpaired translation between image and attribute. BiGAN/ALI learns an inverse mapping between image-and-attribute pairs.
There exists another line of research that is not GAN-based. Deep feature interpolation (DFI) relies on linear interpolation of deep convolutional features. It is also weakly-supervised in the sense that it requires two domains of images (e.g. young or old) with inexact annotations . DFI demonstrates high-fidelity results on facial style transfer. While, the generated pixels look unnatural when the desired attribute intensity takes extreme values, we also find that DFI cannot control the attribute intensity quantitatively. ? (?) considers a binary setting and sets qualitatively the intensity of the attribute. Unlike prior research, our method uses weak supervision in the form of pairwise comparisons and leverages uncertainty together with noise-tolerant adversarial learning to yield a robust performance in image attribute editing.
Pairwise Comparison GAN
In this section, we introduce the proposed method for pairwise weakly-supervised visual attribute editing. Denote an image collection as and ’s underlying absolute attribute values as . Given a set of pairwise comparisons (e.g., or , where ), our goal is to generate a realistic image quantitatively with a different desired attribute intensity, for example, from 20 years old to 50 years old. The proposed framework consists of an Elo rating network followed by a noise-robust conditional GAN.
The designed attribute rating module is motivated by the Elo rating system , which is widely used to evaluate the relative levels of skills between players in zero-sum games. Elo rating from a player is represented as a scalar value which is adjusted based on the outcome of games. We apply this idea to image attribute editing by considering each image as a player and comparison pairs as games with outcomes. Then we learn a rating function.
Elo rating system. The Elo system assumes the performance of each player is normally distributed. For example, if Player A has a rating of and Player B has a rating of , the probability of Player A winning the game against Player B can be predicted by . We use to denote the actual score that Player A obtains after the game, which can be valued as After each game, the player’s rating is updated according to the difference between the prediction and the actual score by , where is a constant.
Image pair rating prediction network. Given an image pair and a certain attribute , we propose to use a neural network for predicting the relative attribute relationship between and . This design allows amortized inference, that is, the rating prediction network can provide ratings for both seen and unseen data. The model structure is illustrated in Figure 2.
The network contains two input branches fed with and . For each image , we propose to learn its rating value by an encoder network . Assuming the rating value of follows a normal distribution, that is , we employ the reparameterization trick , (where ). After obtaining each image’s latent rating and , we formulate the pair-wise attribute comparison prediction as where sigm is the sigmoid function. Then, the predictive probability of winning is obtained by integrating out the latent variables and ,
and . The above integration is intractable, and can be approximated by Monte Carlo, . We denote the ground-truth of and as and . The ranking loss can be formulated with a logistic-type function, that is
Noticing that is biased, an alternative unbiased upper bound can be derived as
In practice, we find that performs slightly better than .
We further consider a Bayesian variant of . The Bayesian neural network is shown to be able to provide the epistemic uncertainty of the model by estimating the posterior over network weights in network parameter training . Specifically, let be an approximation of the true posterior where denotes the parameter of , we measure the difference between and with the KL-divergence. The overall learning objective is the negative evidence lower bound (ELBO) ,
? (?) propose to view dropout together with weight decay as a Bayesian approximation, where sampling from is equivalent to performing dropout and the KL term in Equation 4 becomes regularization (or weight decay) on .
The predictive uncertainty of rating for image can be approximated using:
with a set of sampled outputs: .
Transitivity. Notice that the transitivity does not hold because of the stochasticity in . If we fix to be zero and a non-Bayesian version is used, the Elo rating network becomes a RankNet , and transitivity holds. However, one can still maintain transitivity by avoiding reparameterization and modeling . In practice, we find that reparameterization works better.
Conditional GAN with Noisy Information
The discriminator is discriminating between real data and generated data . At the same time, is trained to fool by producing images that are both realistic and consistent with the given attribute rating. As such, the Bayesian variant of the encoder is required for considering robust conditional adversarial training.
Mutual information maximization. Besides conditioning the discriminator, to further encourage the generative process to be consistent with ratings and thus learn a disentangled representation , we add a reconstruction loss on the predictive ratings:
The above reconstruction loss can be viewed as the conditional entropy between and ,
Thus, minimizing the reconstruction loss is equivalent to maximizing the mutual information between the conditioned rating and the output image.
Full objective. Finally, the full objective can be written as:
where s control the relative importance of corresponding losses. The final objective formulates a minimax problem where we aim to solve:
Analysis of loss functions. ? (?) show that the adversarial training results in minimizing the Jensen-Shannon divergence between the true conditional and the generated conditional. Here, the approximated conditional will converge to the distribution characterized by the encoder . If is optimal, the approximated conditional will converge to the true conditional, we defer the proof in Supplementary.
GAN training. In practice, we find that the conditional generative model trains better if equal-pairs (pairs with approximately equal attribute intensities) are filtered out and only different-pairs (pairs with clearly different intensities) are remained. Comparisons of training CGAN with or without equal-pairs can be found in Supplementary.
Pair Sampling
Active learning strategies such as OHEM can be incorporated in our Elo rating network. In hard example mining, only pairs with small rating differences are queried (hard+diff/all in Table 1). In addition, to maximize the number of different-pairs we also try easy example mining (easy+diff/all in Table 1). As shown, easy examples are inferior to hard examples in terms of both rating correlations and image qualities. The reason might be that easy example mining chooses pairs with drastic differences in attribute intensity, which makes the model hard to train. Hard examples help to learn a better rating function, however, provide less amount of different-pairs for the generative model to capture attribute transitions. We therefore augment hard examples with pseudo-pairs (easy examples but with predicted labels, listed as hard+pseudo-diff/all in Table 1). The augmentation strategy works well, but in following experiments we use randomly sampled pairs because (1) the random strategy is simple and performs equally well, and (2) pseudo-labels are less reliable than queried labels.
Number of pairs. Suppose there are images in the dataset, then the possible number of pairs is upper bounded by . However, if pairs are necessary, there is no benefit of choosing pairwise comparisons over absolute label annotation. Using results from , the following proposition shows that only comparisons are needed to recover an approximate ranking.
For a constant and any , if we measure comparisons chosen uniformly with repetition, the Elo rating network will output a permutation of expected risk at most .
We also provide an empirical study in the Supplementary that supports the above proposition.
Experiments
In this section, we first present a motivating experiment on MNIST. Then we evaluate the PC-GAN in two parts: (1) learning attribute ratings, and (2) conditional image synthesis, both qualitatively and quantitatively.
Dataset. We evaluate PC-GAN on a variety of datasets for image attribute editing tasks:
Annotated MNIST provides annotations of stroke thickness for MNIST dataset.
CACD is a large dataset collected for cross-age face recognition, which includes 2,000 subjects and 163,446 images. It contains multiple images for each person which cover different ages.
UTKFace is also a large-scale face dataset with a long age span, ranging from 0 to 116 years. This dataset contains 23,709 facial images with annotations of age, gender, and ethnicity.
SCUT-FBP is specifically designed for facial beauty perception. It contains 500 Asian female portraits with attractiveness ratings (1 to 5) labeled by 75 human raters.
CelebA is a standard large-scale dataset for facial attribute editing. It consists of over 200k images, annotated with 40 binary attributes.
For the MNIST experiment, stroke thickness is the desired attribute. As illustrated in Figure 4-a, the thickness information is still entangled. But in Figure 4-b, the thickness is correctly disentangled from the rest attributes.
We use CACD and UTK for age progression, SCUT-FBP and CelebA for attractiveness experiment. Since no true relatively labeled dataset is publically available, pairs are simulated from “ground-truth” attribute intensity given in the dataset. The tie margins within which two candidates are considered equal are 10, 10, and 0.4 for CACD, UTK, and SCUT-FBP, respectively. This also simplifies the quantitative evaluation process since one can directly measure the prediction error for absolute attribute intensities. Notice that CelebA only provides binary annotations, from which pairwise comparisons are simulated. Interestingly, the Elo rating network is still able to recover approximate ratings from those binary labels.
Implementation. PC-GAN is implemented using PyTorch . Network architectures and training details are given in Supplementary. For a fair evaluation, the basic modules are kept identical across all baselines.
Rating visualization. Figure 10 presents the predicted ratings learned from CACD, UTKFace, and SCUT-FBP from left to right. The ratings learned from pairwise comparisons highly correlate with the ground-truth labels, which indicates that the rating resembles the attribute intensity well. The uncertainties v.s. ground-truth labels is visualized in Figure 11. The plots show a general trend that the model is more certain about instances with extreme attribute values than those in the middle range, which matches our intuition. Additional attention-based visualizations are given in Supplementary.
Noise resistance. As mentioned previously, not only does pairwise comparison require less annotating effort, it tends to yield more accurate annotations. Consider a simple setting: if all annotators (annotating the absolute attribute value) exhibit the same random noise with a tie margin , then the corresponding pairwise annotation with the same tie margin would absorb the noise. We provide an empirical study of the noise resistance of pairwise comparisons in Supplementary.
Conditional Image Synthesis
Baselines. We consider two unsupervised baselines CycleGAN and BiGAN, two fully-supervised baselines Disc-CGAN and Cont-CGAN, and DFI in a similar weakly-supervised setting.
CycleGAN learns an encoder (or a “generator” from images to attributes) and a generator between images and attributes simultaneously.
ALI/BiGAN learns the encoder (an inverse mapping) with a single discriminator.
Disc-CGAN/IPCGAN [2018b] takes discretized attribute intensities (one-hot embedding) as supervision.
Cont-CGAN uses the same CGAN framework as PC-GAN but ratings are replaced by true labels. It is an upper bound of PC-GAN.
Qualitative results. In Figure 9, we compare our results with all baselines. For each row, we take a source and a target image as inputs and our goal is to edit the attribute value of the source image to be equal to that of the target image. PC-GAN is competitive with fully-supervised baselines while all unsupervised methods fail to change attribute intensities.
More results are shown in Figure 5, 7, 6, where the target rating value is the average of (cluster mean) a batch of (10 to 50) labeled images. From Figure 5, we see aging characteristics like receding hairlines and wrinkles are well learned. Figure 6 shows convincing indications of rejuvenation and age progression. Figure 7 shows results for SCUT-FBP, which is inherently challenging because of the size of the dataset. Compared to datasets such as CACD, SCUT-FBP is significantly smaller, with only 500 images in total (from which we take 400 for training). Training on large datasets, as the CelebA experiment in Figure 8 shows, our model produces convincing results. We also find that the model is capable of learning important patterns that correspond to attractiveness, such as in the hairstyle and the shape of the cheek shown in Figure 7. (The result does not represent the authors’ opinion of attractiveness, but only reflects the statistics of the annotations.)
Quantitative results. For quantitative evaluations, we report in Table 2 classification accuracy (Acc) evaluated on synthesized images. In our experiments, we train classifiers to predict attribute intensities of images into discrete groups (CACD , , up to ; UTK , , up to , SCUT-FBP , , up to ). PC-GAN demonstrates comparable performance with fully-supervised baselines and are significantly better than unsupervised methods. Additional metrics are reported in the Supplementary.
AMT user studies. We also conduct user study experiments. Workers from Amazon Mechanical Turk (AMT) are asked to rate the quality of each face (good or bad) and vote to which age group a given image belongs. Then we calculate the percentage of images rated as good and the classification accuracy. Table 5 shows that PC-GAN is on a par with the fully-supervised counterparts. We conduct hypothesis testing of PC-GAN and Disc-CGAN for image quality rating, , which indicates they are not statistically different with confidence level.
Ablation Studies
Supervision. First, the comparisons in Table 2 serve as an ablation study of full, no, and weak supervision, where PC-GAN is on a par with fully-supervised and significantly better than unsupervised baselines.
GAN loss terms. Second, an ablation study of CGAN loss terms is provided in Table 3. Notice that setting some losses to zero is a special case of our full objective under different s. Although we did not extensively tune ’s values since it is not the main focus of this paper, we conclude that is the most important term in terms of image qualities.
Uncertainty. The ablation study of the effectiveness of adding Bayesian uncertainties to achieve robust conditional adversarial training is given in Table 4. The three variants considered in the table differ in how much the Bayesian neural net is involved in the whole training pipeline: CNN-CGAN is a non-Bayesian Elo rating network plus a normal CGAN, BNN-CGAN learns a Bayesian encoder and yields the average ratings for a given image, and BNN-RCGAN trains a full Bayesian encoder with a noise-robust CGAN. Results confirm that the performance can be boosted by integrating an uncertainty-aware Elo rating network and an extended robust conditional GAN.
Conclusion
In this paper, we propose a noise-robust conditional GAN framework under weak supervision for image attribute editing. Our method can learn an attribute rating function and estimate the predictive uncertainties from pairwise comparisons, which requires less annotation effort. We show in extensive experiments that the proposed PC-GAN performs competitively with the supervised baselines and significantly outperforms the unsupervised baselines.
We would like to thank Fei Deng for valuable discussions on Elo rating networks. This research is supported in part by NSF 1763523, 1747778, 1733843, and 1703883.
References
Appendix A Supplementary
In Supplementary, we first show the analysis of CGAN loss terms and give a proof of Proposition 0.1. Then we provide an empirical study of how the number of pairs varies with the size of the dataset. The preliminary results on noise resistance is also presented. Next, we show qualitative attention visualization of the Elo rating network and report additional quantitative IS and FID scores for baselines and list details of network architectures. Finally, we show additional results on conditional image synthesis.
Appendix B Analysis of Loss Terms
As a standard recall in , the adversarial training results in minimizing the Jensen-Shannon divergence between the true conditional and the generated conditional. We show that the following proposition holds:
where we assume and are sampled independently. We get the optimal discriminator by applying Euler-Lagrange equation,
Finally plugging in yields,
In addition, the reconstruction loss , cycle loss , and identity preserving loss are all non-negative. Minimizing these losses will keep the equilibrium of . If the encoder and the feature extractor are trained properly, achieves its minimum when is optimally trained. ∎
Appendix C Proof of Proposition 0.1
For , we define if and 0 otherwise, measures the extent to which should be prefered over ,
where is the ground-truth and is prediction from Elo ranking network.
as our loss function and from results in , we have the lemma:
For , any , if we sample pairs uniformly with repetition from , with probability ,
where , .
Set , for , there is so that if ,
Appendix D Number of Pairs
To experimentally verify the number of pairs needed to learn a rating, we sampled from UTKFace subsets of sizes , , , , and , and train Elo rating networks with different number of pairs for each subset. As illustrated in Figure 12, to achieve a Spearman correlation above , approximately pairs are needed, where is the size of the subset. comparisons are needed for exact recovery of ranking between objects. Through our ranking network, we need comparisons to learn rating that is close enough to the true attribute strength and also keeping the space between objects. Annotation of absolute attribute strength is very noisy and usually takes annotations because of majority voting (e.g. if 3 workers per instance), our method doesn’t require more effort in annotation and pairwise comparisons are easier to annotate comparing to absolute attribute strength, which will lead to a faster finishing time in crowd-sourcing phase.
Appendix E Noise Resistance
Considering there is noise when annotating the absolute labels. Taking age annotation as an example, we assume annotators will give an age that deviates from the true age by a random noise: , where is the tie margin in Figure 13. As shown, the correlation curve of ratings drops slowly until the noise level is too high. Although only the curve on SCUT-FBP shows superior results over the ground-truth label, the general trend is that the rating curves decrease slower than the absolute label curves. This demonstrates the Elo rating network’s potential of noise resistance.
We choose UTKFace dataset to investigate how conditional synthesis results might be affected by margins. In Table 6, Spearman correlations and Inception Scores evaluated on UTKFace under different margin values are reported.
Appendix F Attention Visualization
The proposed Elo rating network is visualized using Grad-CAM . In Figure 14-a, local regions that are critical for decision making are highlighted: for CACD and UTKFace, aging indicators such as forehead wrinkles, crow’s feet eyes (babies usually have big eyes) are highlighted; for SCUT-FBP, the gradient map highlights facial regions like eyes, nose, pimples etc. Similar to DFI, if viewing the rating as deep features, one can optimize over the input image to obtain a new image with desired attribute intensity. We thus invert the encoders to see what a “typical” image with extreme attribute intensity would look like by optimizing the average face as shown in Figure 14-b.
Appendix G IS and FID Scores
Additional Inception Scores (IS) , Fréchet Inception Distances (FID) are reported in Table 7. Classifiers for evaluating classification accuracies are also used to compute Inception Scores and as auxiliary classifiers in training Disc-CGAN/IPCGAN. The unsupervised baselines have high Inception Scores and low Fréchet Inception Distances but very low classification accuracies since their outputs are almost identical to source images. Collectively, PC-GAN demonstrates comparable performance with fully-supervised baselines and are significantly better than unsupervised methods.
Appendix H Network Architectures
We show the architectures of our Elo ranking network as well as the spatial transformer network in Table 8. Facial attribute classifiers are finetuned ResNet-18 .
Appendix I Additional Results
Additional results of our PC-GAN and two fully-supervised baselines Cont-CGAN and Disc-CGAN/IPCGAN [2018b] on CACD, UTKFace, and SCUT-FBP datasets are given in Figure 15, 16, and 17 respectively. Results for unsupervised baselines are not shown since the changes in outputs are subtle. For CACD, attribute values (from Attr0 to Attr4) correspond to ages of , , , and ; for UTK, attribute values correspond to ages of , , , and ; for SCUT-FBP, attribute values correspond to scores of , , , and , respectively.
PC-GAN, Cont-CGAN and Disc-CGAN perform similarly on CACD. Disc-GAN performs much worse on UTKFace and SCUT-FBP, presumably due to the discretization of attribute strength. For example, in SCUT-FBP, the number of images are unevenly distributed across discretized attribute groups, that is, groups with least and largest attribute strength (attractiveness) have only limited images. In this case, we are more likely to see mode collapse in Disc-CGAN. As a result, Disc-CGAN is outputting same images for Attr0 and Attr4 in Figure 8. PC-GAN and Cont-CGAN have a similar quality in synthesized images in all three datasets, which shows PC-GAN can synthesize images of same qualities using pairwise comparisons.