Wasserstein Introspective Neural Networks

Kwonjoon Lee, Weijian Xu, Fan Fan, Zhuowen Tu

Introduction

Performance within the task of supervised image classification has been vastly improved in the era of deep learning using modern convolutional neural network (CNN) based discriminative classifiers . On the other hand, unsupervised generative models in deep learning were previously attained using methods under the umbrella of graphical models — e.g., the Boltzmann machine or autoencoder architectures. However, the rich representational power seen within convolution-based (discriminative) models is not being directly enjoyed in these generative models. Later, inverting convolutional neural networks in order to convert internal representations into a real image was investigated in . Recently, generative adversarial networks (GAN) and followup works have attracted a tremendous amount of attention in machine learning and computer vision by producing high quality synthesized images by training a pair of competing models against one another in an adversarial manner. While a generator tries to create “fake” images to fool the discriminator, the discriminator attempts to discern between these “real” (given training) and “fake” images. After convergence, the generator is able to produce images faithful to the underlying data distribution.

Before the deep learning era , generative modeling had been an area with a steady pace of development . These models were guided by rigorous statistical theories which, although nice in theory, did not succeed in producing synthesized images with practical quality.

In terms of building generative models from discriminative classifiers, there have been early attempts in . In , a generative model was obtained from a repeatedly trained boosting algorithm using a weak classifier whereas used a strong classifier in order to self-generate negative examples or “pseudo-negatives”.

To address the lack of richness in representation and efficiency in synthesis, convolutional neural networks were adopted in introspective neural networks (INN) to build a single model that is simultaneously generative and discriminative. The generative modeling aspect was studied in where a sequence of CNN classifiers (10−6010-60) were trained, while the power within the classification setting was revealed in in the form of introspective convolutional networks (ICN) that used only a single CNN classifier. Although INN models point to a promising direction to obtain a single model being both a good generator and a strong discriminative classifier, a sequence of CNNs were needed to generate realistic synthesis. As a result, this requirement may serve as a possible bottleneck with respect to training complexity and model size.

Recently, a generic formulation was developed within the GAN model family to incorporate a Wasserstein objective to alleviate the well-known difficulty in GAN training. Motivated by introspective neural networks (INN) and this Wasserstein objective , we propose to adopt the Wasserstein term into the INN formulation to enhance the modeling capability. The resulting model, Wasserstein introspective neural networks (WINN) shows greatly enhanced modeling capability over INN by having 20×20\times reduction in the number of CNN classifiers.

Significance and Related Work

We make several interesting observations for WINN:

A mathematical connection between the WGAN formulation and the INN algorithm is made to better understand the overall objective function within INN.

By adopting the Wasserstein distance into INN, we are able to generate images using a single CNN in WINN with even higher quality than those by INN that uses 20 CNNs (as seen in Figure 2, 4, 5, 6, and 7; the similar underlying CNN architectures are used in WINN and INN). WINN achieves a significant reduction in model complexity over INN, making the generator more practical.

Within texture modeling, INN and WINN are able to inherently model the input image space, making the synthesis of large texture images realistic, whereas GAN projects a noise vector onto the image space making the image patch stitching more difficult (although extensions exist), as demonstrated in Figure 2.

To compare with the family of GAN models, we compute Inception scores using the standard procedure on the CIFAR-10 datasets and observed modest results. Here, we typically train 4-5 cascades to boost the numbers but WINN with one CNN is already promising. Overall, modern GAN variants (e.g., ) still outperform our WINN with better quality images. Some results are shown in Figure 7.

To test the robustness of the discriminative abilities of WINN, we directly make WINN into a discriminative classifier by training it on the standard MNIST and SVHN datasets. Not only are we able to improve over the previous ICN classifier for supervised classification, we also observe a large improvement in robustness against adversarial examples compared with the baseline CNN, ResNet, and the competing ICN.

In terms of other related work, we briefly discuss some existing methods below.

Wasserstein GAN. A closely related work to our WINN algorithm is the Wasserstein generative adversarial networks (WGAN) method . While WINN adopts the Wasserstein distance as motivated by WGAN, our overall algorithm is still within the family of introspective neural networks (INN) . WGAN on the other hand is a variant of GAN with an improvement over GAN by having an objective that is easier to train. The level of difference between WINN and WGAN is similar to that between INN and GAN . The overall comparisons between INN and GAN have been described in .

Generative ConvNets. Recently, there has also been a cluster of algorithms developed in where Langevin dynamics are adopted in generator CNNs. However, the models proposed in do not perform introspection (Figure 1) and their generator and discriminator components are still somewhat separated; thus, their generators are not used as effective discriminative classifiers to perform state-of-the-art classification on standard supervised machine learning tasks. Their training processes are also more complex than those of INN and WINN.

Deep energy models (DEMs) . DEM extends the standard density estimation by using multi-layer neural networks (MLNN) with a rather complex training procedure. The probability model in DEM includes both the raw input and the features computed by MLNN. WINN instead takes a more general and simplistic form and is easier to train (see Eq. (1)). In general, DEM belongs to the minimum description length (MDL) family models in which the maximum likelihood is achieved. WINN, instead, has a formulation being simultaneously discriminative and generative.

Introspective Neural Networks

We first briefly introduce the introspective neural network method (INNg) for generative modeling and its companion model which focuses on the classification aspect. The main motivation behind the INN work is to make a convolutional neural network classifier simultaneously discriminative and generative. A single CNN classifier is trained in an introspective manner to improve the standard supervised classification result , however, a sequence of CNNs (typically 10−6010-60) is needed to be able to synthesize images of good quality .

where η∼N(0,ϵ)\eta\sim\mathcal{N}(0,\epsilon) is a Gaussian distribution and ϵ\epsilon is the step size that is annealed in the sampling process. Overall, we desire

using the iterative reclassification-by-synthesis process guided by Eq. (1).

2 Connection to the Wasserstein distance

The overall training process, reclassification-by-synthesis, is carried out iteratively without an explicit objective function. The generative adversarial network (GAN) model instead has an objective function formulated in a minimax fashion with the generator and discriminator competing against each other. The Wasserstein generative adversarial network (WGAN) work improves GAN by replacing the Jensen-Shannon distance with an efficient approximation of the Earth-Mover distance . Also, there has been further generalization of the GAN family models in .

Let p+(x)≡p(x∣y=+1)p^{+}({\bf x})\equiv p({\bf x}|y=+1) be the target distribution and pW−(x)≡p(x∣y=−1;W)p_{{\bf\mathsf{W}}}^{-}({\bf x})\equiv p({\bf x}|y=-1;{\bf\mathsf{W}}) be the pseudo-negative distribution parameterized by W{\bf\mathsf{W}}. Next, we show a connection between the INN framework and the WGAN formulation , whose objective (rewritten with our notations) can be defined as

where ∣∣f∣∣L≤1||f||_{L}\leq 1 denotes the space of 1-Lipschitz functions. To build the connection between Eq. (3) of WGAN and Eq. (1) of INN, we first present the following lemma.

Considering f(x)=ln⁡p+(x)pW−(x)f({\bf x})=\ln\frac{p^{+}({\bf x})}{p_{{\bf\mathsf{W}}}^{-}({\bf x})} and assuming its 1-Lipschitz property, we have a lower bound on the Wasserstein distance by

where KL(p∣∣q)KL(p||q) denotes the Kullback-Leibler divergence between the two distributions pp and qq, and KL(p+∣∣pW−)+KL(pW−∣∣p+)KL(p^{+}||p_{{\bf\mathsf{W}}}^{-})+KL(p_{{\bf\mathsf{W}}}^{-}||p^{+}) is the Jeffreys divergence.

Proof: see Appendix A. Note that using the Bayes’ rule, the ratio of the generative probabilities p(x∣y=+1)p(x∣y=−1)\frac{p({\bf x}|y=+1)}{p({\bf x}|y=-1)} in Lemma 1 can be turned into the ratio of the discriminative probabilities p(y=+1∣x)p(y=−1∣x)\frac{p(y=+1|{\bf x})}{p(y=-1|{\bf x})} assuming equal priors p(y=+1)=p(y=−1)p(y=+1)=p(y=-1).

The introspective neural network formulation (Eq. (1)) implicitly minimizes a lower bound of the WGAN objective (Eq. (3)).

Wasserstein Introspective Networks

Here we present the formulation for WINN building upon the formulation of the prior introspective learning works presented in Section 3.

We denote our unlabeled input training data as S+={xi∣yi=+1,i=1,…,n}S_{+}=\{{\bf x}_{i}|y_{i}=+1,i=1,\ldots,n\}. Also, we denote the set of all the self-generated pseudo-negative samples up to step tt as S−t={xi∣yi=−1,i=1,…,l}S_{-}^{t}=\{{\bf x}_{i}|y_{i}=-1,i=1,\ldots,l\}. In other words, S−tS_{-}^{t} consists of pseudo-negatives x{\bf x} sampled from our model pWt−(x)p_{{\bf\mathsf{W}}_{t}}^{-}({\bf x}) for t≥1t\geq 1 where Wt{\bf\mathsf{W}}_{t} is the model parameter vector at step tt.

Classification-step. The classification-step can be viewed as training a classifier to approximate the Wasserstein distance between S+S_{+} and S−tS_{-}^{t} for t≥1t\geq 1. Note that we also keep pseudo-negatives from earlier stages – which are essentially the mistakes of the earlier stages – to prevent the classifier forgetting what it has learned in previous stages. We use CNNs parametrized by Wt{\bf\mathsf{W}}_{t} as base classifiers. Let fWt(⋅)f_{{\bf\mathsf{W}}_{t}}(\cdot) denote the output of final fully connected layer (without passing through sigmoid nonlinearity) of the CNN. In the previous introspective learning frameworks , the classifier learning objective was to minimize the following standard cross-entropy loss function on S+∪S−tS_{+}\cup S_{-}^{t}:

where σ(⋅)\sigma(\cdot) denotes the sigmoid nonlinearity. Motivated by Section 3.2, in WINN training we wish to minimize the following Wasserstein loss function by the stochastic gradient descent algorithm via backpropagation:

To enforce the function fWtf_{{\bf\mathsf{W}}_{t}} to be 11-Lipschitz, we add the following gradient penalty term to L(Wt)\mathcal{L}({\bf\mathsf{W}}_{t}):

where x^=αx++(1−α)x−\hat{{\bf x}}=\alpha{\bf x}^{+}+(1-\alpha){{\bf x}}^{-}, x+∼p+{\bf x}^{+}\sim p^{+}, x−∼pWt−{{\bf x}}^{-}\sim p^{-}_{{\bf\mathsf{W}}_{t}}, and α∼U\alpha\sim U.

Synthesis-step. Obtaining increasingly difficult pseudo-negative samples is an integral part of the introspective learning framework, as it is crucial for tightening the decision boundary. To this end, we develop an efficient sampling procedure under the Wasserstein formulation. After the classification-step, we obtain the following distribution of pseudo-negatives:

where Zt=∫exp⁡{fWt(x)}⋅p0−(x)dxZ_{t}=\int\exp\{f_{{\bf\mathsf{W}}_{t}}({\bf x})\}\cdot p^{-}_{0}({\bf x})d{\bf x}; the initial distribution p0−(x)p_{0}^{-}({\bf x}) is a Gaussian distribution G(x;0,σ2)G({\bf x};0,\sigma^{2}) or the distribution defined in Appendix D. We find that the distribution of Appendix D encourages the diversity of sampled images. The following equivalence is shown in :

The sampling strategy of was to carry out gradient ascent on the term ln⁡p(y=+1∣x;Wt)p(y=−1∣x;Wt)\ln\frac{p(y=+1|{\bf x};{\bf\mathsf{W}}_{t})}{p(y=-1|{\bf x};{\bf\mathsf{W}}_{t})}. In Lemma 1 we chose f(x)f({\bf x}) to be ln⁡p+(x)pW−(x)\ln\frac{p^{+}({\bf x})}{p_{{\bf\mathsf{W}}}^{-}({\bf x})}. Using Bayes’ rule, it is easy to see that ∇ln⁡p(y=+1∣x;Wt)p(y=−1∣x;Wt)\nabla\ln\frac{p(y=+1|{\bf x};{\bf\mathsf{W}}_{t})}{p(y=-1|{\bf x};{\bf\mathsf{W}}_{t})} is loosely connected to ∇ln⁡p+(x)pWt−(x)\nabla\ln\frac{p^{+}({\bf x})}{p_{{\bf\mathsf{W}}_{t}}^{-}({\bf x})}. Also, argue that fWt(x)f_{{\bf\mathsf{W}}_{t}}({\bf x}) correlates with the quality of the sample x{\bf x}. This motivates us to use the following sampling strategy. After initializing x{\bf x} by drawing a fair sample from p0−(x)p_{0}^{-}({\bf x}), we increase fWt(x)f_{{\bf\mathsf{W}}_{t}}({\bf x}) using gradient ascent on the image x{\bf x} via backpropagation. Specifically, as shown in , we can obtain fair samples from the distribution pWt−p_{{\bf\mathsf{W}}_{t}}^{-} using the following update rule:

where ϵ\epsilon is a time-varying step size and η\eta is a random variable following the Gaussian distribution N(0,ϵ)N(0,\epsilon). Gaussian noise term is added to make samples cover the full distribution. Inspired by , we found that injecting noise in the image space could be substituted by applying Dropout to the higher layers of CNN. In practice, we were able obtain the samples of enough diversity without step size annealing and noise injection. As an early stopping criterion, we empirically find that the following is effective: (1) we measure the minimum and maximum fWt(⋅)f_{{\bf\mathsf{W}}_{t}}(\cdot) of positive examples; (2) we set the early stopping threshold to a random number from the uniform distribution between these two numbers. Intuitively, by matching the value of fWt(⋅)f_{{\bf\mathsf{W}}_{t}}(\cdot) positives and pseudo-negatives, we expect to obtain pseudo-negative samples that match the quality of positive samples.

2 Expanding model capacity

In practice, we find that the version with the single classifier – which we call WINN-single – is expressive enough to capture the generative distribution under variety of applications. The introspective learning formulation allows us to model more complex distributions by adding a sequence of cascaded classifiers parameterized by (W1,…,WK)({{\bf\mathsf{W}}^{1}},\ldots,{{\bf\mathsf{W}}^{K}}). Then, we can model the distribution as:

In the next sections, we demonstrate the modeling capability of WINN under cascaded classifiers, as well as its agnosticy to the type of base classifier.

3 GAN’s discriminator vs. WINN’s classifier

GAN uses the competing discriminator and the generator whereas WINN maintains a single model being both generative and discriminative. Some general comparisons between GAN and INN have been provided in . Below we make a few additional observations that are worth future exploration and discussions.

First, the generator of GAN is a cost-effective option for image patch synthesis, as it works in a feed-forward fashion. However the generator of GAN is not meant to be trained as a classifier to perform the standard classification task, while the generator in the introspective framework is also a strong classifier. Section 5.6 shows WINN to have significant robustness to external adversarial examples.

Second, the discriminator in GAN is meant to be a critic but not a generator. To show whether or not the discriminator in GAN can also be used as a generator, we train WGAN-GP on the CelebA face dataset. Using the same CNN architecture (ResNet from ) that was used as GAN’s discriminator, we also train a WINN-single model, making GAN’s discriminator and WINN-single to have the identical CNN architecture. Applying the sampling strategy to WGAN-GP’s discriminator allows us to synthesize image form WGAN-GP’s discriminator as well and we show some samples in Figure 3 (a). These synthesized images are not like faces, yet they have been classified by the discriminator of WGAN-GP as “real” faces; this demonstrates the separation between the generator and the discriminator in GAN. In contrast, images synthesized by WINN-single’s CNN classifier are faces like, as shown in Figure 3 (b).

Third, the discriminator of GAN may not be used as a direct discriminative classifier for the standard supervised learning task. As shown and discussed in ICN , the introspective framework has the ability of classification for discriminator.

Experiments

Classification-step. For training the discriminator network, we use Adam with a mini-batch size of 100. The learning rate was set to 0.0001. We set β1=0.0,\beta_{1}=0.0, and β2=0.9\beta_{2}=0.9, inspired by . Each batch consists of 50 positive images sampled from the set of positives S+S_{+} and 50 pseudo-negative images sampled from the set of pseudo-negatives S−S_{-}. In each iteration, we limit total number of training images to 10,00010,000. Synthesis-step. For synthesizing pseudo-negative images via back-propagation, we perform gradient ascent on the image space. In the first cascade, each image is initialized with a noise sampled from the distribution described in Appendix D. In the later cascades, images are initialized with the images sampled from the last cascade. We use Adam with a mini-batch size of 100. The learning rate was set to 0.01. We set β1=0.9,\beta_{1}=0.9, and β2=0.99\beta_{2}=0.99.

2 Texture modeling

We evaluate the texture modeling capability of WINN. For a fair comparison, we use the same 7 texture images presented in where each texture image has a size of 256 ×\times 256. We follow the training method of Section 5.1 except that positive images are constructed by cropping 64 ×\times 64 patches from the source texture image at random positions. We use network architecture of Appendix C. After training is done on the 64×\times64 patch-based model, we try to synthesize texture images of arbitrary size using the anysize-image-generation method following . During the synthesis process, we keep a single working image of size 320×\times320. Note that we expand the image so that center 256×\times256 pixels are covered with equal probability. In each iteration, we sample 200 patches from the working image, and perform gradient ascent on the chosen patches. For the overlapping pixels between patches, we take the average of the gradients assigned to such pixels. We show synthesized texture images in Figure 2 and 4. WINN-single shows a significant improvement over INNg-single and comparable results to INNg (using 20 CNNs). It is worth noting that leverage rich features of VGG-19 network pretrained on ImageNet. WINN and INNg instead train networks from scratch.

3 CelebA face modeling

The CelebA dataset consists of 202,599202,599 face images of celebrities. This dataset has been widely used in the previous generative modeling works since it contains large pose variations and background clutters. The network architecture adopted here is described in Appendix C. In Figure 5, we show some synthesized face images using WINN-single and WINN, as well as those by DCGAN , INNg-single, and INNg . WINN-single attains image quality even higher than that of INNg (12 CNNs).

4 SVHN modeling

SVHN consists of 32×3232\times 32 images from Google Street View. It contains 73,25773,257 training images, 26,03226,032 test images, and 531,131531,131 extra images. We use only the training images for the unsupervised SVHN modeling. We use the ResNet architecture described in . Generated images by WINN-single and WINN (4 CNN classifiers) as well as DCGAN and INN are shown in Figure 6. The improvement of WINN over INNg is evident.

5 CIFAR-10 modeling

CIFAR-10 consists of 50,00050,000 training images and 10,00010,000 test images of size 32×3232\times 32 in 10 classes. We use training images augmented by horizontal flips for unsupervised CIFAR-10 modeling. We use the ResNet given in . Figure 7 shows generated images by various models.

To measure the semantic discriminability, we compute the Inception scores on 50,00050,000 generated images. WINN shows its clear advantage over INN. WINN-5CNNs produces a result close to WGAN but there is still a gap to the state-of-the-art results by WGAN-GP.

6 Image classification and adversarial examples

To demonstrate the robustness of WINN as a discriminative classifier, we present experiments on the supervised classification tasks.

Training Methods. We add the Wasserstein loss term to the ICN loss function, obtaining the following:

where Wt=<wt(0),wt(1)1,...,wt(1)K>{\bf\mathsf{W}}_{t}=<{\bf w}_{t}^{(0)},{\bf w}_{t}^{(1)_{1}},...,{\bf w}_{t}^{(1)_{K}}>, wt(0){\bf w}_{t}^{(0)} denotes the internal parameters for the CNN, and wt(1)k{\bf w}_{t}^{(1)_{k}} denotes the top-layer weights for the kk-th class. In the experiments, we set the weight of the WINN loss, α\alpha, to 0.01. We use the vanilla network architecture resembling as the baseline CNN, which has less filters and parameters than the one in . We also use a ResNet-32 architecture with Layer Normalization on MNIST and SVHN. In the classification-step, we use Adam with a fixed learning rate of 0.001, β1\beta_{1} of 0.00.0. In the synthesis-step, we use the Adam optimizer with a learning rate of 0.020.02 and β1\beta_{1} of 0.90.9. Table 2 shows the errors on MNIST and SVHN.

Robustness to adversarial examples. It is argued in that discriminative CNN’s vulnerability to adversarial examples primarily arises due to its linear nature. Since the reclassification-by-synthesis process helps tighten the decision boundary (Figure 1), one might expect that CNNs trained with the WINN algorithm are more robust to adversarial examples. Note that unlike existing methods for adversarial defenses , our method does not train networks with specific types of adversarial examples. With test images of MNIST and SVHN, we adopt “fast gradient sign method” (ϵ=0.125\epsilon=0.125 for MNIST and ϵ=0.005\epsilon=0.005 for SVHN) to generate adversarial examples clipped to range $,whichdiffersfrom.Weexperimentwithtwonetworkshavingthesamearchitectureandonlydifferingintrainingmethod(thestandardcross−entropylossvs.theWINNprocedure).WecalltheformerasthebaselineCNN.WesummarizetheresultsinTable3.ComparedtoICN,WINNsignificantlyreducestheadversarialerrorto7.99, which differs from . We experiment with two networks having the same architecture and only differing in training method (the standard cross-entropy loss vs. the WINN procedure). We call the former as the baseline CNN. We summarize the results in Table 3. Compared to ICN , WINN significantly reduces the adversarial error to 7.99% and improves the correction rate to 90.00%. In addition, we have adopted the ResNet-32 architecture into WINN. See Table 3 and 4. We still obtain the adversarial error reduction and correction rate improvement on MNIST and SVHN (\epsilon=0.005$) with ResNet-32. Our observation is that WINN is not necessarily improving over a strong baseline for the supervised classification task but its advantage on adversarial attacks is evident.

7 Agnostic to different architectures

In Figure 8, we demonstrate our algorithm being agnostic to the type of classifier, by varying network architectures to ResNet and DenseNet . Little modification was required to adapt two architectures for WINN.

Conclusion

In this paper, we have introduced Wasserstein introspective neural networks (WINN) that produce encouraging results as a generator and a discriminative classifier at the same time. WINN is able to achieve model size reduction over the previous introspective neural networks (INN) by a factor of 2020. In most of the images shown in the paper, we find a single CNN classifier in WINN being sufficient to produce visually appealing images as well as significant error reduction against adversarial examples. WINN is agnostic to the architecture design of CNN and we demonstrate results on three networks including a vanilla CNN, ResNet , and DenseNet networks where popular CNN discriminative classifiers are turned into generative models under the WINN procedure. WINN can be adopted in a wide range of applications in computer vision such as image classification, recognition, and generation. Acknowledgements. This work is supported by NSF IIS-1717431 and NSF IIS-1618477. The authors thank Justin Lazarow, Long Jin, Hubert Le, Ying Nian Wu, Max Welling, Richard Zemel, and Tong Zhang for valuable discussions.

Appendix

Proof. Plugging f(x)=ln⁡p+(x)pW−(x)f({\bf x})=\ln\frac{p^{+}({\bf x})}{p_{{\bf\mathsf{W}}}^{-}({\bf x})} into Eq. (3), we have

The Jeffreys divergence in Eq. (4) of lemma 1 is lower and upper bounded by KL(p+∣∣pW−)KL(p^{+}||p_{{\bf\mathsf{W}}}^{-}) up to some multiplicative constant

(1+pmin+2)KL(p+∣∣pW−)≤JD(p+;pW−)≤(1+2pmin+)KL(p+∣∣pW−){\scriptstyle(1+\frac{p^{+}_{min}}{2})KL(p^{+}||p_{{\bf\mathsf{W}}}^{-})\leq JD(p^{+};p_{{\bf\mathsf{W}}}^{-})\leq(1+\frac{2}{p^{+}_{min}})KL(p^{+}||p_{{\bf\mathsf{W}}}^{-})}, where JD(p+;pW−)=KL(p+∣∣pW−)+KL(pW−∣∣p+)JD(p^{+};p_{{\bf\mathsf{W}}}^{-})=KL(p^{+}||p_{{\bf\mathsf{W}}}^{-})+KL(p_{{\bf\mathsf{W}}}^{-}||p^{+}) is the Jeffreys divergence.

Proof. Based on the Pinsker’s inequality , it is observed that

where ∣pW−−p+∣|p_{{\bf\mathsf{W}}}^{-}-p^{+}| is total variation (TV) distance. From we also have

where pmin+=min⁡xp+(x)p^{+}_{min}=\min_{{\bf x}}p^{+}({\bf x}). Applying the above bounds to KL(p+∣∣pW−)KL(p^{+}||p_{{\bf\mathsf{W}}}^{-}) and using the symmetry of the TV distance ∣pW−−p+∣≡∣p+−pW−∣|p_{{\bf\mathsf{W}}}^{-}-p^{+}|\equiv|p^{+}-p_{{\bf\mathsf{W}}}^{-}|,

Plugging the equation above into the Jeffreys divergence, we observe that KL(p+∣∣pW−)+KL(pW−∣∣p+)KL(p^{+}||p_{{\bf\mathsf{W}}}^{-})+KL(p_{{\bf\mathsf{W}}}^{-}||p^{+}) is upper and lower bounded by by KL(p+∣∣pW−)KL(p^{+}||p_{{\bf\mathsf{W}}}^{-}). □\Box

Now we can look at theorem 1. It was shown in that Eq. (1) reduces KL(p+∣∣pW−)KL(p^{+}||p_{{\bf\mathsf{W}}}^{-}), which bounds the Jeffreys divergence KL(p+∣∣pW−)+KL(pW−∣∣p+)KL(p^{+}||p_{{\bf\mathsf{W}}}^{-})+KL(p_{{\bf\mathsf{W}}}^{-}||p^{+}) as shown in corollary 1. Lemma 1 shows the connection between Jeffreys divergence and the WGAN objective (Eq. (3)) when f(x)=ln⁡p+(x)pW−(x)f({\bf x})=\ln\frac{p^{+}({\bf x})}{p_{{\bf\mathsf{W}}}^{-}({\bf x})}. We therefore can see that the formulation of introspective neural networks (Eq. (1)) connects to a lower bound of the WGAN objective (Eq. (3)). □\Box

C. Texture and CelebA Modeling Architecture. Inspired by , we design a CNN architecture for 64 ×\times 64 image as in Table 5. We use Swish non-linearity after each convolutional layer. We add Layer Normalization after each convolution except the first layer, following .

D. Alternative Initializations. We sample an initial pseudo-negative image by applying an operation defined by the network above to a tensor of size 4×4×5124\times 4\times 512 sampled from UU. The weights of the network are sampled from G(0,0.12)G(0,0.1^{2}). We do not apply any nonlinearities in the network. We add Layer Normalization after each convolution except the last layer.

References