Generative networks as inverse problems with Scattering transforms
Tomás Angles, Stéphane Mallat
Introduction
Generative Adversarial Networks (GANs) and Variational Auto-Encoders (VAEs) allow training generative networks to synthesize images of remarkable quality and complexity from Gaussian white noise. This work shows that one can train generative networks having similar properties to those obtained with GANs or VAEs without learning a discriminator or an encoder. The generator is a deep convolutional network that inverts a predefined embedding operator. To reproduce relevant properties of GAN image synthesis the embedding operator is chosen to be Lipschitz continuous to deformations, and it is implemented with a wavelet Scattering transform. Defining image generators as the solution of an inverse problem provides a mathematical framework, which is closer to standard probabilistic models such as Gaussian autoregressive models.
GANs were introduced by Goodfellow et al. (2014) as an unsupervised learning framework to estimate implicit generative models of complex data (such as natural images) by training a generative model (the generator) and a discriminative model (the discriminator) simultaneously. An implicit generative model of the random vector consists in an operator which transforms a Gaussian white noise random vector into a model of . The operator is called a generative network or generator when it is a deep convolutional network. Radford et al. (2016) introduced deep convolutional architectures for the generator and the discriminator, which result in high-quality image synthesis. They also showed that linearly modifying the vector results in a progressive deformation of the image .
Goodfellow et al. (2014) and Arjovsky et al. (2017) argue that GANs select the generator by minimizing the Jensen-Shannon divergence or the Wasserstein distance calculated from empirical estimations of these distances with generated and training images. However, Arora et al. (2017) prove that this explanation fails to pass the curse of dimensionality since estimates of Jensen-Shannon or Wasserstein distances do not generalize with a number of training examples which is polynomial on the dimension of the images. Therefore, the reason behind the generalization capacities of generative networks remains an open problem.
VAEs, introduced by Kingma & Welling (2014), provide an alternative approach to GANs, by optimizing together with its inverse on the training samples, instead of using a discriminator. The inverse is an embedding operator (the encoder) that is trained to transform into a Gaussian white noise . Therefore, the loss function to train a VAE is based on probabilistic distances which also suffer from the same dimensionality curse shown in Arora et al. (2017). Furthermore, a significant disadvantage of VAEs is that the resulting generative models produce blurred images compared with GANs.
Generative Latent Optimization (GLO) was introduced in Bojanowski et al. (2017) to eliminate the need for a GAN discriminator while restoring sharper images than VAEs. GLO still uses an autoencoder computational structure, where the latent space variables are optimized together with the generator . Despite good results, linear variations of the embedding space variables are not mapped as clearly into image deformations as in GANs, which reduces the quality of generated images.
GANs and VAEs raise many questions. Where are the deformation properties coming from? What are the characteristics of the embedding operator ? Why do these algorithms seem to generalize despite the curse of dimensionality? Learning a stable embedding which maps into a Gaussian white noise is intractable without strong prior information (Arora et al., 2017). This paper shows that this prior information is available for image generation and that one can predefine the embedding up to a linear operator. The embedding must be Lipschitz continuous to translations and deformations so that modifications of the input noise result in deformations of . Lipschitz continuity to deformations requires separating the signal variations at different scales, which leads to the use of wavelet transforms. We concentrate on wavelet Scattering transforms (Mallat, 2012), which linearize translations and provide appropriate Gaussianization. We then define the generative model as an inversion of the Scattering embedding on training data, with a deep convolutional network. The inversion is regularized by the architecture of the generative network, which is the same as the generator of a DCGAN (Radford et al., 2016). Experiments in Section 4 show that these generative Scattering networks have similar properties as GAN generators, and the synthesized images have the same quality as the ones obtained with VAEs or GLOs.
Computing a Generator from an Embedding
Unsupervised learning consists in estimating a model of a random vector of dimension from realizations of . Autoregressive Gaussian processes are simple models computed from an input Gaussian white noise by estimating a parametrized linear operator . This operator is obtained by inverting a linear operator, whose coefficients are calculated from the realizations of . We shall similarly compute models from a Gaussian white noise , by estimating a parametrized operator , but which is a deep convolutional network instead of a linear operator. is obtained by inverting an embedding of the realizations of , with a predefined operator .
Let us denote by the set of all parametrized convolutional network generators defined by a particular architecture. We impose that by minimizing an loss over the convolutional network class :
We shall define , where is a fixed normalized operator, and is an affine operator which performs a whitening and a projection to a lower dimensional space. We impose that are realizations of a Gaussian process and that transforms this process into a lower dimensional Gaussian white noise. We normalize by imposing that and that it is contractive, this is:
The affine operator performs a whitening of the distribution of the by subtracting the empirical mean and normalizing the largest eigenvalues of the empirical covariance matrix . For this purpose, we calculate the eigendecomposition of the covariance matrix . Let be the orthogonal projection in the space generated by the eigenvectors having the largest eigenvalues. We choose so that does not contract too much the distances between the , formally, this means that defines a bi-Lipschitz embedding of these samples, and hence that there exists a constant so that:
This bi-Lipschitz property must be satisfied when is equal to the dimension of and hence when is the identity. The dimension should then be adjusted to avoid reducing too much the distances between the .
We choose with , where is the identity. The resulting embedding is:
This embedding computes a -dimensional whitening of the . The generator network which inverts over the training samples is then computed according to (1).
The existence of a known embedding operator allows us to use the generative network as an associative or content addressable memory. The input can be interpreted as an address, of lower dimension than the generated image . Any training image is associated to the address . The network is optimized to generate the training images from these lower dimensional addresses. The inner network coefficients thus include a form of distributed memory of these training images. Moreover, if the network generalizes then one can approximately reconstruct a realization of the random process from its embedding address . In this sense, the memory is content addressable.
2 Gaussianization and Continuity to Deformations
We now describe the properties of the normalized embedding operator to build a generator having similar properties as GANs or VAEs. We mentioned that we would like to be nearly a Gaussian white noise and since where is affine then should be Gaussian. Therefore, should be realizations of a Gaussian process and hence be concentrated over an ellipsoid.
The normalized embedding operator must also be covariant to translations because it will be inverted by several layers of a deep convolutional generator, which are covariant to translations. Indeed, the generator belongs to which is defined by a DCGAN architecture (Radford et al., 2016). In this family of networks, the first layer is a linear operator which reshapes and adjusts the mean and covariance of the input white noise . The non-stationary part of the process is captured by this first affine transformation. The next layers are all convolutional and hence covariant to translations. These layers essentially invert the normalized embedding operator over the training samples.
The normalized embedding operator must also linearize translations and small deformations. Indeed, if the input is linearly modified then, to reproduce GAN properties, the output should be continuously deformed. Therefore, we require to be Lipschitz continuous to translations and deformations. A translation and a deformation of an image can be written as , where denotes the spatial coordinates. Let be the maximum translation amplitude. Let be the Jacobian of at and be the norm of this Jacobian matrix. The deformation induced by is measured by , which specifies the maximum scaling factor induced by the deformation. The value defines a metric on diffeomorphisms (Mallat, 2012) and thus specifies the deformation “size”. The operator is said to be Lipschitz continuous to translations and deformations over domains of scale if there exists a constant such that for all and all we have:
This inequality implies that translations of that are small relative to and small deformations are mapped by into small quasi-linear variations of .
Another Gaussianization strategy comes from the Central Limit Theorem by averaging nearly independent random variables having variances of the same order of magnitude. This averaging can be covariant to translations if implemented with convolutions with a low-pass filter. The resulting operator will also be Lipschitz continuous to deformations. However, an averaging operator loses all high-frequency information. To define an operator which satisfies the bi-Lipschitz condition (2), we must preserve the high-frequency content despite the averaging. The next section explains how to do so with a Scattering transform.
Generative Scattering Networks
In this section, we show that a Scattering transform (Mallat, 2012; Bruna & Mallat, 2013) provides an appropriate embedding for image generation, without learning. It does so by taking advantage of prior information on natural signals, such as translation and deformation properties. We also specify the architecture of the deep convolutional generator that inverts this embedding, and we summarize the algorithm to perform this regularized inversion.
Since the first order term of a deformation is a local scaling, defining an operator that is Lipschitz continuous to deformations, and hence satisfies (3), requires decomposing the signal at different scales, which is done by a wavelet transform (Mallat, 2012). The linearization of translations at a scale is obtained by an averaging implemented with a convolution with a low-pass filter at this scale. A Scattering transform uses non-linearities to compute interactions between coefficients across multiple scales, which restores some information lost due to the averaging. It also defines a bi-Lipschitz embedding (2) and the averaging at the scale Gaussianizes the random vector . This scale adjusts the trade-off between Gaussianization and contraction of distances due to the averaging.
To obtain an order two Scattering operator , the operator computes sequences of up to two wavelet convolutions and complex modulus:
Therefore, there are channels. A Scattering transform is then obtained by averaging each channel with a Gaussian low-pass filter whose spatial width is proportional to :
Convolutions with are followed by a subsampling of ; as a result, if has pixels then is of dimension where:
The maximum scale is limited by the image width . For our experiments we used and images of size . In this case and . Since for , has more coefficients than . Based on this coefficient counting, we expect to be invertible for , but not for .
If wavelet coefficients become nearly independent when they are sufficiently far away then becomes progressively more Gaussian as the scale increases, because of the spatial averaging by . Indeed, if is independent from for then the Central Limit Theorem proves that converges to a Gaussian distribution when increases. However, as the scale increases, the averaging produces progressively more instabilities which can deteriorate the bi-Lipschitz bounds in (2). This trade-off between Gaussianization and stability defines the optimal choice of the scale .
A Scattering transform is thus an instance of a deep convolutional network whose filters are specified by wavelets and where the non-linearity is chosen to be a modulus. The choice of filters is flexible and other multiscale representations, such as the ones in Portilla & Simoncelli (2000); Lyu & Simoncelli (2009); Malo & Laparra (2010), may also be used.
Following the notations in 2.1, the normalized embedding operator is chosen to be , thus the embedding operator is defined by . A generative Scattering network is a deep convolutional network which implements a regularized inversion of this embedding. Both networks are illustrated in Figure 1.
More specifically, a generative Scattering network is a deep convolutional network which inverts the whitened Scattering embedding on training samples. It is obtained by minimizing the loss , as explained in Section 2.1. The minimization is done with the Adam optimizer (Kingma & Ba, 2014), using the default hyperparameters. The generator illustrated in Figure 1, is a DCGAN generator (Radford et al., 2016), of depth :
The non-linearity is a ReLU. The first operator is linear (fully-connected) plus a bias, it transforms into a array of channels. The next operators for perform a bilinear upsampling of their input, followed by a multichannel convolution along the spatial variables, and the addition of a constant bias for each channel. The operators compute a progressive inversion of , calculated with the convolutional operators for . The last operator does not perform an upsampling. All the convolutional layers have filters of size , with symmetric padding at the boundaries. All experiments are performed with color images of dimension pixels.
Numerical experiments
This section evaluates generative Scattering networks with several experiments. The accuracy of the inversion given by the generative network is first computed by calculating the reconstruction error of training images. We assess the generalization capabilities by computing the reconstruction error on test images. Then, we evaluate the visual quality of images generated by sampling the Gaussian white noise . Finally, we verify that linear interpolations of the embedding variable produce a morphing of the generated images, as in GANs. The code to reproduce the experiments can be found in https://github.com/tomas-angles/generative-scattering-networks.
We consider three datasets that have different levels of variabilities: CelebA (Liu et al., 2015), LSUN (bedrooms) (Yu et al., 2015) and Polygon5. The last dataset consists of images of random convex polygons of at most five vertices, with random colors. All datasets consist of RGB color images with shape . For each dataset, we consider only training images and test images.
In all experiments, the Scattering averaging scale is , which linearizes translations and deformations of up to pixels. For an RGB image , is computed for each of the three color channels. Because of the subsampling by , has a spatial resolution of , with channels, and is thus of dimension . Since it has more coefficients than the input image , we expect it to be invertible, and numerical experiments indicate that this is the case. This dimension is reduced to by the whitening operator. The resulting operator remains a bi-Lipschitz embedding of the training images of the three datasets in the sense of (2). The Lipschitz constant is , and of the distances between image pairs are preserved with a Lipschitz constant smaller than . Further reducing the dimension to , which is often used in numerical experiments (Radford et al., 2016; Bojanowski et al., 2017), has a marginal effect on the Lipschitz bound and on numerical results, but it slightly affects the recovery of high-frequency details.
We now assess the generalization properties of by comparing the reconstructions of training and test samples from ; figures 2 and 3 show such reconstructions. Table 1 gives the average training and test reconstruction errors in dB for each dataset. The training error is between 3dB and 8dB above the test error, which is a sign of overfitting. However, this overfitting is not large compared to the variability of errors from one dataset to the next. Overfitting is not good for unsupervised learning where the intent is to model a probability density, but if we consider this network as an associative memory, it is not a bad property. Indeed, it means that the network performs a better recovery of known images used in training than unknown images in the test, which is needed for high precision storage of particular images.
Polygons are simple images that are much better recovered than faces in CelebA, which are simpler images than the ones in LSUN. This simplicity is related to the sparsity of their wavelet coefficients, which is higher for polygons than for faces or bedrooms. Wavelet sparsity drives the properties of Scattering coefficients which provide localized norms of wavelet coefficients. The network regularizes the inversion by storing the information needed to reconstruct the training images, which is a form of memorization. LSUN images require more memory because their wavelet coefficients are less sparse than polygons or faces; this might explain the difference of accuracies over datasets. However, the link between sparsity and the memory capacity of the network is not yet fully understood. The generative network has itself sparse activations with about of them being equal to zero on average over all images of the three datasets. Sparsity thus seems to be an essential regularization property of the network.
Similarly to VAEs and GLOs, Scattering image generations eliminate high-frequency details, even on training samples. This is due to a lack of memory capacity of the generator. This was verified by reducing the number of training images. Indeed, when using only training images, all high-frequencies are recovered, but there are not enough images to generalize well on test samples. This is different from GAN generations, where we do not observe this attenuation of high-frequencies over generated images. GANs seem to use a different strategy to cope with the memory limitation; instead of reducing precision, they seem to ”forget” some training samples (mode-dropping), as shown in Bojanowski et al. (2017). Therefore, GANs versus VAEs or generative Scattering networks introduce different types of errors, which affect diversity versus recovery precision.
To evaluate the generalization properties of the network on Gaussian white noise , Figure 4 shows images generated from random samplings of . Generated images have strong similarities with the ones in the training set for polygons and faces. The network recovers colored geometric shapes in the case of Polygon5 even tough they are not exactly polygons, and it recovers faces for CelebA with a blurred background. For LSUN, the images are piecewise regular, and most high frequencies are missing; this is due to the complexity of the dataset and the lack of memory capacity of the generative network.
Figure 5 evaluates the deformation properties of the network. Given two input images and , we modify to obtain the interpolated images:
The linear interpolation over the embedding variable produces a continuous deformation from one image to the other while colors and image intensities are also adjusted. It reproduces the morphing properties of GANs. In our case, these properties result from the Lipschitz continuity to deformations of the Scattering transform.
Conclusion
This paper shows that most properties of GANs and VAEs can be reproduced with an embedding computed with a Scattering transform, which avoids using a discriminator as in GANs or learning the embedding as in VAEs or GLOs. It also provides a mathematical framework to analyze the statistical properties of these generators through the resolution of an inverse problem, regularized by the convolutional network architecture and the sparsity of the obtained activations. Because the embedding function is known, numerical results can be evaluated on training as well as test samples.
We report preliminary numerical results with no hyperparameter optimization. The architecture of the convolutional generator may be adapted to the properties of the Scattering operator as increases. Also, the paper uses a “plain” Scattering transform which does not take into account interactions between angle and scale variables, which may also improve the representation as explained in Oyallon & Mallat (2015).
Acknowledgements
This work was funded by the ERC grant InvariantClass 320959.