Watch your Up-Convolution: CNN Based Generative Deep Neural Networks are Failing to Reproduce Spectral Distributions

Ricard Durall, Margret Keuper, Janis Keuper

Introduction

Generative convolutional deep neural networks have recently been used in a wide range of computer vision tasks: generation of photo-realistic images , image-to-image and text-to-image translations , style transfer , image inpainting , transfer learning or even for training semantic segmentation tasks , just to name a few.

The most prominent generative neural network architectures are Generative Adversarial Networks (GAN) and Variational Auto Encoders (VAE) . Both basic approaches try to approximate a latent-space model of the underlying (image) distributions from training data samples. Given such a latent-space model, one can draw new (artificial) samples and manipulate their semantic properties in various dimensions. While both GAN and VAE approaches have been published in many different variations, e.g. with different loss functions , different latent space constraints or various deep neural network (DNN) topologies for the generator networks , all of these methods have to follow a basic data generation principle: they have to transform samples from a low dimensional (often 1D) and low resolution latent space to the high resolution (2D image) output space. Hence, these generative neural networks must provide some sort of (learnable) up-scaling properties.

While all of these generative methods are steering the learning of their model parameters by optimization of some loss function, most commonly used losses are focusing exclusively on properties of the output image space, e.g. using convolutional neural networks (CNN) as discriminator networks for the implicit loss in an image generating GAN. This approach has been shown to be sufficient in order to generate visually sound outputs and is able to capture the data (image) distribution in image-space to some extent. However, it is well known that up-scaling operations notoriously alter the spectral properties of a signal , causing high frequency distortions in the output.

In this paper, we investigate the impact of up-sampling techniques commonly used in generator networks. The top plot of Figure 1 illustrates the results of our initial experiment, backing our working hypotheses that current generative networks fail to reproduce spectral distributions. Figure 1 also shows that this effect is independent of the actual generator network.

We show the practical impact of our findings for the task of Deepfake detection. The term deepfake describes the recent phenomenon of people misusing advances in artificial face generation via deep generative neural networks to produce fake image content of celebrities and politicians. Due to the potential social impact of such fakes, deepfake detection has become a vital research topic of its own. Most approaches reported in the literature, like , are themselves relying on CNNs and thus require large amounts of annotated training data. Likewise, introduces a deep forgery discriminator with a contrastive loss function and incorporates temporal domain information by employing Recurrent Neural Networks (RNNs) on top of CNNs.

1.2 GAN Stabilization

Regularizing GANs in order to facilitate a more stable training and to avoid mode collapse has recently drawn some attention. While stabilize GAN training by unrolling the optimization of the discriminator, propose regularizations via noise as well as an efficient gradient-based approach. A stabilized GAN training based on octave convolutions has recently been proposed in . None of these approaches consider the frequency spectrum for regularization. Yet, very recently, band limited CNNs have been proposed in for image classification with compressed models. In , first observations have been made that hint towards the importance of the power spectra on model robustness, again for image classification. In contrast, we propose to leverage observations on the GAN generated frequency spectra for training stabilization.

2 Contributions

The contributions of our work can be summarized as follows:

We experimentally show the inability of current generative neural network architectures to correctly approximate the spectral distributions of training data.

We exploit these spectral distortions to propose a very simple but highly accurate detector for generated images and videos, i.e. a DeepFake detector that reaches up to 100% accuracy on public benchmarks.

Our theoretical analysis and further experiments reveal that commonly used up-sampling units, i.e. up-convolutions, are causing the observed effects.

We propose a novel spectral regularization term which is able to compensate spectral distortions.

We also show experimentally that using spectral regularization in GAN training leads to more stable models and increases the visual output quality.

The remainder of the paper is organized in as follows: Section 2 introduces common up-scaling methods and analyzes their negative effects on the spectral properties of images. In Section 3, we introduce a novel spectral-loss that allows to train generative networks that are able to compensate the up-scaling errors and generate correct spectral distributions. We evaluate our methods in Section 4 using current architectures on public benchmarks.

The Spectral Effects of Up-Convolutions

In order to analyze effects on spectral distributions, we rely on a simple but characteristic 1D representation of the Fourier power spectrum. We compute this spectral representation from the discrete Fourier Transform F\mathcal{F} of 2D (image) data II of size M×NM\times N,

via azimuthal integration over radial frequencies ϕ\phi

assuming square images→M=N\rightarrow M=N. We are aware that this notation is abusive, since F(I)\mathcal{F}(I) is discrete. However, fully correct discrete notation would only over complicated a side aspect of our work. A discrete implementations of AI is provided on https://github.com/cc-hpc-itwm/UpConv. . Figure 2 gives a schematic impression of this processing step.

2 Up-convolutions in generative DNNs

Generative neural architectures like GANs produce high dimensional outputs, e.g. images, from very low dimensional latent spaces. Hence, all of these approaches need to use some kind of up-scaling mechanism while propagating data through the network. The two most commonly used up-scaling techniques in literature and popular implementations frameworks (like TensorFlow and PyTorch ) are illustrated in Figure 3: up-convolution by interpolation (up+conv) and transposed convolution (transconv) . We use a very simple auto encoder (AE) setup (see Figure 4) for an initial investigation of the effects of up-convolution units on the spectral properties of 2d images after up-sampling. Figure 5 shows the different, but massive impact of both approaches on the frequency spectrum. Figure 6 gives a qualitative result for a reconstructed image and shows that the mistakes in the frequency spectrum are relevant for the visual appearance.

3 Theoretical Analysis

For the theoretic analysis, we consider, without loss of generality, the case of a one-dimensional signal aa and its discrete Fourier Transform a^\hat{a}

If we want to increase aa’s spatial resolution by factor 22, we get

where bj=0b_{j}=0 for ”bed of nails” interpolation (as used by transconv) and bj=aj−1+aj2b_{j}=\frac{a_{j-1}+a_{j}}{2} for bi-linear interpolation (as used by up+conv).

Let us first consider the case of bj=0b_{j}=0, i.e. ”bed of nails” interpolation. There, the second term in Eq. (2.3) is zero. The first term is similar to the original Fourier Transform, yet with the parameter kk being replaced by kˉ\bar{k}. Thus, increasing the spatial resolution by a factors of 22 leads to a scaling of the frequency axes by a factor of 12\frac{1}{2}. Let us now consider the effect from a sampling theory based viewpoint. It is

since the point-wise multiplication with the Dirac impulse comb only removes values for which aup=0a^{up}=0. Assuming a periodic signal and applying the convolution theorem , we get

by Eq. (2.3). Thus, the ”bed of nails upsampling” will create high frequency replica of the signal in a^up\hat{a}^{up}. To remove these frequency replica, the upsampled signal needs to be smoothed appropriately. All observed spatial frequencies beyond N2\frac{N}{2} are potential upsampling artifacts. While it is obvious from a theoretical point of view, we also demonstrate practically in Figure 8 that the correction of such a large frequency band is (assuming medium to high resolution images) is not possible with the commonly used 3×33\times 3 convolutional filters.

In the case of bilinear interpolation, we have bj=aj−1+aj2b_{j}=\frac{a_{j-1}+a_{j}}{2} in Eq. (2.3), which corresponds to an average filtering of the values of aa adjacent to bjb_{j}. This is equivalent to a point-wise multiplication of aupa^{up} spectrum a^up\hat{a}^{up} with a sinc function by their duality and the convolution theorem, which suppresses artificial high frequencies. Yet, the resulting spectrum is expected to be overly low in the high frequency domain.

Learning to Generate Correct Spectral Distributions

The experimental evaluations of our findings in the previous section and their application to detect generated content (see Section 4.1), raise the question if it would be possible to correct the spectral distortion induced by the up-convolution units used in generative networks. After all, usual network topologies contain learnable convolutional filters which follow the up-convolutions and potentially could correct such errors.

Since common generative network architectures are mostly exclusively using image-space based loss functions, it is not possible to capture and correct spectral distortions directly. Hence, we propose to add an additional spectral term to the generator loss:

Notice that MM is the image size and we use normalization by the 0th0^{th} coefficient (AI0AI_{0}) in order to scale the values of the azimuthal integral to $$.

The effects of adding our spectral loss to the AE setup from Section 2.2 for different values of λ\lambda are shown in Figure 7. As expected based on our theoretical analysis in sec. 2.3, the observed effects can not be corrected by a single, learned 3×33\times 3 filter, even for large values λ\lambda. We thus need to reconsider the architecture parameters.

2 Filter Sizes on Up-Convolutions

In Figure 8, we evaluate our spectral loss on the AE from Section 2.2 with respect to filter sizes and the number of convolutional layers. We consider varying decoder filter sizes from 3×33\times 3 to 11×1111\times 11 and 1 or 3 convolutional layers. While the spectral distortions from the up-sampling can not be removed with a single and even not with three 3×33\times 3 convolutions, it can be corrected by the proposed loss when more, larger filters are learned.

Experimental Evaluation

We evaluate the findings of the previous sections in three different experiments, using prominent GAN architectures on public face generation datasets. Section 4.1 shows that common face generation networks produce outputs with strong spectral distortions which can be used to detect artificial or “fake” images. In Section 4.2, we show that our spectral loss is sufficient to compensate artifacts in the frequency domain of the same data. Finally, we empirically show in Section 4.3 that spectral regularization also has positive effects on the training stability of GANs.

In this section, we show that the spectral distortions caused by the up-convolutions in state of the art GANs can be used to easily identify “fake” image data. Using only a small amount of annotated training data, or even an unsupervised setting, we are able to detect generated faces from public benchmarks with almost perfect accuracy.

We evaluate our approach on three different data sets of facial images, providing annotated data with different spacial resolutions:

FaceForensics++ contains a DeepFake detection data set with 363 original video sequences of 28 paid actors in 16 different scenes, as well as over 3000 videos with face manipulations and their corresponding binary masks. All videos contain a trackable, mostly frontal face without occlusions which enables automated tampering methods to generate realistic forgeries. The resolution of the extracted face images varies, but is usually around 80×80×380\times 80\times 3 pixels.

The CelebFaces Attributes (CelebA) dataset consists of 202,599 celebrity face images with 40 variations in facial attributes. The dimensions of the face images are 178×218×3178\times 218\times 3, which can be considered to be a medium resolution in our context.

In order to evaluate high resolution 1024×1024×31024\times 1024\times 3 images, we provide the new Faces-HQ Faces-HQ data has a size of 19GB. Download: https://cutt.ly/6enDLYG. Also refer to . data set, which is a annotated collection of 40k publicly available images from CelebA-HQ , Flickr-Faces-HQ dataset , 100K Faces project and www.thispersondoesnotexist.com.

1.2 Method

Figure 9 illustrates our simple processing pipeline, extracting spectral features from samples via azimuthal integration (see Figure 2) and then using a basic SVM classifierSVM hyper-parameters can be found in the source code for supervised and K-Means for unsupervised fake detection. For each experiment, we randomly select training sets of different sizes and use the remaining data for testing. In order to handle input images of different sizes, we normalize the 1D power spectrum by the 0th0^{th} coefficient and scale the resulting 1D feature vector to a fixed size.

1.3 Results

Figure 15 shows that real and “fake” faces form well delineated clusters in the high frequency range of our spectral feature space. The results of the experiments in Table 3 confirm that the distortions of the power spectrum, caused by the up-sampling units, are a common problem and allow an easy detection of generated content. This simple indicator even outperforms complex DNN based detection methods using large annotated training setsNote: results of all other methods as reported by . The direct comparison of methods might be biased since used the same real data but generated the fake data independently with different GANs..

2 Applying Spectral Regularization

In this section, we evaluate the effectiveness of our regularization approach on the CelebA benchmark, as in the experiment before. Based our theoretic analysis (see Section 2.3) and first AE experiments in Section 3, we extend existing GAN architectures in two ways: first, we add a spectral loss term (see Eq. (3.1)) to the generator loss. We use 10001000 unannotated real samples from the data set to estimate AIrealAI^{real}, which is needed for the computation of the spectral loss (see Eq. (3.1)). Second, we change the convolution layers after the last up-convolution unit to three filter layers with kernel size 5×55\times 5. The bottom plot of Figure1 shows the results for this experiment in direct comparison to the original GAN architectures. Several qualitative results produced without and with our proposed regularization are given in Figure 11.

3 Positive Effects of Spectral Regularization

By regularizing the spectrum, we achieve the direct benefit of producing synthetic images that not only look realistic, but also mimic the behaviour in the frequency domain. In this way, we are one step closer to sample images from the real distribution. Additionally, there is an interesting side-effect of this regularization. During our experiments, we noticed that GANs with a spectral loss term appear to be much more stable in terms of avoiding “mode-collapse” and better convergence. It is well known that GANs can suffer from challenging and unstable training procedures and there is little to no theory explaining this phenomenon. This makes it extremely hard to experiment with new generator variants, or to employ them in new domains, which drastically limits their applicability.

In order to investigate the impact of spectral regularization on the GAN training, we conduct a series of experiments. By employing a set of different baseline architectures, we assess the stability of our spectral regularization, providing quantitative results on the CelebA dataset. Our evaluation metric is the Fréchet Inception Distance (FID) , which uses the Inception-v3 network pre-trained on ImageNet to extract features from an intermediate layer.

Figures 13 and 14 show the FID evolution along the training epochs, using a baseline GAN implementation with different up-convolution units and a corresponding version with spectral loss. These results show an obvious positive effect in terms of the FID measure, where spectral regularization k.pdf a stable and low FID through out the training while unregularized GANs tend to “collapse”. Figure 12 visualizes the correlation between high FID values and failing GAN image generations.

Discussion and Conclusion

We showed that common “state of the art” convolutional generative networks, like popular GAN image generators fail to approximate the spectral distributions of real data. This finding has strong practical implications: not only can this be used to easily identify generated samples, it also implies that all approaches towards training data generation or transfer learning are fundamentally flawed and it can not be expected that current methods will be able to approximate real data distributions correctly. However, we showed that there are simple methods to fix this problem: by adding our proposed spectral regularization to the generator loss function and increasing the filter sizes of the final generator convolutions to at least 5×55\times 5, we were able to compensate the spectral errors. Experimentally, we have found strong indications that the spectral regularization has a very positive effect on the training stability of GANs. While this phenomenon needs further theoretical investigation, intuitively this makes sense as it is known that high frequent noise can have strong effects on CNN based discriminator networks, which might cause overfitting of the generator.

Source code available: https://github.com/cc-hpc-itwm/UpConv

References

Supplemental Material

The supplementary material of our paper contains additional details on the presented experiments, as well as some support experiments that might help to get a better understanding of the spectral properties of up-convolution units.

Using Spectral Distortions to Detect Deepfakes

In this section, we provide more detailed results of the experiments presented in section 4.1 of the paper.

To the best of our knowledge, currently no public dataset is providing high resolution images with annotated fake and real faces. Therefore, we have created our own data set from established sources, called Faces-HQFaces-HQ data has a size of 19GB. Download: https://cutt.ly/6enDLYG. In order to have a sufficient variety of faces, we have chosen to download and label the images available from the CelebA-HQ data set , Flickr-Faces-HQ data set , 100K Faces project and www.thispersondoesnotexist.com. In total, we have collected 40K high quality images, half of them real and the other half fake faces. Table 2 contains a summary. Training Setting: we divide the transformed data into training and testing sets, with 20% for the testing stage and use the remaining 80% as the training set. Then, we train a classifier with the training data and finally evaluate the accuracy on the testing set.

1.2 CelebA

The CelebFaces Attributes (CelebA) dataset consists of 202,599 celebrity face images with 40 variations in facial attributes. The dimensions of the face images are 178x218x3, which can be considered to be a medium-resolution in our context. Training Setting: While we can use the real images from the CelebA dataset directly, we need to generate the fake examples on our own.Therefore we use the real dataset to train one DCGAN , one DRAGAN , one LSGAN and one WGAN-GP to generate realistic fake images. We split the dataset into 162,770 images for training and 39,829 for testing, and we crop and resize the initial 178x218x3 size images to 128x128x3. Once the model is trained, we can conduct the classification experiments on medium-resolution scale.

1.3 FaceForensics++

FaceForensics++ is a collection of image forensic datasets, containing video sequences that have been modified with different automated face manipulation methods. One subset is the DeepFakeDetection Dataset, which contains 363 original sequences from 28 paid actors in 16 different scenes as well as over 3000 manipulated videos using DeepFakes and their corresponding binary masks. All videos contain a trackable, mostly frontal face without occlusions which enables automated tampering methods to generate realistic forgeries. Training Setting: the employed pipeline for this dataset is the same as for Faces-HQ dataset and CelebA, but with an additional block. Since the DeepFakeDetection dataset contains videos, we first need to extract the frame and then crop the inner faces from them. Due to the different content of the scenes of the videos, these cropped faces have different sizes. Therefore, we interpolate the 1D Power Spectrum to a fix size (300) and normalizes it dividing it by the 0th0^{th} frequency component.

2 Experimental Results

The following figures 15, 16 and 17 show the spectral (AI) distributions of all datasets. In all three cases, it is evident that a classifier should be able to separate real and fake samples. Also, based on our theoretical analysis (see section 2.3 in the paper), one can assume that the generators in used Face-HQ and FaceForensics++ datasets used up+conv based up-convolutions or successively blurred the generated images (due to the drop in high frequencies). CelebA based fakes used transconv.

Figure 18 gives some additional data examples and their according spectral properties for the FaceForensics++ data.

2.2 T-SNE Evaluation

Figure 19 shows the clustering properties of our AI features. It is quite obvious that a classifier should not have problems to separate both classes (real and fake).

2.3 Detection Results Depending on the Number of Available Samples

In this section, we show some additional results on the DeepFake detection task (table 1 in the paper). In tables 3, 4 and 5, we focus on the effect of the available number of data samples during training. As shown in the paper, our approach works quite well in an unsupervised setting and needs as little as 16 annotated training samples to achieve 100% classification accuracy in a supervised setting.

Spectral Regularization on Auto-Encoder

In this second section, we show some additional results from our AE experiments (see figure 4 of the paper).

Figure 20 shows the evaluation of the loss (see equations 10 and 11 in the paper) with and without spectral regularization for a decoder with 3 convolutional layers and 3 filters of kernel size 5×55\times 5 each.

These results show that the spectral regularization also has a positive effect on the convergence of the AE and the quality of the generated output images (in terms of MSE).

2 Effect of the Spectral Regularization

Figure 21 shows the impact of the spectral regularization on the AE problem. We can notice how both transconv and up+conv suffer from different behaviour on the frequency spectrum domain, specially in high frequency components. Nevertheless, after applying our spectral regularization technique, the results get much closer to the real 1D Power Spectrum distribution, generating images closer to the real distribution.

3 Effect of different Topologies

In this experiment, we evaluate the impact of different topology design choices. Figure 22 shows statistics of the spectral distributions for some topologies:

DCGAN_v1: a DCGAN topology with spectral regularization and one convolution layer (32 5x5 filters) after the last two up-convolutions.

DCGAN_v2: a DCGAN topology with spectral regularization and two convolution layers (32 5x5 filters) after the last up-convolution.

DCGAN_v3: a DCGAN topology with spectral regularization and one convolution layer (32 5x5 filters) after the every up-convolution.

DCGAN_v4: a DCGAN topology with spectral regularization and three convolution layers (32 5x5 filters) after the last up-convolution.

Following the theoretical analysis and after a rough topology search for verification, we conclude that it is sufficient to add 3 5x5 convolutional layers after the last up-convolution in order to utilize the spectral regularization.