Unmasking DeepFakes with simple Features

Ricard Durall, Margret Keuper, Franz-Josef Pfreundt, Janis Keuper

I Introduction

Over the last years, the increasing sophistication of smartphones and the growth of social networks have led to a gigantic amount of new digital object contents. This tremendous use of digital images has been followed by a rise of techniques to alter image contents. Until recently, such techniques were beyond the reach of most users since they were dull and time-consuming and they required a high domain expertise on computer vision. Nevertheless, thanks to the recent advances of machine learning and the accessibility to large-volume training data, those limitations have gradually faded away. As a consequence, the time for fabrication and manipulation of digital contents has significantly decreased, allowing even amateur users the modification of contents at their will.

In particular, deep generative models have lately been extensively used to produce artificial images with realistic appearance. Theses models are based on deep neural networks which are able to approximate the true data distribution of a given training set. Hence, one can sample from the learned distribution and add variations. Two of the most commonly used and efficient approaches are Variational Autoencoders (VAE) and Generative Adversarial Networks (GAN). Especially GAN approaches have lately been pushing the limits of state-of-the-art results, improving the resolution and quality of images produced . As a result, deep generative models are opening the door to a new vein of AI-based fake image generation leading to a fast dissemination of high quality tampered image content. While significant developments have been made for image forgery detection, it still remains a hard task since most current methods rely on deep learning approaches, which require large amounts of labeled training data.

In this paper, we address the problem of detecting these artificial image contents, more specifically, fake faces. In order to determine the nature of these pictures, we introduce a new machine learning based method. Our approach relies on a classical frequency analysis of the images that reveals different behaviors at high frequencies. Fig. 1 shows how different a certain range of frequency components behave when the images have been artificially generated.

Our method detects such artifacts by analyzing the frequency domain followed by a simple supervised or unsupervised classifier. Notice that this suggested pipeline does not involve nor requires vast quantities of data, which is a very convenient property for those scenarios that suffer from data scarcity. In addition, we introduce a new data set Faces-HQ, which we used to complement the CelebA data set and FaceForensics++ data set , for our experimental evaluation.

Overall, our contributions are summarized as follows:

We introduce a novel classification pipeline for artificial face detection based on a frequency domain analysis.

We provide a public data set (Faces-HQ) of high quality images containing real and fake faces from a set of different public databases.

We demonstrate how we successfully learn to detect forgery: extensive experiments on high and medium-resolution images of the Faces-HQ and CelebA data sets showed 100% accuracy. Additionally, the evaluation of the FaceForensics++ data set with low-resolution videos reached 91% accuracy.

II Related Work

In this section, we briefly review the related seminal work on high-resolution artificial images forensics and deepfake images, as well as forgery detection. In particular, we focus our attention on spotting tampered images generated by GAN-based methods.

Traditional image forensics methods can be classified according to the image features that they target, such as local noise estimation , pattern analysis , illumination modeling and steganalysis feature classification . However, with the deep learning breakthrough, the computer vision community has radically steered towards neural networks techniques. For example, are recent works based on Convolutional Neural Networks (CNN). These CNN-based approaches also aim to capture the aforementioned image features, but in an inexplicit way.

In 2014, Goodfellow et al. introduced an adversarial framework (GAN) which marked a milestone in generative models. In particular, the image generation has been improved significantly, leading to a striking progress on artificial faces among others. As a consequence, new image and video manipulation techniques known as DeepFake have emerged and established themselves online over the last few months. This occurrence of events on digital image forensics has been drawing an ever increasing attention trying to detect GAN generated images or videos.

The lack of eye blinking is one drawback observed, when the videos are artificially created. This is due to the scarcity of training images including photographs with the subject’s eyes closed. Nevertheless, this detection can be circumvented by adding images with closed eyes in training. Finding unnatural head poses, is also an extend technique , in order to detect tampered digital contents. On the other hand, the works analyze the color-space features from GAN generated images and real images, and use the disparity to classify them.

Other approaches , rather than leveraging explicit lacks or failures, rely on CNNs to distinguish GAN’s output from real images. In the same vein, introduces deep forgery discriminator with a contrastive loss function and incorporates temporal domain information by employing Recurrent Neural Networks (RNNs) upon CNNs. While deep learning methods show promising performance, a key concern is that all these methods can be easily learnt by the GAN. In particular, by incorporating them in to the GAN’s discriminator, the generator can be fine-tuned to learn a countermeasure for any differentiable forensic.

III Method

In the following section, we describe our approach in detail. Figure 2) gives an overview of the processing steps.

Frequency domain analysis is of utmost importance in signal processing theory and applications. In particular in the computer vision domain, the repetitive nature or the frequency characteristics of images can be analyzed on a space defined by Fourier transform. Such transformation consists in a spectral decomposition of the input data indicating how the signal’s energy is distributed over a range of frequencies. Methods based on frequency domain analysis have shown wide applications in image processing, such as image analysis, image filtering, image reconstruction and image compression.

The Discrete Fourier Transform (DFT) is a mathematical technique to decompose a discrete signal into sinusoidal components of various frequencies ranging from 0 (i.e., constant frequency, corresponding to the image mean value) up to the maximum representable frequency, given the spatial resolution. It is the discrete analogon of the continuous Fourier Transform for signals sampled on equidistant points. For 2-dimensional data of size M×NM\times N, it can be computed as

The frequency-domain representation of a signal (XkX_{k}) carries information about the signal’s amplitude and phase at each frequency. Fig. 2 depicts the complex output information (power and phase). Notice that the amplitude spectrum is the square root power spectrum.

III-A2 Azimuthal Average

After applying a Fourier Transform to a sample image, the information is represented in a new domain but within the same dimensionality. Therefore, given that we work with images, the output still contains 2D information. We apply azimuthal averaging to compute a robust 1D representation of the FFT power spectrum. It can be seen as a compression, gathering and averaging similar frequency components into a vector of features. In this way, we can reduce the amount of features without losing relevant information. Furthermore, throughout this compression, we achieve a more robust representation of the input. Fig. 4 shows a visual example of such method.

III-B Classifier Algorithms

Classification is the task to learn a general mapping from the attribute space to descrete classes, using specific examples of instances, each represented by a vector of attribute values and their acording lable.

One of the technically simplest (linear) classification algorithms is the Logistic Regression (LR). It is a simple statistical model that employs a logistic function (see Formula (4)) to model a binary dependent variable. The output from the hypothesis hh is the estimated probability. This is used to infer how confident predicted value can be given an input x. Logistic regression is formulated as

The underlying algorithm of maximum likelihood estimation determines the regression coefficient w for the model that accurately predicts and fits the probability of the binary dependent variable. The algorithm stops when the convergence criterion is met or the maximum number of iterations is reached.

III-B2 Support Vector Machines

Support Vector Machines (SVMs) are among the most widely used learning algorithms for (non-linear) data classification. The target of the SVM formulation is to produce a model (based on the training data) which will identify an optimal separating hyperplane, maximizing the margin between different classes. Given a training set of instance-label pairs (xi,yix_{i},y_{i}), i=1,...,li=1,...,l where xi∈Rnx_{i}\in R^{n} and y∈{1,−1}l\textbf{y}\in\{1,-1\}^{l}, Training of SVMs is implemented by the solution of the following optimization problem

where w and bb are the parameters of our classifier, ξ\xi is the slack variable and C>0C>0 the penalty parameter of the error term.

Here training vectors xix_{i} are mapped into a higher dimensional space by the function ϕ\phi. The training objective of SVMs is to find a linear separating hyperplane with the maximal margin in this higher dimensional space.

III-B3 K-Means Clustering

While supervised classification algorithms like SVM and LR rely on labeled training example to learn a classification, we also want to test the detection performance in the absence of any labeled data. Clustering is an unsupervised machine learning technique which finds similarities in the data points and group similar data points together. The key assumption is that nearby points in the feature space exhibit similar qualities and they can be clustered together. Clustering can be done using different techniques like K-means clustering.

The K-means objective function is defined as

where KK and mm are the number of clusters and samples respectively. A common approach to heuristically approximate solutions is to iteratively identify nearby features based on the distances calculated from initial centroids μ\mu. Then, these features are assigned to the closest cluster and the centroids are re-estimated. Since the amount of clusters is determined by the user, it can be easily employed in classification where we divide data into KK clusters with KK equal to or greater than the number of classes.

IV Experiments

In this section, we show results for a series of experiments evaluating the effectiveness of our approach. First, we introduce a new high-resolution data set, called Faces-HQ, together with its training settings and experiments, and we discuss our results in detail. In order to verify our approach, we also evaluate on the CelebA data set , which contains medium-resolution images, and on the FaceForensics++ data set , which contains low-resolution video sequences.

to the best of our knowledge, currently no public data set is providing high resolution images with annotated fake and real faces. Therefore, we have created our own data set from established sources, called Faces-HQFaces-HQ data has a size of 19GB. Download: https://cutt.ly/6enDLYG. In order to have a sufficient variety of faces, we have chosen to download and label the images available from the CelebA-HQ data set , Flickr-Faces-HQ data set , 100K Faces project and www.thispersondoesnotexist.com. In total, we have collected 40K high quality images, half of them real and the other half fake faces. Table I contains a summary.

IV-A2 Training Setting

as shown in Fig. 2, our pipeline is split into two parts. On the one hand, at pre-processing time, we take the whole data set and we transform every sample from the spatial domain to the 1D frequency domain, reducing 1024x1024x3 high quality color images to 722 features (1D Power Spectrum). This method is formed by a Discrete Fourier Transform followed by an azimuthally average. The transformation can be substantially optimized by employing the Fast Fourier Transform. Notice that after applying the transformation, we use only the power spectrum since it already contains enough information for the classifier. A first visualization (see Fig. 5) using t-Distributed Stochastic Neighbor Embedding (t-SNE) reveals a clear clustering of fake and real samples in this feature space.

On the other hand, once the pre-processing step is finished, we start training the classifier engine. First of all, we divide the transformed data into training and testing sets, with 20% for the testing stage and use the remaining 80% as the training set. Then, we train a classifier with the training data and finally evaluate the accuracy on the testing set. Our goal is to distinguish, real and fake faces, thus we need to use a binary classifier.

IV-A3 Method 1D Power Spectrum

looking at Fig. 6, one can observe that there is a certain repetitive behavior or pattern on the 1D Power Spectrum on those images that belong to the same class. Just by checking individual samples, it is possible to conclude that real and fake images behave in noticeable different spectra at high frequencies, and therefore they can be easily classified.

Driven by this phenomenon, we have evaluated a significant subset of images (4000 in total, 1000 of each sub-data set) and we have computed basic statistics to try to find a more general representation that help to simplify the problem. Fig. 1 plots the mean and the standard deviation of each sub-data set and corroborates the observable and distinguishable trend that real and fake images have. Motivated by this observations we have carried out a set of tests to determine the extent to which our approach successfully detects deepfakes and how much data is needed to train the model. In our experiments, we have implemented one classifier based on support vector machines (SVMs) with a radial basis function kernel and one based on logistic regression. We have run an initial experiment using 80% of the data for training and 20% for testing. We have utilized this configuration for different amount of samples (4000, 1000, 100, 20) equally distributed (see Table II).

After testing the effectiveness and efficiency of our transformed features, we have conducted a another round of experiments to determine the impact of different frequency components. Given the 722 features from 1D Power Spectrum, we have analyzed the relevance of different frequencies by grouping them into 28 sub-sections. Table III,Table IV and Table V show the accuracy results on SVM, logistic regression and K-means respectively. The rows indicate where the chunk of frequencies starts, and the column where it ends. For example, there is a chunk with 0.86 accuracy that contains frequencies from 100 to 300.

IV-B CelebA

CelebFaces Attributes (CelebA) data set consists of 202,599 celebrity face images with 40 variations in facial attributes. The dimensions of the face images are 178x218x3, which can be considered to be a medium resolution in our context.

IV-B2 Training Setting

In order to train our forgery detection classifier we need both real and fake images. We used the real images from the CelebA data set. On the same set, we then train a DCGAN to generate realistic but fake images. We split the data set into 162,770 images for training and 39,829 for testing, and we crop and resize the initial 178x218x3 size images to 128x128x3. Once the model is trained, we can conduct the classification experiments on medium-resolution scale.

IV-B3 Results

We follow the same procedure as in the previous experiments. Table VI) shows perfect classification accuracy in the supervised, and also very good results in unsupervised clustering.

IV-C FaceForensics++

FaceForensics++ is a forensics data set consisting of video sequences that have been modified with different automated face manipulation methods. Additionally, it is hosting DeepFakeDetection Data set. In particular, this data set contains 363 original sequences from 28 paid actors in 16 different scenes as well as over 3000 manipulated videos using DeepFakes and their corresponding binary masks. All videos contain a trackable, mostly frontal face without occlusions which enables automated tampering methods to generate realistic forgeries.

IV-C2 Training Setting

The employed pipeline for this data set is the same as for Faces-HQ data set and CelebA, but with an additional block. Since the DeepFakeDetection data set contains videos, we first need to extract the frame and then crop the inner faces from them. Due to the different content of the scenes of the videos, these cropped faces have different sizes.

The pre-processing part from the pipeline is size independent, thus no changes are required. However, this is not true for the classifiers, since they expect a fixed amount of features. Therefore, we have added an extra processing block just before the classifier that interpolates the 1D Power Spectrum to a fix size (300) and normalizes it dividing it by the th frequency component. The rest of the pipeline remains unchanged.

IV-C3 Method 1D Power Spectrum

As in the previous experiments, Fig. 7 shows that deepfake images have a noticeably different frequency characteristic. Despite of having a similar behaviour along the spatial frequency, there is a clear offset between the real the fakes that allows the images to be classified.

Table VII contains the classification accuracy for the supervised algorithms. These results confirm the robustness of frequency components as classification features. Nevertheless, in this case, we have observed a slightly different behaviour with respect to Faces-HQ accuracy results (see Table II). The problem is become harder for low-resolution inputs. Hence, the accuracy starts to decrease when the number of samples is smaller than 1000, specially, for the logistic regression.

The dependency on samples and the non-perfect classification accuracy can be understood by looking at Fig. 8. We can see how the standard deviations from the real and the deepfake statistics overlap with each other, meaning that some samples will be misclassified. As a result, it is not recommendable to reduce the number of features, since now the classifiers are much more sensitive to the number of features.

Finally, we compute the average classification rate per video, applying a simple majority vote over the single frame classifications. Table VIII shows the accuracy test results, which are relatively higher than the previous ones based on a frame by frame evaluation.

V Discussion and Conclusion

In this paper, we described and evaluated the efficacy of a new method to expose AI-generated fake faces images. Our approach is based on a high-frequency component analysis. We performed extensive experiments to demonstrate the robustness of our pipeline independently of the source image. We show that our method is able to detect high- and medium-resolution deepfake images on two data sets with data from various GANs with 100% accuracy. Low-resolution content is harder to identify since the available frequency spectrum is much smaller. Nevertheless, we are able to identify low-resolution fakes in a popular benchmark with 91% accuracy.

References