Defending against GAN-based Deepfake Attacks via Transformation-aware Adversarial Faces
Chaofei Yang, Lei Ding, Yiran Chen, Hai Li
Introduction
Machine learning (ML) has experienced rapid development during the past decade and been widely adopted by many daily applications. The adoption of ML technologies, however, also induces new threats in data privacy and security. As a famous example, Deepfake recently draws increasing attention by offering the capability to generate a fake video of any particular person: in the fake video, an attacker can swap a person’s face with the synthesized face of another person who can be anyone such as a celebrity, a politician, or just a normal person. Although the concept of face-swap has been studied for long, the threat of Deepfake becomes much severer because of the utilization of generative models such as autoencoders and generative adversarial networks (GANs) . These techniques enable the generation of extremely realistic synthetic images with incredible details. The faces are well synthesized such that naked eyes cannot easily distinguish between a fake video and an authentic one. Deepfake can be potentially used for deceiving identity verification or defaming a person.
Many algorithms have been proposed to detect Deepfake videos . These algorithms usually adopt state-of-the-art neural network (NN) models and rely on techniques such as noise channel analysis, frame consistency detection, data augmentation, etc. These detectors are all mitigation strategies after the high-quality fake content is generated and therefore are simply passive measures against Deepfake attacks. Besides, recent research shows that these detectors are vulnerable . To the best of our knowledge, there are no reliable defense methods against Deepfake attacks yet.
This work aims to take an offensive measure to impede the generation of high-quality fake images or videos. Poisoning attack could be a solution, which compromises ML models in training phase. If the attackers can only access elaborate poisoned faces images and use them to train their Deepfake models, the synthesized faces generated by these models may not have a satisfying quality. Rather than directly implementing poisoning attacks in high complexity, we propose an effective defense method via transformation-aware adversarial faces. Here, we focus on GAN-based Deepfake attacks as GANs are the mainstream generative models which are adopted in many state-of-the-art image translation works . Specifically, we generate adversarially perturbed faces of person A based on the discriminator from a pre-trained Deepfake model. The performance of new Deepfake models trained based on these adversarial faces is degraded. This is reflected by the low quality of the synthesized faces, which are more obviously fake to human observers and can be easily detected based on various metrics. Therefore, the faces of person A are protected against Deepfake attacks.
Our major contributions are summarized as follows:
To the authors’ best knowledge, it is the first offensive method that utilizes adversarial faces to defend against GAN-based Deepfake attacks.
We propose an adversarial face generation method to protect individuals’ faces by considering random differentiable image transformations during the training of Deepfake models. This method can consistently yield more artifacts in synthesized faces, making the recognition of the induced faked images and videos much easier.
We identify the increments of adversarial and edge losses as the major causes of incurring significant degradations in the quality of synthesized faces.
We demonstrate the effectiveness and robustness of our defense method through extensive experiments using multiple pairs of faces with different resolutions, under white-, gray-, and black-box settings based on various metrics.
Related works
Detection of Deepfake videos. Much of the research surrounding Deepfake seeks to detect synthesized videos based on various features. For example, recurrent neural networks (RNNs) can be used to extract temporal information, i.e., frame-level features, for Deepfake detection . Yang et al. propose to use 3D head pose estimation as a feature and adopt the support vector machine (SVM) to classify if the face is fake or not. The use of noise analysis has also been investigated , where a two-stream NN is employed to extract both macro face features and local noise features. Li et al. propose to use image discrepancies across the blending boundary for face forgery detection by leveraging noise analysis. Another mainstream series of detection algorithms rely on data augmentation. Leveraging the imperfection of synthesized videos, e.g., warping artifacts, one can effectively distinguish Deepfake videos from benign counterparts . The training of such detectors requires elaborate augmented data with similar imperfections. The utilization of steganalysis features for data augmentation has also been explored .
Additionally, we can use more generic image manipulation detection algorithms against Deepfake videos. A recent study indicates that convolutional neural network (CNN)-generated images are easy to spot . The results show that such images share common systematic flaws, which can be grasped by dedicated ML models with the assist of data augmentation. Pixel-level image forgery detection can also be achieved. Researchers formulate this problem as a local anomaly detection problem and solve it with the aid of an elaborate score and a long short-term memory solution .
Defense against Deepfake videos. Adversarial examples against face detectors can be generated so that no valid faces can be detected, thus defending against Deepfake attacks . However, the attackers can still rely on manually extracted face regions to train Deepfake models. Therefore, such defense is essentially infeasible to provide enough protection. The lack of valid defenses against Deepfakes motivates us to investigate this subject.
2 GAN-based Deepfake model
The Deepfake model we considered in this paper is an open-source project based on GAN . This model is revised from the original Deepfake model (which is primarily based on autoencoder ) and includes many dedicated features and blocks. The backbone of the generator is borrowed from the CycleGAN . There are also self-attention blocks in both generator and discriminator. The detailed structure can be found in Appendix A.
Reconstruction loss is used to enhance the training of the generator to generate high-quality faces. Edge loss is used to improve the details of the edges by punishing mismatches between real and fake faces. Cyclic loss is evolved from the cycle consistency loss from CycleGAN , which helps to regularize the structured data. Perceptual loss extracts features from multiple layers of a holdout CNN model to determine the high-level differences in content or style .
Adversarial faces as a defense
The difficulty of the defense against Deepfake attacks largely depends on the defenders’ prior knowledge about the Deepfake model and data. Here, we consider different settings from both data and model’s perspectives. The white-box setting indicates that the defenders have full knowledge of both data (both target faces and source faces ) and the attackers’ Deepfake model (including the model structure, training schemes, and hyperparameters). The black-box setting indicates that the defenders have zero knowledge of source faces and the Deepfake model. Note that the defenders shall have the target faces at hand because those are the faces the defenders aim to protect. The gray-box setting is between the above two settings where the defenders have access to either source faces or the Deepfake model. In this paper, we will discuss and evaluate our defense methods under the white-box setting and then extend it to the gray- and black-box settings.
2 Principle and feasibility of adversarial faces as a defense
A straight-forward offensive method to defend against Deepfake attacks could be the poisoning attack that compromises the training of Deepfake models. Thus the trained Deepfake models can not generate meaningful results. However, the implementation of poisoning attacks can be very challenging even for simple feedforward NNs , let alone autoencoders and GANs . Recently, there is an interesting observation that the poisoned instances can be seen as adversarial examples . Based on this phenomenon, we further explore the use of adversarial faces as poisoned samples for defending against Deepfake attacks.
3 Transformation-aware adversarial face generation
A typical adversarial example can be obtained using the fast gradient sign method (FGSM) as
However, adversarial faces based on naïve FGSM cannot achieve satisfying defense performance. The reason is that such adversarial faces are not robust to image transformations (e.g., resizing, cropping, warping), which are commonly applied during the training of Deepfake models. The effect of the adversarial perturbations will be compromised after these image transformation operations.
Algorithm 1 shows the pseudocode of the proposed transformation-aware adversarial face generation algorithm using projected gradient descent (PGD) . Among various adversarial attacks such as FGSM, PGD, momentum iterative FGSM , CW attack , and universal adversarial attack , PGD with a multi-iteration scheme becomes the mainstream algorithm because of its simplicity and effectiveness. The transformation function is embedded inside the PGD loop so that this function varies slightly with iterations. Such randomness improves the robustness of the generated adversarial faces. The transformation function can also be customized based on needs, as long as being differentiable. Meanwhile, the algorithm can be easily incorporated with other adversarial attack methods with trivial modifications. The comparison of the defense performance of different adversarial attack methods can be found in Appendix B.
4 Ensemble method and more
The transformation-aware adversarial face generation is effective under the white-box setting. Furthermore, we extend it to the black-box setting, under which the distribution of the source faces is unknown. We could apply the defense based on source faces from the distribution of a random person, but this may not effectively approximate the unknown distribution. Instead, we propose to generate adversarial faces from pre-trained Deepfake models based on faces sampled from multiple distributions that are different from the actual source domain used by the attackers. Algorithm 2 shows the pseudocode of the ensemble version of the proposed method. These ensemble adversarial faces can better approximate the optimum adversarial faces generated based on the unknown domain , due to the fact that adversarial examples transfer between models . Therefore, such adversarial faces can be more robust under both gray- and black-box settings. In addition, we propose a “random” method that overlaps the perturbations of random directions to the target faces. We also clip these perturbations so that they are within the same perturbation bound (similar to PGD). This method is thus independent of both the model and data, and follows the same procedure under all settings.
Evaluation
We experiment on a subset of faces in Faceforensics++ (400 images for each face), as the training of Deepfake models is time-consuming (typically 30 GPU hours for one model). Specifically, we randomly select four faces as the target faces to protect, where we have two male faces (M1 and M2) and two female faces (F1 and F2). For the male faces, we swap them with four other male faces labeled as M3, M4, M5, and M6 in order to achieve the best swapping results to human eyes. Similarly, we swap the female faces with four other female faces labeled as F3, F4, F5, and F6. For each pair of faces (16 pairs in total), for example, M1 as the target vs. M3 as the source, we (from the attackers’ perspective) train 8 different models as follows:
Original is the original Deepfake model trained with real faces of M1 and M3. PGD-01 and PGD-005 are two white-box Deepfake models trained with transformation-aware adversarial faces of M1 (perturbation bound and , respectively) and real M3. These adversarial faces are generated from a pre-trained Deepfake model (with the same architecture) based on real M1 and M3. Ensemble is similar to PGD-01. The difference is that the adversarial M1 are generated from pre-trained Deepfake models (with the same architecture) based on real M1 and M7/F1/F4 which are different from the source faces to be swapped from (e.g., M3). Thus, applying Ensemble is under the gray-box setting. Random is a Deepfake model trained with randomly perturbed faces M1 with perturbation bound and real faces M3. Lite is a Deepfake model trained with real faces M1 and M3 but with a different architecture than Original. This model has fewer channels (e.g., half) in most layers and serves as a lite version for limited computing resource. Lite-Ens is the lite model trained with faces generated from Ensemble. In this case, both the data and model are unknown to the defender. This model is used to explore the generalization ability of our defense method under the black-box setting. Lite-Random is similar to Random but trained with the lite model.
We perform the experiments on two different resolutions of the Deepfake model, i.e., and , to evaluate the robustness of our defense. In this work, we focus on the raw swapped faces instead of the faces after post-processing because 1) post-processing can vary dramatically from project to project, and 2) current state-of-the-art detectors generally handle poorly on post-processed images/videos. We also find that training with full adversarial faces is better than a mix of adversarial and real faces, of which the comparison results are included in Appendix C. In this section, we present only the results on full adversarial faces, i.e., all target faces are adversarial.
2 Analyzing the degradations of synthesized faces
Figure 2 compares the visualization results of face degradations of two different target faces with 5 different models (see Appendix D for more results). The real and protected target faces, corresponding synthesized faces, and spectrums of Fast Fourier Transform (FFT) are presented in the figure. For both resolutions and target faces, significantly more artifacts can be observed in the synthesized faces generated by models with defense compared to the Original model without defense. This phenomenon is also confirmed by comparing the high-frequency regions of the FFT spectrums. The average intensities in these regions are significantly higher (lighter color) in models with defense, indicating more noise-like signals in corresponding synthesized faces. Moreover, the spectrums sometimes show clear patterns in the high-frequency regions, especially when the resolution is 128. Note that compared with PGD-01, PGD-005 usually incurs less degradation but also provides a lower perturbation bound, implying a trade-off between usability and security.
3 Evaluation using generic image manipulation detector
We also evaluate our defense using a state-of-the-art generic image forgery detection model – ManTra-Net . ManTra-Net generates a pixel-level detection mask reflecting the probability of manipulation in the original image. Figure 4 shows the ManTra-Net masks of different synthesized faces. More results can be found in Appendix E. For real faces cropped from video frames, ManTra-Net tends to generate noise-like patterns, indicating no specific focus. For synthesized faces generated by Original model ManTra-Net starts to focus on the center region of the face including the eyes and mouth. Interestingly, for synthesized faces generated by PGD-01 model, the detector is able to clearly segment out the region of eyes and mouth in the corresponding masks.
Conclusion
In this paper, we propose to defend against Deepfake attacks via transformation-aware adversarial faces. We show that training a Deepfake model with adversarially perturbed face images can lead to a significant degradation in the quality of synthesized faces. This degradation can be visually detectable and easily identified by various metrics. We also identify the adversarial and edge losses as the major indicators of such degradation. Extensive experiment results of multiple faces under white-box, gray-box, and black-box settings demonstrate the effectiveness and robustness of our defense method based on various metrics.
Broader Impact
Deepfake can have potentially drastic security consequences if applied inappropriately . It pushes the need for identity protection to the next level. The proposed methods can help defend against such security threats by degrading the performance of the GAN-based Deepfake models. Celebrities or even normal individuals can benefit from this defense. We also see opportunities for research applying our methods to beneficial purposes, such as investigating whether adversarial examples could protect audio data from Deepfake manipulation. However, there are also potential risks of applying the proposed methods. For example, if the Deepfake model is retrained over time with more unprotected facial images, or new types of Deepfake models are developed, the attackers may still generate high-quality fake faces. This could lead to a false sense of security and unfavorable impacts on the defenders. Additionally, applications that learn from facial images run the risk of performance decrease. On the other hand, many industries can benefit from Deepfake, such as film industry, entertainment and games, educational media, and so on . Such benign uses of Deepfake may be affected by adversarial faces.
References
Appendix A Architectures of the GAN-based Deepfake model
The generative network used in this work was revised from CycleGAN . Here are notations used in defining architectures. “c3s1-k” denotes a Convolution-InstanceNorm-ReLU layer with k filters and stride 1. “sa” is the self-attention layer in SAGAN . “up3-k” denotes an upscale block that consists of a convolutional layer with k filters and a PixelShuffle layer. “Rk” denotes a residual block that consists of two convolutional layers, both of which have k filters. “Dk” denotes a dense layer with the output dimension as k. Below are the generator and discriminator architectures.
Generator architecture: Encoder: c3s1-64, c3s2-128, c3s2-256, sa, c3s2-512, sa, c3s2-1024, D1024, D16384, up4-512. Decoder: up4-256, up4-128, sa, up4-64, R64, sa, c5s1-3.
Discriminator architecture: c3s2-64, c3s2-128, sa, c3s2-256, sa, c5s1-1.