ID-Reveal: Identity-aware DeepFake Video Detection

Davide Cozzolino, Andreas Rössler, Justus Thies, Matthias Nießner, Luisa Verdoliva

Introduction

Recent advancements in synthetic media generation allow us to automatically manipulate images and videos with a high level of realism. To counteract the misuse of these image synthesis and manipulation methods, the digital media forensics field got a lot of attention . For instance, during the past two years, there has been intense research on DeepFake detection, that has been strongly stimulated by the introduction of large datasets of videos with manipulated faces .

However, despite excellent detection performance, the major challenge is how to generalize to previously unseen methods. For instance, a detector trained on face swapping will drastically drop in performance when tested on a facial reenactment method. This unfortunately limits practicality as we see new types of forgeries appear almost on a daily basis. As a result, supervised detection, which requires extensive training data of a specific forgery method, cannot immediately detect a newly-seen forgery type.

This mismatch and generalization issue has been addressed in the literature using different strategies, ranging from applying domain adaptation or active learning to strongly increasing augmentation during training or by means of ensemble procedures . A different line of research is relying only on pristine videos at training time and detecting possible anomalies with respect to forged ones . This can help to increase the generalization ability with respect to new unknown manipulations but does not solve the problem of videos characterized by a different digital history. This is quite common whenever a video is spread over social networks and posted multiple times by different users. In fact, most of the platforms often reduce the quality and/or the video resolution.

Note also that current literature has mostly focused on face-swapping, a manipulation that replaces the facial identity of a subject with another one, however, a very effective modification is facial reenactment , where only the expression or the lips movements of a person are modified (Fig. 2). Recently, the MIT Center for Advanced Virtuality created a DeepFake video of president Richard Nixon https://moondisaster.org. The synthetic video shows Nixon giving a speech he never intended to deliver, by modifying only the lips movement and the speech of the old pristine video. The final result is impressive and shows the importance to develop forgery detection approaches that can generalize on different types of facial manipulations.

To better highlight this problem, we carried out an experiment considering the winning solution of the recent DeepFake Detection Challenge organized by Facebook on Kaggle platform. The performers had the possibility to train their models using a huge dataset of videos (around 100k fake videos and 20k pristine ones with hundreds of different identities). In Fig. 3, we show the results of our experiment. The model was first tested on a dataset of real and deepfake videos including similar face-swapping manipulations, then we considered unseen face-swapping manipulations and finally videos manipulated using facial reenactment. One can clearly observe the significant drop in performance in this last situation. Furthermore, the test on low quality compressed videos shows an additional loss and the final value for the accuracy is no more than a random guess.

It is also worth noting that current approaches are often used as black-box models and it is very difficult to predict the result because in a realistic scenario it is impossible to have a clue about the type of manipulation that occurred. The lack of reliability of current supervised deep learning methods pushed us to take a completely different perspective, avoiding to answer to a binary question (real or fake?) and instead focusing on wondering if the face under test preserves all the biometric traits of the involved subject.

Following this direction, our proposed method turns out to be able to generalize to different manipulation methods and also shows robustness w.r.t. low-quality data. It can reveal the identity of a subject by highlighting inconsistencies of facial features such as temporal consistent motion. The underlying CNN architecture comprises three main components: a facial feature extractor, a temporal network to detect biometric anomalies (temporal ID network) and a generative adversarial network that tries to predict person-specific motion based on the expressions of a different subject. The networks are trained only on real videos containing many different subjects . During test time, in addition to the test video, we assume to have a set of pristine videos of the target person. Based on these pristine examples, we compute a distance metric to the test video using the embedding of the temporal ID network (Fig. 1). Overall, our main contributions are the following:

We propose an example-based forgery detection approach that detects videos of facial manipulations based on the identity of the subject, especially the person-specific face motion.

An extensive evaluation that demonstrates the generalization to different types of manipulations even on low-quality videos, with a significant average improvement of more than 1515% w.r.t. state of the art.

Related Work

Digital media forensics, especially, in the context of DeepFakes, is a very active research field. The majority of the approaches rely on the availability of large-scale datasets of both pristine and fake videos for supervised learning. A few approaches detect manipulations as anomalies w.r.t. features learned only on pristine videos. Some of these approaches verify if the behavior of a person in a video is consistent with a given set of example videos of this person. Our approach ID-Reveal is such an example-based forgery detection approach. In the following, we discuss the most related detection approaches.

Afchar et al. presented one of the first approaches for DeepFake video detection based on supervised learning. It focuses on mesoscopic features to analyze the video frames by using a network with a low number of layers. Rössler et al. investigated the performance of several CNNs architectures for DeepFake video detection and showed that very deep networks are more effective for this task, especially, on low-quality videos. To train the networks, the authors also published a large-scale dataset. The best performing architecture XceptionNet was applied frame-by-frame and has been further improved by follow-up works. In an attention mechanism is included, that can also be used to localize the manipulated regions, while in Kumar et al. a triplet loss has been applied to improve performance on highly compressed videos.

Orthogonally, by exploiting artifacts that arise along the temporal direction it is possible to further boost performance. To this end, Guera et al. propose using a convolutional Long Short Term Memory (LSTM) network. Masi et al. propose to extract features by means of a two-branch network that are then fed into the LSTM: one branch takes the original information, while the other one works on the residual image. Differently in a 3D CNN structure is proposed together with an attention mechanism at different abstraction levels of the feature maps.

Most of these methods achieve very good performance when the training set comprises the same type of facial manipulations, but performance dramatically impairs on unseen tampering methods. Indeed, generalization represents the Achilles’ heel in media forensics. Augmentation can be of benefit to generalize to different manipulations as shown in . In particular, augmentation has been extensively used by the best performing approaches during the DeepFake detection challenge . Beyond the classic augmentation operations, some of them were particularly useful, e.g., by including cut-off based strategies on some specific parts of the face. In addition to augmentation, ensembling different CNNs have been also used to improve performance during this challenge . Another possible way to face generalization is to learn only on pristine videos and interpret a manipulation as an anomaly. This can improve the detection results on various types of face manipulations, even if the network never saw such forgeries in training. In the authors extract the camera fingerprint information gathered from multiple frames and use those for detection. Other approaches focus on specific operations used in current DeepFake techniques. For example, in the aim is to detect the blending operation that characterizes the face boundaries for most current synthetic face generation approaches.

A different perspective to improve generalization is presented in , where few-shot learning strategies are applied. Thus, these methods rely on the knowledge of a few labeled examples of a new approach and guide the training process such that new embeddings can be properly separated from previous seen manipulation methods and pristine samples in a short retraining process.

Features based on physiological signals

Other approaches look at specific artifacts of the generated videos that are related to physiological signals. In it is proposed a method that detects eye blinking, which is characterized by a specific frequency and duration in videos of real people. Similarly, one can also use inconsistencies on head pose or face warping artifacts as identifiers for tampered content. Recent works are also using heart beat and other biological signals to find inconsistencies both in spatial and along the temporal direction.

Identity-based features

The idea of identity-based approaches is to characterize each individual by extracting some specific biometric traits that can be hardly reproduced by a generator . The work by Agarwal et al. is the first approach that exploits the distinct patterns of facial and head movements of an individual to detect fake videos. In inconsistencies between the mouth shape dynamics and a spoken phoneme are exploited. Another related work is proposed in to detect face-swap manipulations. The technique uses both static biometric based on facial identity and temporal ones based on facial expressions and head movements. The method includes standard techniques from face recognition and a learned behavioral embedding using a CNN powered by a metric-learning objective function. In contrast, our proposed method extracts facial features based on a 3D morphable model and focuses on temporal behavior through an adversarial learning strategy. This helps to improve the detection of facial reenactment manipulations while still consistently be able to spot face swapping ones.

Proposed Method

ID-Reveal is an approach for DeepFake detection that uses prior biometric characteristics of a depicted identity, to detect facial manipulations in video content of the person. Any manipulated video content based on facial replacement results in a disconnect between visual identity as well as biometrical characteristics. While facial reenactment preserves the visual identity, biometrical characteristics such as the motion are still wrong. Using pristine video material of a target identity we can extract these biometrical features and compare them to the characteristics computed on a test video that is potentially manipulated. In order to be able to generalize to a variety of manipulation methods, we avoid training on a specific manipulation method, instead, we solely train on non-tampered videos. Additionally, this allows us to leverage a much larger training corpus in comparison to the facial manipulation datasets .

Our proposed method consists of three major components (see Fig. 4). Given a video as input, we extract a compact representation of each frame using a 3D Morphable model (3DMM) . These extracted features are input to the Temporal ID Network which computes an embedded vector. During test time, a metric in the embedding space is used to compare the test video to the previously recorded biometrics of a specific person. However, in order to ensure that the Temporal ID Network is also based on behavioral instead of only visual information, we utilize a second network, called 3DMM Generative Network, which is trained jointly in an adversarial fashion (using the Temporal ID Network as discriminator). In the following, we will detail the specific components and training procedure.

Temporal ID Network

The Temporal ID Network NT\mathcal{N}_{T} processes the temporal sequence of 3DMM features through convolution layers that work along the temporal direction in order to extract the embedded vector yc,i(t)=NT[xc,i(t)]y_{c,i}(t)=\mathcal{N}_{T}\left[x_{c,i}(t)\right]. To evaluate the distance between embedded vectors, we adopt the squared Euclidean distance, computing the following similarity:

As a metric learning loss, similar to the Distance-Based Logistic Loss , we adopt a log-loss on a suitably defined probability . Specifically, for each embedded vector yc,i(t)y_{c,i}(t), we build the probability through softmax processing as:

Thus, we are considering all the similarities with respect to the pivot vector yc,i(t)y_{c,i}(t) in our probability definition pc,i(t)p_{c,i}(t). Note that to obtain a high probability value it is only necessary that at least one similarity with the same individual is much larger than similarities with other individuals. Indeed, the loss proposed here is a less restrictive loss compared to the current literature, where the aim is to achieve a high similarity for all the coherent pairs . The adopted metric learning loss is then obtained from the probabilities through the log-loss function:

In order to tune hyper-parameters during training, we also measure the accuracy of correctly identifying a subject. It is computed by counting the number of times where at least one similarity with the same individual is larger than all the similarities with other individuals. The Temporal ID Network is first trained alone using the previously described loss, and afterward it is fine-tuned together with the 3DMM Generative Network, which we describe in the following paragraph.

DMM Generative Network

The 3DMM Generative Network NG\mathcal{N}_{G} is trained to generate 3DMM features similar to the features that we may extract from a manipulated video. Specifically, the generative network has the goal to output features that are coherent to the identity of an individual, but with the expressions of another subject. The generative network NG\mathcal{N}_{G} works frame-by-frame and generates a 3DMM feature vector by combining two input feature vectors. Let xcx_{c} and xkx_{k} are the 3DMM feature vectors respectively of the individuals cc and kk, then, NG[xk,xc]\mathcal{N}_{G}\left[x_{k},x_{c}\right] is the generated feature vector with appearance of the individual cc and expressions of individual kk. During training, we use batches that contain N×MN\times M videos of NN different individuals each with MM videos. In our experiments, we chose M=N=8M=N=8. To train the generative network NG\mathcal{N}_{G}, we apply it to pairs of videos of these NN identities. Specifically, for each identity cc, we compute an averaged 3DMM feature vector x‾c\overline{x}_{c}. Based on this averaged input feature x‾c\overline{x}_{c} and a frame feature xi(t)x_{i}(t) of a video of person ii (which serves as expression conditioning), we generate synthetic 3DMM features using the generator NG\mathcal{N}_{G}:

The 3DMM Generative Network is trained based on the following loss:

Where Lcycle\mathcal{L}_{cycle} is a cycle consistency used in order to preserve the expression. Specifically, the 3DMM Generative Network is applied twice, firstly to transform a 3DMM feature vector of the individual ii to identity cc and then to transform the generated 3DMM feature vector to identity ii again, we should obtain the original 3DMM feature vector. The Lcycle\mathcal{L}_{cycle} is defined as:

The adversarial loss Ladv\mathcal{L}_{adv} is based on the Temporal ID Network, i.e., it tries to fool the Temporal ID Network by generating features that are coherent for a specific identity. Since the generator works frame-by-frame, it can deceive the Temporal ID Network by only altering the appearance of the individual and not the temporal patterns. The adversarial loss Ladv\mathcal{L}_{adv} is computed as:

where the probabilities pc,i∗(t)p^{*}_{c,i}(t) are computed using the equation 2, but considering the similarities evaluated between generated features and real ones:

Indeed, the generator aims to increase the similarity between the generated features for a given individual and the real features of that individual. During training, the Temporal ID Network is trained to hinder the generator, through a loss obtained as:

where the loss Linv\mathcal{L}_{inv}, contrary to Labv\mathcal{L}_{abv}, is used to minimized the probabilities pc,i∗(t)p^{*}_{c,i}(t). Therefore, it is defined as:

Overall, the final objective of the adversarial game is to increase the ability of the Temporal ID Network to distinguish real identities from fake ones.

Identification

Given a test sequence depicting a single identity as well as a reference set of pristine sequences of the same person, we apply the following procedure: we first embed both the test as well as our reference videos using the Temporal ID Network pipeline. We then compute the minimum pairwise Euclidean distance of each reference video and our test sequence. Finally, we compare this distance to a fixed threshold τid\tau_{id} to decide whether the behavioral properties of our testing sequence coincide with its identity, thus, evaluating the authenticity of our test video. The source code and the trained network of our proposal are publicly available https://github.com/grip-unina/id-reveal.

Results

To analyze the performance of our proposed method, we conducted a series of experiments. Specifically, we discuss our design choices w.r.t. our employed loss functions and the adversarial training strategy based on an ablation study applied on a set of different manipulation types and different video qualities. In comparison to state-of-the-art DeepFake video detection methods, we show that our approach surpasses these in terms of generalizability and robustness.

Our approach is trained using the VoxCeleb2 development dataset consisting of multiple video clips of several identities. Specifically, we use 51205120 subjects for the training-set and 512512 subjects for the validation-set. During training, each batch contains 6464 sequences of 9696 frames. The 6464 sequences are formed by M=8M=8 sequences for each individual, with a total of N=8N=8 different individuals extracted at random from the training-set. Training is performed using the ADAM optimizer , with a learning rate of 10−410^{-4} and 10−510^{-5} for the Temporal ID Network and the 3DMM Generative Network, respectively. The parameters λcycle\lambda_{cycle}, λinv\lambda_{inv} and τ\tau for our loss formulation are set to 1.01.0, 0.0010.001 and 0.080.08 respectively. We first train the Temporal ID Network for 300 epochs (with an epoch size of 25002500 iterations) and choose the best performing model based on the validation accuracy. Using this trained network, we enable our 3DMM Generative Network and continue training for a fixed 100 epochs. For details on our architectures, we refer to the supplemental document. For all experiments, we use a fixed threshold of τid=1.1\tau_{id}=\sqrt{1.1} to determine whether the behavioral properties of a test video coincide with those of our reference videos. This threshold is set experimentally based on a one-time evaluation on 44 real and 44 fake videos from the original DFD using the averaged squared euclidean distance of real and manipulated videos.

2 Ablation Study

In this section, we show the efficacy of the proposed loss and the adversarial training strategy. For performance evaluation of our approach, we need to know the involved identity (the source identity for face-swapping manipulations and the target identity for facial reenactment ones). Based on this knowledge, we can set up the pristine reference videos used to compute the final distance metric. To this end, we chose a controlled dataset that includes several videos of the same identity, i.e., the recently created dataset of the Google AI lab, called DeepFake Dataset (DFD) . The videos contain 2828 paid actors in 1616 different contexts, furthermore, for each subject there are pristine videos provided (varying from 99 to 1616). In total, there are 363363 real and 30683068 DeepFakes videos. Since the dataset only contains face-swapping manipulations, we generated 320320 additional videos that include 160160 Face2Face and 160160 Neural Textures videos. Some examples are shown in Fig. 5.

Performance is evaluated at video level using a leave-one-out strategy for the reference-dataset. In detail, for each video under test, the reference dataset only contains pristine videos with a different context from the one under test. The evaluation is done both on high quality (HQ) compressed videos (constant rate quantization parameter equal to 23) using H.264 and low quality (LQ) compressed videos (quantization parameter equal to 40). This scenario helps us to consider a realistic situation, where videos are uploaded to the web, but also to simulate an attacker who further compresses the video to hide manipulation traces.

We compare the proposed loss to the triplet loss and the multi-similarity loss (MSL) . For these two losses, we adopt the cosine distance instead of the Euclidean one as proposed by the authors. Moreover, hyper-parameters are chosen to maximize the accuracy to correctly identify a subject in the validation set. Results for facial reenactment (FR) and face swapping (FS) in terms of Area Under Curve (AUC) and accuracy both for HQ and LQ videos are shown in Tab. 1. One can observe that our proposed loss gives a consistent improvement over the multi-similarity loss (5.55.5% on average) and the triplet loss (2.82.8% on average) in terms of AUC. In addition, coupled with the adversarial training strategy performance, it gets better for the most challenging scenario of FR videos with a further improvement of around 33% for AUC and of 66% (on average) in terms of accuracy.

3 Comparisons to State of the Art

We compare our approach to several state-of-the-art DeepFake video detection methods. All the techniques are compared using the accuracy at video level. Hence, if a method works frame-by-frame, we average the probabilities obtained from 32 frames uniformly extracted from the video, as it is also done in .

The methods used for our comparison are frame-based methods: MesoNet , Xception , FFD (Facial Forgery Detection) , Efficient-B7 ; ensemble methods: ISPL (Image and Sound Processing Lab) , Seferbekov ; temporal-based methods: Eff.B1 + LSTM, ResNet + LSTM and an Identity-based method: A&B (Appearance and Behavior) . A detailed description of these approaches can be found in the supplemental document. In order to ensure a fair comparison, all supervised approaches (frame-based, ensemble and temporal-based methods) are trained on the same dataset of real and fake videos, while the identity-based ones (A&B and our proposal) are trained instead on VoxCeleb2 .

Generalization and robustness analysis

To analyze the ability to generalize to different manipulation methods, training and test come from different datasets. Note that we will focus especially on generalizing from face swapping to facial reenactment.

In a first experiment we test all the methods on the DFD Google dataset that contains both face swapping and facial reenactment manipulations, as described in Section 4.2. In this case all supervised approaches are trained on DFDC with around 100k fake and 20k real videos. This is the largest DeepFake dataset publicly available and includes five different types of manipulations https://www.kaggle.com/c/deepfake-detection-challenge. Experiments on (HQ) videos, with a compression factor of 23, and on low-quality (LQ) videos, where the factor is 40 are presented in terms of accuracy and AUC in Tab. 2. Most methods suffer from a huge performance drop when going from face-swapping to facial-reenactment, with an accuracy that often borders 50%, equivalent to coin tossing. The likely reason is that the DFDC training set includes mostly face-swapping videos, and methods with insufficient generalization ability are unable to deal with different manipulations. This does not hold for ID-Reveal and A&B, which are trained only on real data and, hence, have an almost identical performance with both types of forgeries. For facial reenactment videos, this represents a huge improvement with respect to all competitors. In this situation it is possible to observe a sharp performance degradation of most methods in the presence of strong compression (LQ videos). This is especially apparent with face-swapping, where some methods are very reliable on HQ videos but become almost useless on LQ videos. On the contrary, ID-Reveal suffers only a very small loss of accuracy on LQ videos, and outperforms all competitors, including A&B, by a large margin.

In another experiment, we use FaceForensics++ (HQ) for training the supervised methods, while the identity based methods are always trained on the VoxCeleb2 dataset . For testing, we use the preview DFDC Facebook dataset and CelebDF . The preview DFDC dataset is composed only of face-swapping manipulations of 68 individuals. For each subject there are 3 to 39 pristine videos with 3 videos for each context. We consider 44 individuals which have at least 12 videos (4 contexts); obtaining a total of 920 real videos and 2925 fake videos. CelebDF contains 890 real videos and 5639 face-swapping manipulated videos. The videos are related to 59 individuals except for 300 real videos that do not have any information about the individual, hence, they cannot be included in our analysis. Results in terms of accuracy and AUC at video-level are shown in Tab. 3. One can observe that also in this scenario our method achieves very good results for all the datasets, with an average improvement with respect to the best supervised approach of about 16% on LQ videos. Even the improvement with respect to the identity based approach A&B is significant, around 14% on HQ videos and 13% on LQ ones. Again the performance of supervised approaches worsens in unseen conditions of low-quality videos, while our method preserves its good performance.

To gain better insights on both generalization and robustness, we want to highlight the very different behavior of supervised methods when we change the fake videos in training. Specifically, for HQ videos if the manipulation (in this case neural textures and face2face) is included in training and test, then performance are very high for all the methods, but they suddenly decrease if we exclude those manipulations from the training, see Fig. 6. The situation is even worse for LQ videos. Identity-based methods do not modify their performance since they do not depend at all on which manipulation is included in training.

Conclusion

We have introduced ID-Reveal, an identity-aware detection approach leveraging a set of reference videos of a target person and trained in an adversarial fashion. A key aspect of our method is the usage of a low-dimensional 3DMM representation to analyze the motion of a person. While this compressed representation of faces contains less information than the original 2D images, the gained type of robustness is a very important feature that makes our method generalize across different forgery methods. Specifically, the 3DMM representation is not affected by different environments or lighting situations, and is robust to disruptive forms of post-processing, e.g., compression. We conducted a comprehensive analysis of our method and in comparison to state of the art, we are able to improve detection qualities by a significant margin, especially, on low-quality content. At the same time, our method improves generalization capabilities by adopting a training strategy that solely focuses on non-manipulated content.

Acknowledgment

We gratefully acknowledge the support of this research by a TUM-IAS Hans Fischer Senior Fellowship, a TUM-IAS Rudolf Mößbauer Fellowship and a Google Faculty Research Award. In addition, this material is based on research sponsored by the Defense Advanced Research Projects Agency (DARPA) and the Air Force Research Laboratory (AFRL) under agreement number FA8750-20-2-1004. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of DARPA and AFRL or the U.S. Government. This work is also supported by the PREMIER project, funded by the Italian Ministry of Education, University, and Research within the PRIN 2017 program.

Appendix

In this appendix, we report the details of our architectures used for the Temporal ID Network and the 3DMM Generative Network (Sec. A). Moreover, we briefly describe the state of the art DeepFake methods we compare to, (see Sec. B). In Sec. C and Sec. D, we present additional results to prove the generalization capability of our method. In Sec. E, we include scatter plots that show the separability of videos of different subjects in the embedding space. Finally, we analyze a real case on the web (see Sec. F).

Appendix A Architectures

We leverage a convolution neural network architecture that works along the temporal direction and is composed by eleven layers (see Fig.7 (a)). We use Group Normalization and LeakyReLU non-linearity for all layers except the last one. Moreover, we adopt à-trous convolutions (also called dilated convolution), instead of classic convolutions in order to increase the receptive fields without increasing the trainable parameters. The first layer increases the number of channels from 62 to 512, while the successive ones are inspired by the ResNet architecture and include a residual block as shown in Fig.7 (c). The parameters KK and DD of the residual blocks are the dimension of the filter and the dilatation factor of the à-trous convolution, respectively. The last layer reduces the channels from 512 to 128. The receptive field of the whole network is equal to 51 frames which is around 2 seconds.

DMM Generative Network

As described in the main paper, the 3DMM Generative Network is fed by two 3DMM feature vectors. The two feature vectors are concatenated which results in a single input vector of 124 channels. The network is formed by five layers: a layer to increase the channels from 124 to 512, three residual blocks, and a last layer to decrease the channels from 512 to 62. The output is summed to the input 3DMM feature vector to obtain the generated 3DMM feature vector (see Fig.7 (b)). All the convolutions have a dimension of filter equal to one in order to work frame-by-frame.

Appendix B Comparison methods

In the main paper, we compare our approach with several state of the art DeepFake detection methods, that are described in following:

MesoNet : is one of the first CNN methods proposed for DeepFake detection which uses dilated convolutions with inception modules.

Xception : is a relatively deep neural network that is achieving a very good performance compared to other CNNs for video DeepFake detection .

FFD (Facial Forgery Detection) : is a variant of Xception, including an attention-based layer, in order to focus on high-frequency details.

Efficient-B7 : has been proposed by Tan et al. and is pre-trained on ImageNet using the strategy described in , where the network is trained with injected noise (such as dropout, stochastic depth, and data augmentation) on both labeled and unlabeled images.

Ensemble methods

ISPL (Image and Sound Processing Lab) : employs an ensemble of four variants of Efficienet-B4. The networks are trained using different strategies, such as self-attention mechanism and triplet siamese strategy. Data augmentation is performed by applying several operations, like downscaling, noise addition and JPEG compression.

Seferbekov : is the algorithm proposed by the winner of the Kaggle competition (Deepfake Detection Challenge) organized by Facebook . It uses an ensemble of seven Efficientnet-B7 that work frame-by-frame. The networks are pre-trained using the strategy described in . The training leverages data augmentation, that, beyond some standard operations, includes a cut-out that drops specific parts of the face.

Temporal-based methods

ResNet + LSTM: is a method based on Long Short Term Memory (LSTM) . In detail, a ResNet50 is used to extract frame-level features from 20 frames uniformly extracted from the video. These features are provided to a LSTM that classifies the whole video.

Eff.B1 + LSTM: This is a variant of the approach described above, where the ResNet architecture is replaced by EfficientNet-B1.

Identity-based methods

A&B (Appearance and Behavior) : is an identity-based approach that includes a face recognition network and a network that is based on head movements. The behavior recognition system encodes the information about the identity through a network that works on a sequence of attributes related to the movement .

Note that all the techniques are compared at video level. Hence, if a method works frame-by-frame, we average the probabilities obtained from 32 frames uniformly extracted from the video. Furthermore, to validate this choice, we compare averaging with the maximum strategy. Results are reported in Tab. 4 using the same experimental setting of Tab. 2 of the main paper. The results prove the advantage to use the averaging operation with respect to the maximum value: the increase in terms of AUC is around 0.04, while the accuracy increases (on average) of about 3%.

Appendix C Additional results

To show the ability of our method to be agnostic to the type of manipulation, we test our proposal on additional datasets, that are not included in the main paper. In Tab. 5 we report the analysis on the dataset FaceForensics++ (FF++) . Results are split for facial reenactment (FR) and face swapping (FS) manipulations. It is important to underline that this dataset does not provide information about multiple videos of the same subject, therefore, for identity-based approaches, the first 6 seconds of each pristine video are used as reference dataset, while the last 6 seconds are used to evaluate the performance (we only consider videos of at least 14 seconds duration, thus, obtaining 360 videos for each manipulation method). For the FF++ dataset, our method obtains always better performance in both the cases of high-quality videos and low-quality ones.

As a further analysis, we test our method on a recent method of face reenactment, called FOMM (First-Order Motion Model) . Using the official code of FOMM, we created 160 fake videos using the pristine videos of DFD, some examples are in Fig. 8. Our approach on these videos achieves an accuracy of 85.6%, and an AUC of 0.94 which further underlines the generalization of our method with respect to a new type of manipulation.

Appendix D Robustness to different contexts

We made additional experiments to understand that for our method it is not necessary that the reference videos are similar to the manipulated ones in terms of environment, lighting, or distance from the subject. To this end, we show results in Fig. 9 obtained for the DFD FR and DFD FS datasets, where information about the video context (kitchen, podium, outside, talking, meeting, etc.) is available. While the reference videos and the under-test videos differ, our method shows robust performance. Results seem only affected by the variety of poses and expressions present in the reference videos (the last reference video in the table contains the most variety in motion, thus yielding better results).

Appendix E Visualization of the embedded vectors

In this section, we include scatter plots that show the 2D orthogonal projection of the extracted temporal patterns. In particular, in Fig. 10 we show the scatter plots of embedded vectors extracted from 4 seconds long video snippets relative to two actors for the DFD dataset by using Linear Discriminant Analysis (LDA) and selecting the 2-D orthogonal projection that maximize the separations between the real videos of two actors and between real videos and fake ones. We can observe that in the embedding space the real videos relative to different actors are perfectly separated. Moreover, also the manipulated videos relative to an actor are well separated from the real videos of the same actor.

Appendix F A real case on the web

We applied ID-Reveal to videos of Nicolas Cage downloaded from YouTube. We tested on three real videos, four DeepFakes videos, one imitator (a comic interpreting Nicolas Cage) and a DeepFake applied on the imitator. We evaluate the distributions of distance metrics that are computed as the minimum pairwise squared Euclidean distance in the embedding space of 4 seconds long video snippets extracted from the pristine reference video and the video under test. In Fig. 11, we report these distributions using a violin plot.

We can observe that the lowest distances are relative to real videos (green). For the DeepFakes (red) all distances are higher and, thus, can be detected as fakes. An interesting case is the video related to the imitator (purple), that presents a much lower distance since he is imitating Nicolas Cage. A DeepFake driven by the imitator strongly reduces the distance (pink), but is still detected by our method.

References