The Creation and Detection of Deepfakes: A Survey

Yisroel Mirsky, Wenke Lee

Introduction

A deepfake is content, generated by an artificial intelligence, that is authentic in the eyes of a human being. The word deepfake is a combination of the words ‘deep learning’ and ‘fake’ and primarily relates to content generated by an artificial neural network, a branch of machine learning.

The most common form of deepfakes involve the generation and manipulation of human imagery. This technology has creative and productive applications. For example, realistic video dubbing of foreign films,https://variety.com/2019/biz/news/ ai-dubbing-david-beckham-multilingual-1203309213/ education though the reanimation of historical figures (Lee, 2019), and virtually trying on clothes while shopping.https://www.forbes.com/sites/forbestechcouncil/2019/05/21/gans-and-deepfakes-could-revolutionize-the-fashion-industry/ There are also numerous online communities devoted to creating deepfake memes for entertainment,https://www.reddit.com/r/SFWdeepfakes/ such as music videos portraying the face of actor Nicolas Cage.

However, despite the positive applications of deepfakes, the technology is infamous for its unethical and malicious aspects. At the end of 2017, a Reddit user by the name of ‘deepfakes’ was using deep learning to swap faces of celebrities into pornographic videos, and was posting them onlinehttps://www.vice.com/en_us/article/gydydm/gal-gadot-fake-ai-porn. The discovery caused a media frenzy and a large number of new deepfake videos began to emerge thereafter. In 2018, BuzzFeed released a deepfake video of former president Barak Obama giving a talk on the subject. The video was made using the Reddit user’s software (FakeApp), and raised concerns over identity theft, impersonation, and the spread of misinformation on social media. Fig. presents an information trust chart for deepfakes, inspired by (Facebook, 2018).

Following these events, the subject of deepfakes gained traction in the academic community, and the technology has been rapidly advancing over the last few years. Since 2017, the number of papers published on the subject rose from 3 to over 250 (2018-20).

To understand where the threats are moving and how to mitigate them, we need a clear view of the technology’s, challenges, limitations, capabilities, and trajectory. Unfortunately, to the best of our knowledge, there are no other works which present the techniques, advancements, and challenges, in a technical and encompassing way. Therefore, the goals of this paper are (1) to provide the reader with an understanding of how modern deepfakes are created and detected, (2) to inform the reader of the recent advances, trends, and challenges in deepfake research, (3) to serve as a guide to the design of deepfake architectures, and (4) to identify the current status of the attacker-defender game, the attacker’s next move, and future work that may help give the defender a leading edge.

We achieve these goals through an overview of human visual deepfakes (Section 2), followed by a technical background which identifies technology’s basic building blocks and challenges (Section 3). We then provide a chronological and systematic review for each category of deepfake, and provide the networks’ schematics to give the reader a deeper understanding of the various approaches (Sections 4 and 5). Finally, after reviewing the countermeasures (Section 6), we discuss their weaknesses, note the current limitations of deepfakes, suggest alternative research, consider the adversary’s next steps, and raise awareness to the spread of deepfakes to other domains (Section 7).

Scope. In this survey we will focus on deepfakes pertaining to the human face and body. We will not be discussing the synthesis of new faces or the editing of facial features because they do not have a clear attack goal associated with them. In Section 7.3 we will discuss deepfakes with a much broader scope, note the future trends, and exemplify how deepfakes have spread to other domains and media such as forensics, finance, and healthcare.

We note to the reader that deepfakes should not be confused with adversarial machine learning, which is the subject of fooling machine learning algorithms with maliciously crafted inputs (Fig. 2). The difference being that for deepfakes, the objective of the generated content is to fool a human and not a machine.

Overview & Attack Models

“Believable media generated by a deep neural network”

In the context of human visuals, we identify four categories: reenactment, replacement, editing, and synthesis. Fig. 3 illustrates some examples facial deepfakes in each of these categories and their sub-types. Throughout this paper we denote ss and tt as the source and the target identities. We also denote xsx_{s} and xtx_{t} as images of these identities and xgx_{g} as the deepfake generated from ss and tt.

A reenactment deepfake is where xsx_{s} is used to drive the expression, mouth, gaze, pose, or body of xtx_{t}:

reenactment is where xsx_{s} drives the expression of xtx_{t}. It is the most common form of reenactment since these technologies often drive target’s mouth and pose as well, providing a wide range of flexibility. Benign uses are found in the movie and video game industry where the performances of actors are tweaked in post, and in educational media where historical figures are reenacted.

reenactment, also known as ‘dubbing’, is where the mouth of xtx_{t} is driven by that of xsx_{s}, or an audio input asa_{s} containing speech. Benign uses of the technology includes realistic voice dubbing into another language and editing.

reenactment is where direction of xtx_{t}’s eyes, and the position of the eyelids, are driven by those of xsx_{s}. This is used to improve photographs or to automatically maintain eye contact during video interviews (et al., 2017).

reenactment is where the head position of xtx_{t} is driven by xsx_{s}. This technology has primarily been used for face frontalization of individuals in security footage, and as a means for improving facial recognition software (Tran et al., 2018).

reenactment, a.k.a. pose transfer and human pose synthesis, is similar to the facial reenactments listed above except that’s its the pose of xtx_{t}’s body being driven.

The Attack Model. Reenactment deep fakes give attackers the ability to impersonate an identity, controlling what he or she says or does. This enables an attacker to perform acts of defamation, cause discredability, spread misinformation, and tamper with evidence. For example, an attacker can impersonate tt to gain trust the of a colleague, friend, or family member as a means to gain access to money, network infrastructure, or some other asset. An attacker can also generate embarrassing content of tt for blackmailing purposes or generate content to affect the public’s opinion of an individual or political leader. The technology can also be used to tamper surveillance footage or some other archival imagery in an attempt to plant false evidence in a trial. Finally, the attack can either take place online (e.g., impersonating someone in a real-time conversation) or offline (e.g., fake media spread on the Internet).

2. Replacement

A replacement deepfake is where the content of xtx_{t} is replaced with that of xsx_{s}, preserving the identity of ss.

is where the content of xtx_{t} is replaced with that of xsx_{s}. A common type of transfer is facial transfer, used in the fashion industry to visualize an individual in different outfits.

is where the content transferred to xtx_{t} from xsx_{s} is driven by xtx_{t}. The most popular type of swap replacement is ‘face swap’, often used to generate memes or satirical content by swapping the identity of an actor with that of a famous individual. Another benign use for face swapping includes the anonymization of one’s identity in public content in-place of blurring or pixelation.

The Attack Model. Replacement deepfakes are well-known for their harmful applications. For example, revenge porn is where an attacker swaps a victim’s face onto the body of a porn actress to humiliate, defame, and blackmail the victim. Face replacement can also be used as a short-cut to fully reenacting tt by transferring tt’s face onto the body of a look-alike. This approach has been used as a tool for disseminating political opinions in the past (Schwartz, 2018).

3. Editing & Synthesis

An enchantment deepfake is where the attributes of xtx_{t} are added, altered, or removed. Some examples include the changing a target’s clothes, facial hair, age, weight, beauty, and ethnicity. Apps such as FaceApp enable users to alter their appearance for entertainment and easy editing of multimedia. The same process can be used by and attacker to build a false persona for misleading others. For example, a sick leader can be made to look healthy (Hao, 2019), and child or sex predators can change their age and gender to build dynamic profiles online. A known unethical use of editing deepfakes is the removal of a victim’s clothes for humiliation or entertainment (Samuel, 2019).

Synthesis is where the deepfake xgx_{g} is created with no target as a basis. Human face and body synthesis techniques such as (Karras et al., 2019) (used in Fig. 3) can create royalty free stock footage or generate characters for movies and games. However, similar to editing deepfakes, it can also be used to create fake personas online.

Although human image editing and synthesis are active research topics, reenactment and replacement deepfakes are the greatest concern because they give an attacker control over one’s identity(Hall, 2018; Chesney and Citron, 2018; Antinori, 2019). Therefore, in this survey we will be focusing on reenactment and replacement deepfakes.

Technical background

Although there are a wide variety of neural networks, most deepfakes are created using variations or combinations of generative networks and encoder decoder networks. In this section we provide a brief introduction to these networks, how they are trained, and the notations which we will be using throughout the paper.

Neural networks are non-linear models for predicting or generating content based on an input. They are made up of layers of neurons, where each layer is connected sequentially via synapses. The synapses have associated weights which collectively define the concepts learned by the model. To execute a network on an nn-dimensional input xx, a process known as forward-propagation is performed where xx propagated through each layer and an activation function is used to summarize a neuron’s output (e.g., the Sigmoid or ReLU function).

Concretely, let l(i)l^{(i)} denote the ii-th layer in the network MM, and let ∥l(i)∥\|l^{(i)}\| denote the number of neurons in l(i)l^{(i)}. Finally, let the total number of layers in MM be denoted as LL. The weights which connect l(i)l^{(i)} to l(i+1)l^{(i+1)} are denoted as the ∥l(i)∥\|l^{(i)}\|-by-∥l(i+1)∥\|l^{(i+1)}\| matrix W(i)W^{(i)} and ∥l(i+1)∥\|l^{(i+1)}\| dimensional bias vector b⃗(i)\vec{b}^{(i)}. Finally, we denote the collection of all parameters θ\theta as the tuple θ≡(W,b)\theta\equiv(W,b), where WW and bb are the weights of each layer respectively. Let a(i+1)a^{(i+1)} denote the output (activation) of layer l(i)l^{(i)} obtained by computing f(W(i)⋅a⃗(i)+b(i)⃗)f\left(W^{(i)}\cdot\vec{a}^{(i)}+\vec{b^{(i)}}\right) where ff is often the Sigmoid or ReLU function. To execute a network on an nn-dimensional input xx, a process known as forward-propagation is performed where xx is used to activate l(1)l^{(1)} which activates l(2)l^{(2)} and so on until the activation of l(L)l^{(L)} produces the mm-dimensional output yy.

To summarize this process, we consider MM a black box and denote its execution as M(x)=yM(x)=y. To train MM in a supervised setting, a dataset of paired samples with the form (xi,yi)(x_{i},y_{i}) is obtained and an objective loss function L\mathcal{L} is defined. The loss function is used to generate a signal at the output of MM which is back-propagated through MM to find the errors of each weight. An optimization algorithm, such as gradient descent (GD), is then used to update the weights for a number of epochs. The function L\mathcal{L} is often a measure of error between the input xx and predicted output y′y^{\prime}. As a result the network learns the function M(xi)≈yiM(x_{i})\approx y_{i} and can be used to make predictions on unseen data.

Some deepfake networks use a technique called one-shot or few-shot learning which enables a pre-trained network to adapt to a new dataset X′X^{\prime} similar to XX on which it was trained. Two common approaches for this are to (1) pass information on x′∈X′x^{\prime}\in X^{\prime} to the inner layers of MM during the feed-forward process, and (2) perform a few additional training iterations on a few samples from X′X^{\prime}.

2. Loss Functions

where y′[c]y^{\prime}[c] is the predicted probability of xix_{i} belonging to the cc-th class.

Other popular loss functions used in deepfake networks include the L1 and L2 norms L1=∣x−xg∣1\mathcal{L}_{1}=|x-x_{g}|^{1} and L2=∣x−xg∣2\mathcal{L}_{2}=|x-x_{g}|^{2}. However, L1 and L2 require paired images (e.g., of ss and tt with same expression) and perform poorly when there are large offsets between the images such as different poses or facial features. This often occurs in reenactment when xtx_{t} has a different pose than xsx_{s} which is reflected in xgx_{g}, and ultimately we’d like xgx_{g} to match the appearance of xtx_{t}.

One approach to compare two unaligned images is to pass them through another network (a perceptual model) and measure the difference between the layer’s activations (feature maps). This loss is called the perceptual loss (Lperc\mathcal{L}_{perc}) and is described in (Johnson et al., 2016) for image generation tasks. In the creation of deepfakes, Lperc\mathcal{L}_{perc} is often computed using a face recognition network such as VGGFace. The intuition behind Lperc\mathcal{L}_{perc} is that the feature maps (inner layer activations) of the perceptual model act as a normalized representation of xx in the context of how the model was trained. Therefore, by measuring the distance between the feature maps of two different images, we are essentially measuring their semantic difference (e.g., how similar the noses are to each other and other finer details.) Similar to Lperc\mathcal{L}_{perc}, there is a feature matching loss (LFM\mathcal{L}_{FM}) (Salimans et al., 2016) which uses the last output of a network. The idea behind LFM\mathcal{L}_{FM} is to consider the high level semantics captured by the last layer of the perceptual model (e.g., the general shape and textures of the head).

Another common loss is a type of content loss (LC\mathcal{L}_{C}) (Gatys et al., 2015) which is used to help the generator create realistic features, based on the perspective of a perceptual model. In LC\mathcal{L}_{C}, only xgx_{g} is passed through the perceptual model and the difference between the network’s feature maps are measured.

3. Generative Neural Networks (for deepfakes)

Deep fakes are often created using combinations or variations of six different networks, five of which are illustrated in Fig. 4.

An ED consists of at least two networks, an encoder EnEn and decoder DeDe. The ED has narrower layers towards its center so that when it’s trained as De(En(x))=xgDe(En(x))=x_{g}, the network is forced to summarize the observed concepts. The summary of xx, given its distribution XX, is En(x)=eEn(x)=e, often referred to as an encoding or embedding and E=En(X)E=En(X) is referred to as the ‘latent space’. Deepfake technologies often use multiple encoders or decoders and manipulate the encodings to influence the output xgx_{g}. If an encoder and decoder are symmetrical, and the network is trained with the objective De(En(x))=xDe(En(x))=x, then the network is called an autoencoder and the output is the reconstruction of xx denoted x^\hat{x}. Another special kind of ED is the variational autorencoder (VAE) where the encoder learns the posterior distribution of the decoder given XX. VAEs are better at generating content than autoencoders because the concepts in the latent space are disentangled, and thus encodings respond better to interpolation and modification.

In contrast to a fully connected (dense) network, a CNN learns pattern hierarchies in the data and is therefore much more efficient at handling imagery. A convolutional layer in a CNN learns filters which are shifted over the input forming an abstract feature map as the output. Pooling layers are used to reduce the dimensionality as the network gets deeper and up-sampling layers are used to increase it. With convolutional, pooling, and upsampling layers, it is possible to build an ED CNNs for imagery.

The GAN was first proposed in 2014 by Goodfellow et al. in (Goodfellow et al., 2014). A GANs consist of two neural networks which work against each other: the generator GG and the discriminator DD. GG creates fake samples xgx_{g} with the aim of fooling DD, and DD learns to differentiate between real samples (x∈Xx\in X) and fake samples (xg=G(z)x_{g}=G(z) where z∼Nz\sim N). Concretely, there is an adversarial loss used to train DD and GG respectively:

This zero-sum game leads to GG learning how to generate samples that are indistinguishable from the original distribution. After training, DD is discarded and GG is used to generate content. When applied to imagery, this approach produces photo realistic images.

Numerous of variations and improvements of GANs have been proposed over the years. In the creation of deepfakes, there are two popular image translation frameworks which use the fundamental principles of GANs:

The pix2pix framework enables paired translations from one image domain to another (Isola et al., 2017). In pix2pix, GG tries to generate the image xgx_{g} given a visual context xcx_{c} as an input, and DD discriminates between (x,xc)(x,x_{c}) and (xg,xc)(x_{g},x_{c}). Moreover, GG is a an ED CNN with skip connections from EnEn to DeDe (called a U-Net) which enables GG to produce high fidelity imagery by bypassing the compression layers when needed. Later, pix2pixHD was proposed (Wang et al., 2018b) for generating high resolution imagery with better fidelity.

An improvement of pix2pix which enables image translation through unpaired training (Zhu et al., 2017). The network forms a cycle consisting of two GANs used to convert images from one domain to another, and then back again to ensure consistency with a cycle consistency loss (Lcyc\mathcal{L}_{cyc}).

An RNN is type of neural network that can handle sequential and variable length data. The network remembers is internal state after processing x(i−1)x^{(i-1)} and can use it to process x(i)x^{(i)} and so on. In deepfake creation, RNNs are often used to handle audio and sometimes video. More advanced versions of RNNs include long short-term memory (LSTM) and gate reccurent units (GRU).

4. Feature Representations

Most deep fake architectures use some form of intermediate representation to capture and sometimes manipulate ss and tt’s facial structure, pose, and expression. One way is to use the facial action coding system (FACS) and measure each of the face’s taxonomized action units (AU) (Ekman et al., 2002). Another way is to use monocular reconstruction to obtain a 3D morphable model (3DMM) of the head from a 2D image, where the pose and expression are parameterized by a set of vectors and matrices. Then use the parameters or a 3D rendering of the head itself. Some use a UV map of the head or body to give the network a better understanding of the shape’s orientation.

Another approach is to use image segmentation to help the network separate the different concepts (face, hair, etc). The most common representation is landmarks (a.k.a. key-points) which are a set of defined positions on the face or body which can be efficiently tracked using open source CV libraries. The landmarks are often presented to the networks as a 2D image with Gaussian points at each landmark. Some works separate the landmarks by channel to make it easier for the network to identity and associate them. Similarly, facial boundaries and body skeletons can also used.

For audio (speech), the most common approach is to split the audio into segments, and for each segment, measure the Mel-Cepstral Coefficients (MCC) which captures the dominant voice frequencies.

5. Deepfake Creation Basics

To generate xgx_{g}, reenactment and face swap networks follow some variation of this process (illustrated in Fig. 5): Pass xx through a pipeline that (1) detects and crops the face, (2) extracts intermediate representations, (3) generates a new face based on some driving signal (e.g., another face), and then (4) blends the generated face back into the target frame.

In general there are six approaches to driving an image:

Let a network work directly on the image and perform the mapping itself.

Train an ED network to disentangle the identity from the expression, and then modify/swap the encodings of the target the before passing it through the decoder.

Add an additional encoding (e.g., AU or embedding) before passing it to the decoder.

Convert the intermediate face/body representation to the desired identity/expression before generation (e.g., transform the boundaries with a secondary network or render a 3D model of the target with the desired expression).

Use the optical flow field from subsequent frames in a source video to drive the generator.

Create composite of the original content (hair, scene, etc) with a combination of the 3D rendering, warped image, or generated content, and pass the composite through another network (such as pix2pix) to refine the realism.

6. Generalization

A deepfake network may be trained or designed to work with only a specific set of target and source identities. An identity agnostic model is sometimes hard to achieve due to correlations learned by the model between ss and tt during training.

Let EE be some model or process for representing or extracting features from xx, and let MM be a trained model for performing replacement or reenactment. We identify three primary categories in regard to generalization:

A model that uses a specific identity to drive a specific identity: xg=Mt(Es(xs))x_{g}=M_{t}(E_{s}(x_{s}))

A model that uses any identity to drive a specific identity: xg=Mt(E(xs))x_{g}=M_{t}(E(x_{s}))

A model that uses any identity to drive any identity: xg=M(E1(xs),E2(xt))x_{g}=M(E_{1}(x_{s}),E_{2}(x_{t}))

7. Challenges

The following are some challenges in creating realistic deepfakes:

Generative networks are data driven, and therefore reflect the training data in their outputs. This means that high quality images of a specific identity requires a large number of samples of that identity. Moreover, access to a large dataset of the driver is typically much easier to obtain than the victim. As a result, over the last few years, researchers have worked hard to minimize the amount of training data required, and to enable the execution of a trained model on new target and source identities (unseen during training).

One way to train a neural network is to present the desired output to the model for each given input. This process of data pairing is a laborious an sometimes impractical when training on multiple identities and actions. To avoid this issue, many deepfake networks either (1) train in a self-supervised manner by using frames selected from the same video of tt, (2) use unpaired networks such as Cycle-GAN, or (3) utilize the encodings of an ED network.

Sometimes the identity of the driver (e.g., ss in reenactment) is partially transferred to xgx_{g}. This occurs when training on a single input identity, or when the network is trained on many identities but data pairing is done with the same identity. Some solutions proposed by researchers include attention mechanisms, few-shot learning, disentanglement, boundary conversions, and AdaIN or skip connections to carry the relevant information to the generator.

Occlusions are where part of xsx_{s} or xtx_{t} is obstructed with a hand, hair, glasses, or any other item. Another type of obstruction is the eyes and mouth region that may be hidden or dynamically changing. As a result, artifacts appear such as cropped imagery or inconsistent facial features. To mitigate this, works such as (Nirkin et al., 2019; Pumarola et al., 2019; Siarohin et al., 2019b) perform segmentation and in-painting on the obstructed areas.

Deepfake videos often produce more obvious artifacts such as flickering and jitter (Vougioukas et al., 2019a). This is because most deepfake networks process each frame individually with no context of the preceding frames. To mitigate this, some researchers either provide this context to GG and DD, implement temporal coherence losses, use RNNs, or perform a combination thereof.

Reenactment

In this section we present a chronological review of deep learning based reenactment, organized according to their class of identity generalization. Table 1 provides a summary and systematization of all the works mentioned in this section. Later, in Section 7, we contrast the various methods and identify the most significant approaches.

Expression reenactment turns an identity into a puppet, giving attackers the most flexibility to achieve their desired impact. Before we review the subject, we note that expression reenactment has been around long before deepfakes were popularized. In 2003, researchers morphed models of 3D scanned heads (Blanz et al., 2003). In 2005, it was shown how this can be done without a 3D model (Chang and Ezzat, 2005), and through warping with matching similar textures (Garrido et al., 2014). Later, between 2015 and 2018, Thies et al. demonstrated how 3D parametric models can be used to achieve high quality and real-time results with depth sensing and ordinary cameras ((Thies et al., 2015) and (Thies et al., 2016, 2018)).

Regardless, today deep learning approaches are recognized as the simplest way to generate believable content. To help the reader understand the networks and follow the text, we provide the model’s network schematics and loss functions in figures 6-8.

In 2017, the authors of (Xu et al., 2017) proposed using a CycleGAN for facial reenactment, without the need for data pairing. The two domains where video frames of ss and tt. However, to avoid artifacts in xgx_{g}, the authors note that both domains must share a similar distributions (e.g., poses and expressions).

In 2018, Bansal et al. proposed a generic translation network based on CycleGAN called Recycle-GAN (Bansal et al., 2018). Their framework improves temporal coherence and mitigates artifacts by including next-frame predictor networks for each domain. For facial reenactment, the authors train their network to translate the facial landmarks of xsx_{s} into portraits of xtx_{t}.

1.2. Many-to-One (Multiple Identities to a Single Identity)

In 2017, the authors of (Bao et al., 2017) proposed CVAE-GAN, a conditional VAE-GAN where the generator is conditioned on an attribute vector or class label. However, reenactment with CVAE-GAN requires manual attribute morphing by interpolating the latent variables (e.g., between target poses).

Later, in 2018, a large number of source-identity agnostic models were published, each proposing a different method to decoupling ss from tt:Although works such as (Olszewski et al., 2017) and (Zhou and Shi, 2017) achieved fully agnostic models (many-to-many) in 2017, their works were on low resolution or partial faces.

Facial Boundary Conversion. One approach was to first convert the structure of source’s facial boundaries to that of the target’s before passing them through the generator (Wu et al., 2018). In their framework ‘ReenactGAN’, the authors use a CycleGAN to transform the boundary bsb_{s} to the target’s face shape as btb_{t} before generating xgx_{g} with a pix2pix-like generator.

Temporal GANs. To improve the temporal coherence of deepfake videos, the authors of (Tulyakov et al., 2018) proposed MoCoGAN: a temporal GAN which generates videos while disentangling the motion and content (objects) in the process. Each frame is generated using a target expression label zcz_{c}, and a motion embedding zM(i)z_{M}^{(i)} for the ii-th frame, obtained from a noise seeded RNN. MoCoGAN uses two discriminators, one for realism (per frame) and one for temporal coherence (on the last TT frames).

In (Wang et al., 2018a), the authors proposed a framework called Vid2Vid, which is similar to pix2pix but for videos. Vid2Vid considers the temporal aspect by generating each frame based on the last LL source and generated frames. The model also considers optical flow to perform next-frame occlusion prediction (due to moving objects). Similar to pix2pixHD, a progressive training strategy is to generate high resolution imagery. In their evaluations, the authors demonstrate facial reenactment using the source’s facial boundaries. In comparison to MoCoGAN, Vid2Vid is more practical since it the deepfake is driven by xsx_{s} (e.g., an actor) instead of crafted labels.

The authors of (Kim et al., 2018) took temporal deepfakes one step further achieving complete facial reenactment (gaze, blinking, pose, mouth, etc.) with only one minute of training video. Their approach was to extract the source and target’s 3D facial models from 2D images using monocular reconstruction, and then for each frame, (1) transfer the facial pose and expression of the source’s 3D model to the target’s, and (2) produce xgx_{g} with a modified pix2pix framework, using the last 11 frames of rendered heads, UV maps, and gaze masks as the input.

1.3. Many-to-Many (Multiple IDs to Multiple IDs)

Label Driven Reenactment. The first attempts at identity agnostic models were made in 2017, where the authors of (Olszewski et al., 2017) used a conditional GAN (CGAN) for the task. Their approach was to (1) extract the inner-face regions as (xt,xs)(x_{t},x_{s}), and then (2) pass them to an ED to produce xgx_{g} subjected to L1\mathcal{L}_{1} and Ladv\mathcal{L}_{adv} losses. The challenge of using a CGAN was that the training data had to be paired (images of different identities with the same expression).

Going one step further, in (Zhou and Shi, 2017) the authors reenacted full portraits at low resolutions. Their approach was to decoupling the identities was to use a conditional adversarial autoencoder to disentangle the identity from the expression in the latent space. However, their approach is limited to driving xtx_{t} with discreet AU expression labels (fixed expressions) that capture xsx_{s}. A similar label based reenactment was presented in the evaluation of StarGAN (Choi et al., 2018); an architecture similar to CycleGAN but for NN domains (poses, expressions, etc).

Later, in 2018, the authors of (Pham et al., 2018) proposed GATH which can drive xtx_{t} using continuous action units (AU) as an input, extracted from xsx_{s}. Using continuous AUs enables smoother reenactments over previous approaches (Olszewski et al., 2017; Zhou and Shi, 2017; Choi et al., 2018). Their generator is ED network trained on the loss signals from using three other networks: (1) a discriminator, (2) an identity classifier, and (3) a pretrained AU estimator. The classifier shares the same hidden weights as the discriminator to disentangle the identity from the expressions.

Self-Attention Modeling. Similar to (Pham et al., 2018), another work called GANimation (Pumarola et al., 2019) reenacts faces through AU value inputs estimated from xsx_{s}. Their architecture uses an AU based generator that uses a self attention model to handle occlusions, and mitigate other artifacts. Furthermore, another network penalizes GG with an expression prediction loss, and shares its weights with the discriminator to encourage realistic expressions. Similar to CycleGAN, GANimation uses a cycle consistency loss which eliminates the need for image pairing.

Instead of relying on AU estimations, the authors of (Sanchez and Valstar, 2018) propose GANnotation which uses facial landmark images. Doing so enables the network to learn facial structure directly from the input but is more susceptible to identity leakage compared to AUs which are normalized. GANotation generates xgx_{g} based on (xt,ls)(x_{t},l_{s}), where lsl_{s} is the facial landmarks of xsx_{s}. The model uses the same self attention model as GANimation, but proposes a novel “triple consistency loss” to minimize artifacts in xgx_{g}. The loss teaches the network how to deal with intermediate poses/expressions not found in the training set. Given ls,ltl_{s},l_{t} and lzl_{z} sampled randomly from the same video, the loss is computed as

3D Parametric Approaches. Concurrent to the work of (Kim et al., 2018), other works also leveraged 3D parametric facial models to prevent identity leakage in the generation process. In (Shen et al., 2018a), the authors propose FaceID-GAN which can reenacts tt at oblique poses and high resolution. Their ED generator is trained in tandem with a 3DMM face model predictor, where the model parameters of xtx_{t} are used to transform xsx_{s} before being joined with the encoder’s embedding. Furthermore, to prevent identity leakage from xsx_{s} to xgx_{g}, FaceID-GAN incorporates an identification classifier within the adversarial game. The classifier has 2N2N outputs where the first NN outputs (corresponding to training set identities) are activated if the input is real and the rest are activated if it’s fake.

Later, the authors of (Shen et al., 2018a) proposed FaceFeat-GAN which improves the diversty of the faces while preserving the identity (Shen et al., 2018b). The approach is to use a set of GANs to learn facial feature distributions as encodings, and then use these generators to create new content with a decoder. Concretely, three encoder/predictor neural networks PP, QQ, and II, are trained on real images to extract feature vectors from portraits. PP predicts 3DMM parameters pp, QQ encodes the image as qq capturing general facial features using feedback from II, and II is an identity classifier trained to predict label yiy_{i}. Next two GANs, seeded with noise vectors, produce p′p^{\prime} and q′q^{\prime} while a third GAN is trained to reconstruct xtx_{t} from (p,q,yi)(p,q,y_{i}) and xgx_{g} from (p′,q′,yi)(p^{\prime},q^{\prime},y_{i}). To reenact xtx_{t}, (1) yty_{t} is predicted using II (even if the identity was previously unseen), (2) zpz_{p} and zqz_{q} are selected empirically to fit xsx_{s}, and (3) the third GAN’s generator uses (p′,q′,yt)(p^{\prime},q^{\prime},y_{t}) to create xgx_{g}. Although FaceFeat-GAN improves image diversty, it is less practical than FaceID-GAN since the GAN’s input seed zz be selected empirically to fit xsx_{s}.

In (Nagano et al., 2018), the authors present paGAN, a method for complete facial reenactment of a 3D avatar, using a single image of the target as input. An expression neutral image of xtx_{t} is used to generate a 3D model which is then driven by xsx_{s}. The driven model is used to create inputs for a U-Net generator: the rendered head, its UV map, its depth map, a masked image of xtx_{t} for texture, and a 2D mask indicating the gaze of xsx_{s}. Although paGAN is very efficient, the final deepfake is 3D rendered which detracts from the realism.

Using Multi-Modal Sources. In (Wiles et al., 2018) the authors propose X2Face which can reenact xtx_{t} with xsx_{s} or some other modality such as audio or a pose vector. X2Face uses two ED networks: an embedding network and a driving network. First the embedding network encodes 1-3 examples of the target’s face to vtv_{t}: the optical flow field required to transform xtx_{t} to a neutral pose and expression. Next, xtx_{t} is interpolated according to mtm_{t} producing xt′x_{t}^{{}^{\prime}}. Finally, the driving network maps xsx_{s} to the vector map vsv_{s}, crafted to interpolate xt′x_{t}^{{}^{\prime}} to xgx_{g}, having the pose and expression of xsx_{s}. During training, first L1\mathcal{L}_{1} loss is used between xtx_{t} and xgx_{g}, and then an identity loss is used between xsx_{s} and xgx_{g} using a pre-trained identity model trained on the VGG-Face Dataset. All interpolation is performed with a tensorflow interpolation layer to enable back propagation using xt′x_{t}^{{}^{\prime}} and xgx_{g}. The authors also show how the embedding of driving network can be mapped to other modalities such as audio and pose.

In 2019, nearly all works pursued identity agnostic models:

Facial Landmark & Boundary Conversion. In (Zhang et al., 2019a), the authors propose FaceSwapNet which tries to mitigate the issue of identity leakage from facial landmarks. First two encoders and a decoder are used to transfer the expression in landmark lsl_{s} to the face structure of ltl_{t}, denoted lgl_{g}. Then a generator network is used to convert xtx_{t} to xgx_{g} where lgl_{g} is injected into the network with AdaIn layers like a Style-GAN. The authors found that it is crucial to use triplet perceptual loss with an external VGG network.

In (Fu et al., 2019), the authors propose a method for high resolution reenactment and at oblique angles. A set of networks encode the source’s pose, expression, and the target’s facial boundary for a decoder that generates the reenacted boundary bgb_{g}. Finally, an ED network generates xgx_{g} using an encoding of xtx_{t}’s texture in its embedding. A multi-scale loss is used to improve quality and the authors utilize a small labeled dataset by training their model in a semi-supervised way.

In (Nirkin et al., 2019), the authors present FSGAN: a face swapping and facial reenactment model which can handle occlusions. For reenactment a pix2pixHD generator receives xtx_{t} and the source’s 3D facial landmarks lsl_{s}, represented as a 256x256x70 image (one channel for each of the 70 landmarks). The output is xgx_{g} and its segmentation map mgm_{g} with three channels (background, face, and hair). The generator is trained recurrently where each output is passed back as input for several iterations while lsl_{s} is interpolated incrementally from lsl_{s} to ltl_{t}. To improve results further, delaunay Triangulation and barycentric coordinate interpolation are used to generate content similar to the target’s pose. In contrast to other facial conversion methods (Zhang et al., 2019a; Fu et al., 2019), FSGAN uses fewer neural networks enabling real time reenactment at 30fps.

Latent Space Manipulation. In (Tripathy et al., 2019), the authors present a model called ICFace where the expression, pose, mouth, eye, and eyebrows of xtx_{t} can be driven independently. Their architecture is similar to a CycleGAN in that one generator translates xtx_{t} into a neutral expression domain as xtηx_{t}^{\eta} and another generator translates xtηx_{t}^{\eta} into an expression domain as xgx_{g}. Both generators ar conditioned on the target AU.

In (et al., 2019c) the authors propose an Additive Focal Variational Auto-encoder (AF-VAE) for high quality reenactment. This is accomplished by separating a C-VAE’s latent code into an appearance encoding eae_{a} and identity-agnostic expression coding exe_{x}. To capture a wide variety of factors in eae_{a} (e.g., age, illumination, complexion, …), the authors use an additive memory module during training which conditions the latent variables on a Gaussian mixture model, fitted to clustered set of facial boundaries. Subpixel convolutions were used in the decoder to mitigate artifacts and improve fidelity.

Warp-based Approaches. In the past, facial reenactment was done by warping the image xtx_{t} to the landmarks lsl_{s} (Averbuch-Elor et al., 2017). In (Geng et al., 2019), the authors propose wgGAN which uses the same approach but creates high-fidelity facial expressions by refining the image though a series of GANs: one for refining the warped face and another for in-painting the occlusions (eyes and mouth). A challenge with wgGAN is that the warping process is sensitive to head motion (change in pose).

In (Zhang et al., 2019b), the authors propose a system which can also control the gaze: a decoder generates xgx_{g} with an encoding of xtx_{t} as the input and a segmentation map of xsx_{s} as reenactment guidance via SPADE residual blocks. The authors blend xgx_{g} with a warped version, guided by the segmentation, to mitigate artifacts in the background.

To overcome issue of occlusions in the eyes and mouth, the authors of (Gu et al., 2020) use multiple images of tt as a reference, in contrast to (Geng et al., 2019) and (Zhang et al., 2019b) which only use one. In their approach (FLNet), the model is provided with NN samples of tt (XtX_{t}) having various mouth expressions, along with the landmark deltas between XtX_{t} and xsx_{s} (LtL_{t}). Their model is an ED (configured like GANimation (Pumarola et al., 2019)) which produces (1) NN encodings for a warped xgx_{g}, (2) an appearance encoding, and (3) a selection (weight) encoding. The encodings are then coverted into images using seperate CNN layers and merged together through masked multiplication. The entire model is trained end-to-end in a self supervised manner using frames of tt taken from different videos.

Motion-Content Disentanglement. In (Otberdout et al., 2019) the authors propose a GAN to reenact neutral expression faces with smooth animations. The authors describe the animations as temporal curves in 2D space, summarized as points on a spherical manifold by calculating their square-root velocity function (SRVF). A WGAN is used to complete this distribution given target expression labels, and a pix2pix GAN is used to convert the sequences of reconstructed landmarks into a video frames of the target.

In contrast to MoCoGAN (Tulyakov et al., 2018), the authors of (Wang et al., 2020a) propose ImaGINator: a conditional GAN which fuses both motion and content and uses with transposed 3D convolutions to capture the distinct spatio-temporal relationships. The GAN also uses a temporal discriminator, and to increase diversity, the authors train the temporal discriminator with some videos using the wrong label.

A challenge with works such as (Otberdout et al., 2019) and (Wang et al., 2020a) is that they are label driven and produce videos with a set number of frames. This makes the deepfake creation process manual and less practical. In contrast, the authors of (Siarohin et al., 2019a) propose Monkey-Net: a self supervised network for driving an image with an arbitrary video sequence. Similar to MoCoGAN (Tulyakov et al., 2018), the authors decouple the source’s content and motion. First a series of networks produce a motion heat map (optical flow) using the source and target’s key-points, and then an ED generator produces xgx_{g} using xsx_{s} and the optical flow (in its embedding).

Later in (Siarohin et al., 2019b), the authors extend Monkey-Net by improving the object appearance when large pose transformations occur. They accomplish this by (1) modeling motion around the keypoints using affine transformations, (2) updating the key-point loss function accordingly, and (3) having the motion generator predict an occlusion mask on the preceding frame for in-painting inference. Their work has been implemented as a free real-time reenactment tool for video chats, called Avitarify.https://github.com/alievk/avatarify

1.4. Few-Shot Learning

Towards the end of 2019 and into the beginning of 2020, researchers began looking into minimizing the amount of training data further via one-shot and few-shot learning.

In (Zakharov et al., 2019), the authors propose a few-shot model which works well at oblique angles. To accomplish this, the authors perform meta-transfer learning, where the network is first trained on many different identities and then fine-tuned on the target’s identity. Then, an identity encoding of xtx_{t} is obtained by averaging the encodings of kk sets of (xt,lt)(x_{t},l_{t}). Then a pix2pix GAN is used to generate xgx_{g} using lsl_{s} as an input, and the identity encoding via AdaIN layers. Unfortunately, the authors note that their method is sensitive to identity leakage.

In (Wang et al., 2019a) the authors of Vid2Vid (Section 4.1.2) extend their work with few-shot learning. They use a network weight generation module which utilizes an attention mechanism. The module learns to extract appearance patterns from a few samples of xtx_{t} which are injected into the video synthesis layers. In contrast to FLNet (Gu et al., 2020), (Zakharov et al., 2019), and (Wang et al., 2019a) which merge the multiple representations of tt before passing it through the generator. This approach is more efficient because it involves fewer passes through the model’s networks.

In (Ha et al., 2020), the authors propose MarioNETte which alleviates identity leakage when the pose of xsx_{s} is different than xtx_{t}. In contrast to other works which encode the identity separately or use of AdaIN layers, the authors use an image attention block and target feature alignment. This enables the model to better handle the differences between face structures. Finally, the identity is also preserved using a novel landmark transformer inspired by (Blanz and Vetter, 1999).

2. Mouth Reenactment (Dubbing)

In contrast to expression reenactment, mouth reenactment (a.k.a., video or image dubbing) is concerned with driving a target’s mouth with a segment of audio. Fig. 9 presents the relevant schematics for this section.

Obama Puppetry. In 2017, the authors of (Suwajanakorn et al., 2017) created a realistic reenactment of former president Obama. This was accomplished by (1) using a time delayed RNN over MFCC audio segments to generate a sequence of mouth landmarks (shapes), (2) generating the mouth textures (nose and mouth) by applying a weighted median to images with similar mouth shapes via PCA-space similarity, (3) refining the teeth by transferring the high frequency details other frames in the target video, and (4) by using dynamic programming to re-time the target video to match the source audio and blend the texture in.

Later that year, the authors of (Kumar et al., 2017) presented ObamaNet: a network that reenacts an individual’s mouth and voice using text as input instead of audio like (Suwajanakorn et al., 2017). The process is to (1) convert the source text to audio using Char2Wav (Sotelo et al., 2017), (2) generate a sequence of mouth-keypoints using a time-delayed LSTM on the audio, and (3) use a U-Net CNN to perform in-painting on a composite of the target video frame with a masked mouth and overlayed keypoints.

Later in 2018, Jalalifar et al. (Jalalifar et al., 2018) proposed a network that synthesizes the entire head portrait of Obama, and therefore does not require pose re-timing and can trained end-to-end, unlike (Suwajanakorn et al., 2017) and (Kumar et al., 2017). First, a bidirectional LSTM coverts MFCC audio segments into sequence of mouth landmarks, and then a pix2pix like network generates frames using the landmarks and a noise signal. After training, the pix2pix network is fine-tuned using a single video of the target to ensure consistent textures.

3D Parametric Approaches. Later on in 2019, the authors of (Fried et al., 2019) proposed a method for editing a transcript of a talking heads which, in turn, modifies the target’s mouth and speech accordingly. The approach is to (1) align phenomes to asa_{s}, (2) fit a 3D parametric head model to each frame of XtX_{t} like (Kim et al., 2018), (3) blend matching phenomes to create any new audio content, (4) animate the head model with the respective frames used during the blending process, and (5) generate XgX_{g} with a CGAN RNN using composites as inputs (rendered mouths placed over the original frame).

The authors of (Thies et al., 2019a) had a different approach: (1) animate a the reconstructed 3D head with the predicted blend shape parameters from asa_{s} using a DeepSpeech model for feature extraction, (2) use Deferred Neural Rendering (Thies et al., 2019b) to generate the mouth region, and then (3) use a network to blend the mouth into the original frame. Compared to previous works, the authors found that their approach only requires 2-3 minutes of video while producing very realistic results. This is because neural rendering can summarize textures with a high fidelity and operate on UV maps –mitigating artifacts in how the textures are mapped to the face.

2.2. Many-to-Many (Multiple IDs to Multiple IDs)

One of the first works to perform identity agnostic video dubbing was (Shimba et al., 2015). There the authors used an LSTM to map MFCC audio segments to the face shape. The face shapes were represented as the coefficients of an active appearance model (AAM), which were then used to retrieve the correct face shape of the target.

Improvements in Lip-sync. Noting a human’s sensitivity to temporal coherence, the authors of (Song et al., 2018) use a GAN with three discriminators: on the frames, video, and lip-sync. Frames are generated by (1) encoding each MFCC audio segment as(i)a_{s}^{(i)} and xtx_{t} with separate encoders, (2) passing the encodings through an RNN, and (3) decoding the outputs as xg(i)x_{g}^{(i)} using a decoder.

In (Yu et al., 2019c) the authors try to improve the lipsyncing with a textual context. A time-delayed LSTM is used to predict mouth landmarks given MFCC segments and the spoken text using a text-to-speech model. The target frames are then converted into sketches using an edge filter and the predicted mouth shapes are composited into them. Finally, a pix2pix like GAN with self-attention is used to generate the frames with both video and image conditional discriminators.

Compared to direct models such as direct models (Song et al., 2018; Yu et al., 2019c), the authors of (Chen et al., 2019) improve the lip-syncing by preventing the model from learning irrelevant correlations between the audiovisual signal and the speech content. This was accomplished with LSTM audio-to-landmark network and a landmark-to-identity CNN-RNN used in sequence. There, the facial landmarks are compressed with PCA and the attention mechanism from (Pumarola et al., 2019) is used to help focus the model on the relevant patterns. To improve synchronization further, the authors proposed a regression based discriminator which considers both sequence and content information.

EDs for Preventing Identity Leakage. The authors in (Zhou et al., 2019a) mitigate identity leakage by disentangling the speech and identity latent spaces using adversarial classifiers. Since their speech encoder is trained to project audio and video into the same latent space, the authors show how xgx_{g} can be driven using xsx_{s} or asa_{s}.

In (Jamaludin et al., 2019), the authors propose Speech2Vid which also uses separate encoders for audio and identity. However, to capture the identity better, the identity encoder EnIEn_{I} uses a concatenation of five images of the target, and there are skip connections from the EnIEn_{I} to the decoder. To blend the mouth in better, a third ‘context’ encoder is used to encourage in-painting. Finally, a VDSR CNN is applied to xgx_{g} to sharpen the image.

A disadvantage with (Zhou et al., 2019a) and (Jamaludin et al., 2019) is that they cannot control facial expressions and blinking. To resolve this, the authors in (Vougioukas et al., 2019a) generate frames with a stride transposed CNN decoder on GRU-generated noise, in addition to the audio and identity encodings. Their video discriminator uses two RNNs for both the audio and video. When applying the L1 loss, the authors focus on the lower half of the face to encourage better lip sync quality over facial expressions.

Later in (Vougioukas et al., 2019b), the same authors improve the temporal coherence by splitting the video discriminator into two: (1) for temporal realism in mouth to audio synchronization, and (2) for temporal realism in overall facial expressions. Then in (Kefalas et al., 2019), the authors tune their approach further by fusing the encodings (audio, identity, and noise) with a polynomial fusion layer as opposed to simply concatenating the encodings together. Doing so makes the network less sensitive to large facial motions compared to (Vougioukas et al., 2019b) and (Jamaludin et al., 2019).

3. Pose Reenactment

Most deep learning works in this domain focus on the problem of face frontalization. However, there are some works which focus on facial pose reenactment.

In (Hu et al., 2018) the authors use a U-Net to convert (xt,lt,ls)(x_{t},l_{t},l_{s}) into xgx_{g} using a GAN with two discriminators: one conditioned with the neutral pose image, and the other conditioned with the landmarks. In (Tran et al., 2018), the authors propose DR-GAN for pose-invariant face recognition. To adjust the pose of xtx_{t}, the authors use an ED GAN which encodes xtx_{t} as ete_{t}, and then decodes (et,ps,z)(e_{t},p_{s},z) as xgx_{g}, where psp_{s} is the source’s pose vector and zz is a noise vector. Compared to (Hu et al., 2018), (Tran et al., 2018) has the flexibility of manipulating the encodings for different tasks and the authors improve the quality of xgx_{g} by averaging multiple examples of the identity encoding before passing it through the decoder (similar to (Gu et al., 2020; Zakharov et al., 2019; Wang et al., 2019a)). In (Cao et al., 2019), the authors suggest using two GANs: The first frontalizes the face and produces a UV map, and second rotates the face, given the target angle as an injected embedding. The result is that each model performs a less complex operation and can therefore the models collectively can produce a higher quality image.

4. Gaze Reenactment

There are only a few deep learning works which have focused on gaze reenactment. In (Ganin et al., 2016) the authors convert a cropped eye xtx_{t}, its landmarks, and the source angle, to a flow (vector) field using a 2-scale CNN. xgx_{g} is then generated by applying a flow field to xtx_{t} to warping it to the source angle. The authors then correct the illumination of xgx_{g} with a second CNN. A challenge with (Ganin et al., 2016) is that the head must be frontal to avoid inconsistencies due to pose and perspective. To mitigate this issue, the authors of (Yu et al., 2019b) proposed the Gaze Redirection Network (GRN). In GRN, the target’s cropped eye, head pose, and source angle are encoded separately and then passed though an ED network to generate an optical flow field. The field is used to warp xtx_{t} into xgx_{g}. To overcome the lack of training data and the challenge of data pairing, the authors (1) pre-train their network on 3D synthesized examples, (2) further tune their network on real images, and then (3) fine tune their network on 3-10 examples of the target.

5. Body Reenactment

Several facial reenactment papers from Section 4.1 discuss body reenactment too. For example, Vid2Vid (Wang et al., 2018a, 2019a), MocoGAN (Tulyakov et al., 2018), and others (Siarohin et al., 2019a, b). In this section, we focus on methods which specifically target body reenactment. Schematics for some of these architectures can be found in Fig. 10.

In the work (Liu et al., 2019a), the authors perform facial reenactment with the upper-body as well (arms and hands). The approach is to (1) use a pix2pixHD GAN to convert the source’s facial boundaries to the targets, (2) and then paste them onto a captured pose skeleton of the source, and (3) use a pix2pixHD GAN to generate xgx_{g} from the composite.

5.2. Many-to-One (Multiple Identities to a Single Identity)

Dance Reenactment. In (Chan et al., 2019) the authors make people dance using a target specific pix2pixHD GAN with a custom loss function. The generator receives an image of the captured pose skeleton and the discriminator receives the current and last image conditioned on their poses. The quality of face is then improved with a residual predicted by an additional pix2pixHD GAN, given the face region of the pose. A many-to-one relationship is achieved by normalizing the input pose to that of the target’s.

The authors of (Liu et al., 2019c) then tried to overcome artifacts which occur in (Chan et al., 2019) such stretched limbs due to incorrectly detected pose skeletons. They used photogrammetry software on hundreds of images of the target, and then reenacted the 3D rendering of the target’s body. The rendering, partitioned depth map, and background are then passed to a pix2pix model for image generation, using an attention loss.

Another artifact in (Chan et al., 2019) was that the model could not generalize well to unseen poses. To improve the generalization, the authors of (Aberman et al., 2019) trained their network on many identities other than ss and tt. First they trained the GAN on paired data (the same identity doing different poses) and then later added another discriminator to evaluate the temporal coherence given (1) xg(i)x_{g}^{(i)} driven by another video, and (2) the optical flow predicted version.

A challenge with the previous works was that they required a lots of training data. This was reduced from about an hour of video footage to only 3 minutes in (Zhou et al., 2019b) by segmenting and orienting the limbs of xtx_{t} according to xsx_{s} before the generation step. Then a pix2pixHD GAN uses this composition and the last kk frames’ poses to generate the body. Finally, another pix2pixHD GAN is used to blend the body into the background.

5.3. Many-to-Many (Multiple IDs to Multiple IDs)

Pose Alignment. In (Siarohin et al., 2018) the authors try to resolve the issue of misalignment when using pix2pix like architectures. They propose ‘deformable skip connections’ which help orient the shuttled feature maps according to the source pose. The authors also propose a novel nearest neighbor loss instead of using L1 or L2 losses. To modify unseen identities at test time, an encoding of xtx_{t} is passed to the decoder’s inner layers.

Although the work of (Siarohin et al., 2018) helps align the general images, artifacts can still occur when xsx_{s} and xtx_{t} have very different poses. To resolve this, the authors of (Zhu et al., 2019) use novel Pose-Attentional Transfer blocks (PATB) inside their GAN-based generator. The architecture passes xtx_{t} and the poses psp_{s} concatenated with ptp_{t} through separate encoders which are passed though a series of PATBs before being decoded. The PATBs progressively transfer regional information of the poses to regions of the image to ultimately create a body that has better shape and appearance consistency.

Pose Warping. In (Neverova et al., 2018) the authors use a pre-trained DensePose network (Alp Güler et al., 2018) to refine a predicted pose with a warped and in-painted DensePose UV spatial map of the target. Since the spatial map covers all surfaces of the body, the generated image has improved texture consistency. In contrast to (Zhu et al., 2019; Siarohin et al., 2018) which uses feature mappings to alleviate misalignment, the authors of (Zablotskaia et al., 2019) use warping which reduces the complexity of the network’s task. Their model, called DwNet, uses a ‘warp module’ in an ED network to encode xt(i−1)x_{t}^{(i-1)} warped to ps(i)p_{s}^{(i)}, where pp is a UV body map of a pose obtained a DensePose network.

A challenge with the alignment techniques of the previous works is that the body’s 3D shape and limb scales are not considered by the network resulting in identity leakage from xsx_{s}. In (Liu et al., 2019b), the authors counter this issue with their Liquid Warping GAN. This is accomplished by predicting target and source’s 3D bodies with the model in (Kanazawa et al., 2018) and then by translating the two through a novel liquid warping block (LWB) in their generator. Specifically, the estimated UV maps of xsx_{s} and xtx_{t}, along with their calculated transformation flow, are passed through a three stream generator which produces (1) the background via in-painting, (2) a reconstruction of the xsx_{s} and its mask for feature mapping, and (3) the reenacted foreground and its mask. The latter two streams use a shared LWB to help the networks address multiple sources (appearance, pose, and identity). The final image is obtained through masked multiplication and the system is trained end-to-end.

Background Foreground Compositing. In (Balakrishnan et al., 2018), the authors break the process down into three stages, trained end-to-end: (1) use a U-Net to segment xtx_{t}’s body parts and then orient them according to the source pose psp_{s}, (2) use a second U-Net to generate the body xgx_{g} from the composite, and (3) use a third U-Net to perform in-painting on the background and paste xgx_{g} into it. The authors of (et al., 2018) then streamlined this process by using a single ED GAN network to disentangle the foreground appearance (body), background appearance, and pose. Furthermore, by using an ED network, the user gains control over each of these aspects. This is accomplished by segmenting each of these aspects before passing them through encoders. To improve the control over the compositing, the authors of (De Bem et al., 2019) used a CVAE-GAN. This enabled the authors to change the pose and appearance of bodies individually. The approach was to condition the network on heatmaps of the predicted pose and skeleton.

5.4. Few-Shot Learning

In (Lee et al., 2019), the authors demonstrate the few-shot learning technique of (Finn et al., 2017) on a pix2pixHD network and the network of (Balakrishnan et al., 2018). Using just a few sample images, they were able to transfer the resemblance of a target to new videos in the wild.

Replacement

The network schematics and summary of works for replacement deepfakes can be found in Fig. 12 and Table 2 respectively.

At first, face swapping was a manual process accomplished using tools such as Photoshop. More automated systems first appeared between 2004-08 in (Blanz et al., 2004) and (Bitouk et al., 2008). Later, fully automated methods were proposed in (Vlasic et al., 2006; Dale et al., 2011; Kemelmacher-Shlizerman, 2016) and (Nirkin et al., 2018) using methods such as warping and reconstructed 3D morphable face models.

Online Communities. After the Reddit user ‘deepfakes’ was exposed in the media, researchers and online communities began finding improved ways to perform face swapping with deep neural networks. The original deepfake network, published by the Reddit user, is an ED network (visualized in Fig. 11). The architecture consists of one encoder EnEn and two decoders DesDe_{s} and DetDe_{t}. The components are trained concurrently as two autoencoders: Des(En(xs))=x^sDe_{s}(En(x_{s}))=\hat{x}_{s} and Det(En(xt))=x^tDe_{t}(En(x_{t}))=\hat{x}_{t}, where xx is a cropped face image. As a result, EnEn learns to map ss and tt to a shared latent space, such that

Currently, there are a number of open source face swapping tools on GitHub based on the original network. One of the most popular is DeepFaceLab (iperov, 2019). Their current version offers a wide variety of model configurations, including adversarial training, residual blocks, a style transfer loss, and masked loss to improve the quality of the face and eyes. To help the network map the target’s identity into arbitrary face shapes, the training set is augmented with random face warps.

Another tool called FaceSwap-GAN (shaoanlu, 2018) follows a similar architecture, but uses a denoising autoencoder with a self-attention mechanisms, and offers cycle-consistency loss which can reduce the identity leakage and increase the image fidelity. The decoders in FaceSwap-GAN also generate segmentation masks which helps the model handle occlusions and is used to blend xgx_{g} back into the target frame. Finally, (dee, 2017) is another open source tool that provides a GUI. Their software comes with 10 popular implementations, including that of (iperov, 2019), and multiple variations of the original Redit user’s code.

1.2. One-to-Many (Single Identity to Multiple Identities)

In (Korshunova et al., 2017), the authors use a modified style transfer with CNN, where the content is xtx_{t} and the style is the identity of xsx_{s}. The process is (1) align xtx_{t} to a reference xsx_{s}, (2) transfer the identity of ss to the image using a multi scale CNN, trained with style loss on images of ss, and (3) align the output to xtx_{t} and blend the face back in with a segmentation mask.

1.3. Many-to-Many (Multiple IDs to Multiple IDs)

One of the first identity agnostic methods was (Olszewski et al., 2017), mentioned in Section 4.1.3. However, to train this CGAN, one needs a dataset of paired faces with different identities having the same expression.

Disentanglement with EDs. However, To provide more control over the In (Bao et al., 2018) the authors us an ED to disentangle the identity from the attributes (pose, hair, background, and lighting) during the training process. The identity encodings are the last pooling layer of a face classifier, and the attribute encoder is trained using a weighted L2 lossand a KL divergence loss to mitigate identity leakage. The authors also show that they can adjust attributes, expression, and pose via interpolation of the encodings. Instead of swapping identities, the authors of (Sun et al., 2018) wanted to variably obfuscate the target’s identity. To accomplish this, the authors used an ED to predict the 3D head parameters which where either modified or replaced with the source’s. Finally a GAN was used to in-paint the face of xtx_{t} given the modified head model parameters.

Disentanglement with VAEs. In (Natsume et al., 2018b), the authors propose RSGAN: a VAE-GAN consisting of two VAEs and a decoder. One VAE encodes the hair region and the other encodes the face region, where both are conditioned on a predicted attribute vector cc describing xx. Since VAEs are used, the facial attributes can be edited through cc.

In contrast to (Natsume et al., 2018b), the authors of (Natsume et al., 2018a) use a VAE to prepare the content for the generator, and use a network to perform the blending via in-painting. A single VAE-ED network is run on xsx_{s} and then xtx_{t} producing encodings for the face of xsx_{s} and the landmarks of xtx_{t}. To perform a face swap, a generator receives the masked portrait of xtx_{t} and performs in-painting on the masked face. The generator uses the landmark encodings in its embedding layer. During training, randomly generated faces are used with triplet loss on the encodings to preserve identities.

Face Occlusions. FSGAN (Nirkin et al., 2019), mentioned Section 4.1.3, is also capable of face swapping and can handle occlusions. After the face reenactment generator produces xrx_{r}, a second network predicts the target’s segmentation mask mtm_{t}. Then (xr⟨f⟩,mt)(x_{r}^{\langle f\rangle},m_{t}) is passed to a third network that performs in-painting for occlusion correction. Finally a fourth network blends the corrected face into xtx_{t} while considering ethnicity and lighting. Instead of using interpolation like (Nirkin et al., 2019), the authors of (Li et al., 2019a) propose FaceShifter which uses novel Adaptive Attentional Denormalization layers (AAD) to transfer localized feature maps between the faces. In contrast to (Nirkin et al., 2019), FaceShifter reduces the number of operations by handling the occlusions through a refinement network trained to consider the delta between the original xtx_{t} and a reconstructed x^t\hat{x}_{t}.

1.4. Few-Shot Learning

The same author of FaceSwap-GAN (shaoanlu, 2018) also hosts few-shot approach online dubbed “One Model to Swap Them All” (Shaoanlu, 2019). In this version the generator receives (xs⟨f⟩,xt⟨f⟩,mt)(x_{s}^{\langle f\rangle},x_{t}^{\langle f\rangle},m_{t}) where its encoder is conditioned on VGGFace2 features of xtx_{t} using FC-AdaIN layers, and its decoder is conditioned on xtx_{t} and the face structure mtm_{t} via layer concatenations and SPADE-ResBlocks respectively. Two discriminators are used: one on image quality given the face segmentation and the other on the identities.

2. Transfer

Although face transfers precede face swaps, today there are very few works that use deep learning for this task. However, we note that face a transfer is equivalent to performing self-reenactment on a face swapped portrait. Therefore, high quality face a transfers can be achieved by combining a method from Section 4.1 and Section 5.1.

In 2018, the authors of (Moniz et al., 2018) proposed DepthNets: an unsupervised network for capturing facial landmarks and translating the pose from one identity to another. The authors use a Siamese network to predict a transformation matrix that maps the xsx_{s}’s 3D facial landmarks to the corresponding 2D landmarks of xtx_{t}. A 3D renderer (OpenGL) is then used to warp xs⟨f⟩x_{s}^{\langle f\rangle} to the source pose ltl_{t}, and the composition is refined using a CycleGAN. Since warping is involved, the approach is sensitive to occlusions.

Later in 2019, the authors of (Xiao et al., 2019) proposed a self-supervised network which can change the identity of an object within an image. Their ED disentangles the identity from an objects pose using a novel disentanglement loss. Furthermore to handle misaligned poses, an L1 loss is computed using a pixel mapped version of xgx_{g} to xsx_{s} (using the weights of the identity encoder). Similarly, the authors of (Li et al., 2019d) proposed a method disentangled identity transfer. However neither (Xiao et al., 2019) or (Li et al., 2019d) were explicitly performed on faces.

Countermeasures

In general, countermeasures to malicious deepfakes can be categorized as either detection or prevention. We will now briefly discuss each accordingly. A summary and systematization of the deepfake detection methods can be found in Table 3.

The subject of image forgery detection is a well researched subject (Zheng et al., 2019). In our review of detection methods, we will focus on works which specifically deal with detecting deepfakes of humans.

Deepfakes often generate artifacts which may be subtle to humans, but can be easily detected using machine learning and forensic analysis. Some works identify deepfakes by searching for specific artifacts. We identify seven types of artifacts: Spatial artifacts in blending, environments, and forensics; temporal artifacts in behavior, physiology, synchronization, and coherence.

Blending (spatial). Some artifacts appear where the generated content is blended back into the frame. To help emphasize these artifacts to a learner, researchers have proposed edge detectors, quality measures, and frequency analysis (Agarwal et al., 2017; Zhang et al., 2017; Akhtar and Dasgupta, [n.d.]; Mo et al., 2018; Durall et al., 2019). In (Li et al., 2020a) the authors follow a more explicit approach to detecting the boundary. They trained a CNN network to predict an image’s blending boundary and a label (real or fake). Instead of using a deepfake dataset, the authors trained their network on a dataset of face swaps generated by splicing similar faces found through facial landmark similarity. By doing so, the model has the advantage that is focuses on the blending boundary and not other artifacts caused by the generative model.

Environment (spatial). The content of a fake face can be anomalous in context to the rest of the frame. For example, residuals from face warping processes (Li and Lyu, 2019b, a; Li et al., 2019e), lighting (Straub, 2019), and varying fidelity (Korshunov and Marcel, 2018a) can indicate the presence of generated content. In (Li et al., 2020b), the authors follow a different approach by contrasting the generated foreground to the (untampered) background using a patch and pair CNN. The authors of (Nirkin et al., 2020) also contrast the fore/background but enable a network to identify the distinguishing features automatically. They accomplish this by (1) encoding the face and context (hair and background) with an ED and (2) passing the difference between the encodings with the complete image (encoded) to a classifier.

Forensics (spatial). Several works detect deepfakes by analyzing subtle features and patterns left by the model. In (Yu et al., 2019a) and (Marra et al., 2019), the authors found that GANs leave unique fingerprints and show how it is possible to classify the generator given the content, even in the presence of compression and noise. In (Koopman et al., 2018) the authors analyze a camera’s unique sensor noise (PRNU) to detect pasted content. To focus on the residuals, the authors of (Masi et al., 2020) use a two stream ED to encode the color image and a frequency enhanced version using “Laplacian of Gaussian layers” (LoG). The two encodings are then fed through an LSTM which then classifies the video based on a sequence of frames.

Instead of searching for residuals, the authors of (Yang et al., 2019) search for imperfections and found that deepfakes tend to have inconsistent head poses. Therefore, they detect deepfakes by predicting and monitoring facial landmarks. The authors of (Wang et al., 2020b) had a different approach by training classifier to focus on the imperfections instead of the residuals. This was accomplished by using a dataset generated using a ProGAN instead of other GANs since the ProGAN’s images contain the least amount of frequency artifacts. In contrast to (Wang et al., 2020b), the authors in (Guo et al., 2020) use a network to emphasize the residuals and suppress the imperfections in a preprocessing step for a classifier. Their network uses adaptive convolutional layers that predict residuals to maximize the artifacts’ influence. Although this approach may help the network identify artifacts better, it may not generalize as well to new types of artifacts.

Behavior (temporal). With large amounts of data on the target, mannerisms and other behaviors can be monitored for anomalies. For example, in (Agarwal et al., 2019) the authors protect world leaders from a wide variety of deepfake attacks by modeling their recorded stock footage. Recently, the authors of (Mittal et al., 2020) showed how behavior can be used with no reference footage of the target. The approach is to detect discrepancies in the perceived emotion extracted from the clip’s audio and video content. The authors use a custom Siamese network to consider the audio and video emotions when contrasted to real and fake videos.

Physiology (temporal). In 2014, researchers hypothesized that generated content will lack physiological signals and identified computer generated faces by monitoring their heart rate (Conotter et al., 2014). Regarding deepfakes, (Ciftci and Demir, 2019) monitored blood volume patterns (pulse) under the skin, and (Li et al., 2018) took a more robust approach by monitoring irregular eye blinking patterns. Instead of detecting deepfakes, the authors of (Ciftci et al., 2020) use the pulse signal to help determine the model used to create the deepfake.

Synchronization (temporal). Inconsistencies are also a revealing factor. In (Korshunov and Marcel, 2018b) and (et al., 2019b), the authors noticed that video dubbing attacks can be detected my correlating the speech to landmarks around the mouth. Later, in (Agarwal et al., 2020), the authors refined the approach by detecting when visemes (mouth shapes) are inconsistent with the spoken phonemes (utternaces). In particular, they focus on phonemes where the mouth is fully closed (B, P, M) since deepfakes in the wild tend to fail in generating these visemes.

Coherence (temporal). As noted in Section 4.1, realistic temporal coherence is challenging to generate, and some authors capitalize on the resulting artifacts to detect the fake content. For example, (Guera and Delp, 2018) uses an RNN to detect artifacts such as flickers and jitter, and (Sabir et al., 2019) uses an LSTM on the face region only. In (Chan et al., 2019) a classifier is trained pairs of sequential frames and in (Amerini et al., 2019) the authors refine the network’s focus by monitoring the frames’ optical flow. Later the same authors use an LSTM to predict the next frame, and expose deepfakes when the reconstruction error is high (Amerini and Caldelli, 2020).

1.2. Undirected Approaches

Instead of focusing on a specific artifact, some authors train deep neural networks as generic classifiers, and let the network decide which features to analyze. In general, researchers have taken one of two approaches: classification or anomaly detection.

Classification. In (Marra et al., 2018; Rossler et al., 2019; Nguyen et al., 2019b), it was shown that deep neural networks tend to perform better than traditional image forensic tools on compressed imagery. Various authors then demonstrated how standard CNN architectures can effectively detect deepfake videos (Afchar et al., 2018; Do et al., 2018; Tariq et al., 2018; Ding et al., 2019). In (Hsu et al., 2020), the authors train the CNN as a Siamese network using contrasting examples of real and fake images. In (Fernando et al., 2019), the authors were concerned that a CNN can only detect the attacks on which they trained. To close this gap, the authors propose using Hierarchical Memory Network (HMN) architecture which considers the contents of the face and previously seen faces. The network encodes the face region which is then processed using a bidirectional GRU while applying an attention mechanism. The final encoding is then passed to a memory module, which compares it to recently seen encodings and makes a prediction. Later, in (Rana and Sung, 2020), the authors use an ensemble approach and leverage the predictions of seven deepfake CNNs by passing their predicitons to a meta classifer. Doing so produces results which are more robust (fewer false positives) than using any single model. In (de Lima et al., 2020), the authors tried a variety of different classic spatio-temproal networks and feature extractors as a baseline for temporal deepfake detection. They found that a 3D CNN, which looks at multiple frames at once, out performs both recurrent networks and a the state of the art ID3 architecture.

To localize the tampered areas, some works train networks to predicting masks learned from a ground truth dataset, or by mapping the neural activations back to the raw image (Nguyen et al., 2019a; Du et al., 2019; Stehouwer et al., 2019; Li et al., 2019c).

In general, we note that the use of classifiers to detect deepfakes is problematic since an attacker can evade detection via adversarial machine learning. We will discuss this issue further in Section 7.2.

Anomaly Detection. In contrast to classification, anomaly detection models are trained on the normal data and then detect outliers during deployment. By doing so, these methods do not make assumptions on how the attacks look and thus generalize better to unknown creation methods. The authors of (Wang et al., 2019b) follow this approach by measuring the neural activation (coverage) of a face recognition network. By doing so, the model is able to overcome noise and other distortions, by obtaining a stronger signal from than just using the raw pixels. Similarly, in (Khalid and Woo, 2020) a one-class VAE is trained to used to reconstruct real images. Then, for new images, an anomaly score is computed by taking the MSE between mean component of the encoded image and the mean component of the reconstructed image. Alternatively, the authors of (Bao et al., 2018) measure an input’s embedding distance to real samples using an ED’s latent space. The difference between these works is that (Wang et al., 2019b) and (Khalid and Woo, 2020) rely on a model’s inability to process unknown patterns while (Bao et al., 2018) contrasts the model’s representations.

Instead of using a neural network directly, the authors of (Fernandes et al., 2020) use a state of the art attribution based confidence metric (ABC). To detect a fake image, the ABC is used to determine if the image fits the training distribution of a pretrained face recognition network (e.g., VGG).

2. Prevention & Mitigation

Data Provenance. To prevent deepfakes, some have suggested that data provenance of multimedia should be tracked through distributed ledgers and blockchain networks (Fraga-Lamas and Fernandez-Carames, 2019). In (et al., 2019a) the authors suggest that the content should be ranked by participants and AI. In contrast, (Hasan and Salah, 2019) proposes that the content should authenticated and managed as a global file system over Etherium smart contracts.

Counter Attacks. To combat deepfakes, the authors of (Li et al., 2019f) show how adversarial machine learning can be used to disrupt and corrupt deepfake networks. The authors perform adversarial machine learning to add crafted noise perturbations to xx, which prevents deepfake technologies from locating a proper face in xx. In a different approach, the authors of (Shan et al., 2020) use adversarial noise to change the identity of the face so that web crawlers will not be able find the image of tt to train their model.

Discussion

In general, there is a different cost and payoff for each deepfake creation method. However, the most effective and threatening deepfakes are those which are (1) the most practical to implement [Training Data, Execution Speed, and Accessibility] and (2) are the most believable to the victim [Quality]:

Models trained on numerous samples of the target often yield better results (e.g., (Chan et al., 2019; Fried et al., 2019; iperov, 2019; Jalalifar et al., 2018; Kumar et al., 2017; Liu et al., 2019a; Suwajanakorn et al., 2017; Wu et al., 2018)). For example, in 2017, (Suwajanakorn et al., 2017) produced an extremely believable reenactment of Obama which exceeds the quality of recent works. However, these models require many hours footage for training, and are therefore are only suitable for exposed targets such as actors, CEOs, and political leaders. An attacker who wants to commit defamation, impersonation, or a scam on an arbitrary individual will need to use a many-to-many or few-shot approach. On the other hand, most of these methods rely on a single reference of tt and are therefore prone to generating artifacts. This is because the model must ‘imagine’ missing information (e.g., different poses and occlusions). Therefore, approaches which provide the model with a limited number of reference samples (Gu et al., 2020; Ha et al., 2020; Tran et al., 2018; Wang et al., 2019a; Wiles et al., 2018; Yu et al., 2019b; Zakharov et al., 2019) strike the best balance between data and quality.

The trade-off between these aspects depends on whether the attack is online (interactive) or offline (stored media). Social engineering attacks involving deepfakes are likely to be online and thus require real-time speeds. However, high resolution models have many parameters and sometimes use several networks (e.g., (Fu et al., 2019)) and some process multiple frames to provide temporal coherence (e.g., (Bansal et al., 2018; Kim et al., 2018; Wang et al., 2018a)). Other methods may be slowed down due to their pre/post-processing steps, such as warping (Geng et al., 2019; Gu et al., 2020; Zhang et al., 2019b), UV mapping or segmentation prediction (Cao et al., 2019; Olszewski et al., 2017; Nagano et al., 2018; Yu et al., 2019b), and the use of refinement networks (Chan et al., 2019; Geng et al., 2019; Kim et al., 2018; Li et al., 2019a; Moniz et al., 2018; Thies et al., 2019a). To the best of our knowledge, (Jamaludin et al., 2019; Nirkin et al., 2019; Nagano et al., 2018; Korshunova et al., 2017) and (Siarohin et al., 2019b) are the only papers which claim to generate real time deepfakes, yet they subjectively tend to be blurry or distort the face. Regardless, a victim is likely fall for an imperfect deepfake in a social engineering attack when placed under pressure in a false pretext (Workman, 2008). Moreover, it is likely that an attacker will implement a complex method at a lower resolution to speed up the frame rate. In which case, methods that have texture artifacts would be preferred over those which produce shape or identity flaws (e.g., (Siarohin et al., 2019b) vs (Zakharov et al., 2019)). For attacks that are not real-time (e.g, fake news), resolution and fidelity is critical. In these cases, works that produce high quality images and videos with temporal coherence are the best candidates (e.g., (Ha et al., 2020; Wang et al., 2018a)).

We also note that availability and reproducibility are key factors in the proliferation of new technologies. Works that publish their code and datasets online (e.g., (Wiles et al., 2018; Wu et al., 2018; Sanchez and Valstar, 2018; Tulyakov et al., 2018; Kefalas et al., 2019; Siarohin et al., 2019b)) are more likely to be used by researchers and criminals compared to those which are unavailable (Kim et al., 2018; Pham et al., 2018; Zhang et al., 2019a; Wang et al., 2020a; Ha et al., 2020; Fried et al., 2019; Thies et al., 2019a; Aberman et al., 2019; Nirkin et al., 2019) or require highly specific or private datasets (Yu et al., 2019b; Ganin et al., 2016; Nagano et al., 2018). This is because the payoff in implementing a paper is minor compared to using a functional and effective method available online. Of course, this does not include state-actors who have plenty of time and funding.

We have also observed that approaches which augment a network’s inputs with synthetic ones produce better results in terms of quality and stability. For example, by rotating limbs (Liu et al., 2019a; Zhou et al., 2019b), refining rendered heads (Nagano et al., 2018; Fried et al., 2019; Yu et al., 2019c; Balakrishnan et al., 2018; Thies et al., 2019a; Wang et al., 2018b), providing warped imagery (Zablotskaia et al., 2019; Neverova et al., 2018; Moniz et al., 2018; Geng et al., 2019) and UV maps (Cao et al., 2019; Kim et al., 2018; Otberdout et al., 2019; Zablotskaia et al., 2019; Gu et al., 2020). This is because the provided contextual information reduces the problem’s complexity for the neural network.

Given these considerations, in our opinion, the most significant and available deepfake technologies today are (Siarohin et al., 2019b) for facial reenactment because of it’s efficiency and practicality; (Chen et al., 2019) for mouth reenactment because of its quality; and (iperov, 2019) for face replacement because its high fidelity and wide spread use. However, this is a subjective opinion based on the samples provided online and in the respective papers. A comparative research study, where the methods are trained on the same dataset and evaluated by a number of people is necessary to determine the best quality deepfake in each category.

1.2. Research Trends

Over the last few years there has been a shift towards identity agnostic models and high resolution deepfakes. Some notable advancements include (1) unpaired self-supervised training techniques to reduce the amount of initial training data, (2) one/few-shot learning which enables identity theft with a single profile picture, (3) improvements of face quality and identity through AdaIN layers, disentanglement, and pix2pixHD network components, (4) fluid and realistic videos through temporal discriminators and optical flow prediction, and (5) the mitigation of boundary artifacts by using secondary networks to blend composites into seamless imagery (e.g., (Wang et al., 2018b; Fried et al., 2019; Thies et al., 2019a)).

Another large advancement in this domain was the use of perceptual loss on a pre-trained VGG Face recognition network. The approach boosts the facial quality significantly, and as a result, has been adopted in popular online deepfake tools (dee, 2017; shaoanlu, 2018). Another advancement being adopted is the use of a network pipeline. Instead of enforcing a set of global losses on a single network, a pipeline of networks is used where each network is tasked with a different responsibility (conversion, generation, occlusions, blending, etc.) This give more control over the final output and has been able to mitigate most of the challenges mention in Section 3.7.

1.3. Current Limitations

Aside from quality, there are a few limitations with the current deepfake technologies. First, for reenactment, content is always driven and generated with a frontal pose. This limits the reenactment to a very static performance. Today, this is avoided by face swapping the identity onto a lookalike’s body, but a good match is not always possible and this approach has limited flexibility. Second, reenactments and replacements depend on the driver’s performance to deliver the identity’s personality. We believe that next generation deepfakes will utilize videos of the target to stylize the generated content with the expected expressions and mannerisms. This will enable a much more automatic process of creating believable deepfakes. Finally, a new trend is real-time deepfakes. Works such as (Jamaludin et al., 2019; Nirkin et al., 2019) have achieved real-time deepfakes at 30fps. Although real-time deepfakes are an enabler for phishing attacks, the realism is not quite there yet. Other limitations include the coherent rendering of hair, teeth, tongues, shadows, and the ability to render the target’s hands (especially when touching the face). Regardless, deepfakes are already very convincing (Rossler et al., 2019) and are improving at a rapid rate. Therefore, it is important that we focus on effective countermeasures.

2. The Deepfake Arms Race

Like any battle in cyber security, there is an arms race between the attacker and defender. In our survey, we observed that the majority deepfake detection algorithms assume a static game with the adversary: They are either focused on identifying a specific artifact, or do not generalize well to new distributions and unseen attacks (Cozzolino et al., 2018). Moreover, based on the recent benchmark of (Li et al., 2019e), we observe that the performance of state-of-the-art detectors are decreasing rapidly as the quality of the deepfakes improve. Concretely, the three most recent benchmark datasets (DFD by Google (Nick Dufour, 2019), DFDC by Facebook (Dolhansky et al., 2019), and Celeb-DF by (Li et al., 2019e)) were released within one month of each other at the end of 2019. However, the deepfake detectors only achieved an AUC of 0.860.86,0.760.76, and 0.660.66 on each of them respectively. Even a false alarm rate of 0.0010.001 is far too low considering the millions of images published online daily.

Evading Artifact-based Detectors. To evade an artifact-based detector, the adversary only needs to mitigate a single flaw to evade detection. For example, GG can generate the biological signals monitored by (Li et al., 2018; Ciftci and Demir, 2019) by adding a discriminator which monitors these signals. To avoid anomalies in extensive the neuron activation (Wang et al., 2019b), the adversary can add a loss which minimizes neuron coverage. Methods which detect abnormal poses and mannerisms (Agarwal et al., 2019) can be evaded by reenacting the entire head and by learning the mannerisms from the same databases. Models which identify blurred content (Mo et al., 2018) are affected by noise and sharpening GANs (Jalalifar et al., 2018; Kim et al., 2016), and models which search for the boundary where the face was blended in (Li et al., 2019b; Agarwal et al., 2017; Zhang et al., 2017; Akhtar and Dasgupta, [n.d.]; Mo et al., 2018; Durall et al., 2019) do not work on deepfakes passed through refiner networks, which use in-painting, or those which output full frames (e.g., (Nagano et al., 2018; Kim et al., 2018; Zhou et al., 2019b; Natsume et al., 2018a; Nirkin et al., 2019; Li et al., 2019a; Liu et al., 2019c; Zablotskaia et al., 2019)). Finally, solutions which search for forensic evidence (Yu et al., 2019a; Marra et al., 2019; Koopman et al., 2018) can be evaded (or at least raise the false alarm rate) by passing xgx_{g} through filters, or by performing physical replication or compression.

Evading Deep Learning Classifiers. There are a number of detection methods which apply deep learning directly to the task of deepfake detection (e.g., (Afchar et al., 2018; Do et al., 2018; Tariq et al., 2018; Ding et al., 2019; Fernando et al., 2019)). However, an adversary can use adversarial machine learning to evade detection by adding small perturbations to xgx_{g}. Advances in adversarial machine learning has shown that these attacks transfer across multiple models regardless of the training data used (Papernot et al., 2016). Recent works have shown how these attacks not only work on deepfakes classifiers (Neekhara et al., 2020) but also work with no knowledge of the classifier or it’s training set (Carlini and Farid, 2020).

Moving Forward. Nevertheless, deepfakes are still imperfect, and these methods offer a modest defense for the time being. Furthermore, these works play an important role in understanding the current limitations of deepfakes, and raise the difficulty threshold for malicious users. At some point, it may become too time-consuming and resource-intensive a common attacker to create a good-enough fake to evade detection. However, we argue that solely relying on the development of content-based countermeasures is not sustainable and may lead to a reactive arms-race. Therefore, we advocate for more out-of-band approaches for detecting a preventing deepfakes. For example, the establishment of content provenance and authenticity frameworks for online videos (Fraga-Lamas and Fernandez-Carames, 2019; et al., 2019a; Hasan and Salah, 2019), and proactive defenses such as the use of adversarial machine learning to protect content from tampering (Li et al., 2019f).

3. Deepfakes in other Domains

In this survey, we put a focus on human reenactment and replacement attacks; the type of deepfakes which has made the largest impact so far (Hall, 2018; Antinori, 2019). However, deepfakes extend beyond human visuals, and have spread many other domains. In healthcare, the authors of (Mirsky et al., 2019) showed how deepfakes can be used to inject tor remove medical evidence in CT and MRI scan for insurance fraud, disruption, and physical harm. In (Jia et al., 2018) it was shown how one’s voice can be cloned with only five seconds of audio, and in Sept. 2019 a CEO was scammed out of $250K via a voice clone deepfake (Demiani, 2019). The authors of (Bontrager et al., 2018) have shown how deep learning can generate realistic human fingerprints that can unlock multiple users’ devices. In (Schreyer et al., 2019) it was shown how deepfakes can be applied to financial records to evade the detection of auditors. Finally, it has been shown how deepfakes of news articles can be generated (Zellers et al., 2019) and that deepfake tweets exist as well (Fagni et al., 2020).

These examples demonstrate that deepfakes are not just attack tools for misinformation, defamation, and propaganda, but also sabotage, fraud, scams, obstruction of justice, and potentially many more.

4. What’s on the Horizon

We believe that in the coming years, we will see more deepfakes being weaponized for monetization. The technology has proven itself in humiliation, misinformation, and defamtion attacks. Moreover, the tools are becoming more practical (dee, 2017) and efficient (Jia et al., 2018). Therefore, is seems natural that malicious users will find ways to use the technology for a profit. As a result, we expect to see an increase in deepfake phishing attacks and scams targeting both companies and individuals.

As the technology matures, real-time deepfakes will become increasingly realistic. Therefore, we can expect that the technology will be used by hacking groups to perform reconnaissance as part of an APT, and by state actors to perform espionage and sabotage by reenacting of officials or family members.

To keep ahead of the game, we must be proactive and consider the adversary’s next step, not just the weaknesses of the current attacks. We suggest that more work be done on evaluating the theoretical limits of these attacks. For example, by finding a bound on a model’s delay can help detect real-time attacks such as (Jia et al., 2018), and determining the limits of GANs like (Agarwal and Varshney, 2019) can help us devise the appropriate strategies. As mentioned earlier, we recommend further research on solutions which do not require analyzing the content itself. Moreover, we believe it would be beneficial for future works to explore the weaknesses and limitations of current deepfakes detectors. By identifying and understanding these vulnerabilities, researchers will be able to develop stronger countermeasures.

Conclusion

Not all deepfakes are malicious. However, because the technology makes it so easy to create believable media, malicious users are exploiting it to perform attacks. These attacks are targeting individuals and causing psychological, political, monetary, and physical harm. As time goes on, we expect to see these malicious deepfakes spread to many other modalities and industries.

In this survey we focused on reenactment and replacement deepfakes of humans. We provided a deep review of how these technologies work, the differences between their architectures, and what is being done to detect them. We hope this information will be helpful to the community in understanding and preventing malicious deepfakes.

References