Age Progression/Regression by Conditional Adversarial Autoencoder

Zhifei Zhang, Yang Song, Hairong Qi

Introduction

Face age progression (i.e., prediction of future looks) and regression (i.e., estimation of previous looks), also referred to as face aging and rejuvenation, aims to render face images with or without the “aging” effect but still preserve personalized features of the face (i.e., personality). It has tremendous impact to a wide-range of applications, e.g., face prediction of wanted/missing person, age-invariant verification, entertainment, etc. The area has been attracting a lot of research interests despite the extreme challenge in the problem itself. Most of the challenges come from the rigid requirement to the training and testing datasets, as well as the large variation presented in the face image in terms of expression, pose, resolution, illumination, and occlusion. The rigid requirement on the dataset refers to the fact that most existing works require the availability of “paired” samples, i.e., face images of the same person at different ages, and some even require paired samples over a long range of age span, which is very difficult to collect. For example, the largest aging dataset “Morph” only captured images with an average time span of 164 days for each individual. In addition, existing works also require the query image to be labeled with the true age, which can be inconvenient from time to time. Given the training data, existing works normally divide them into different age groups and learn a transformation between the groups, therefore, the query image has to be labeled in order to correctly position the image.

Although age progression and regression are equally important, most existing works focus on age progression. Very few works can achieve good performance of face rejuvenating, especially for rendering baby face from an adult because they are mainly surface-based modeling which simply remove the texture from a given image . On the other hand, researchers have made great progress on age progression. For example, the physical model-based methods parametrically model biological facial change with age, e.g., muscle, wrinkle, skin, etc. However, they suffer from complex modeling, the requirement of sufficient dataset to cover long time span, and are computationally expensive; the prototype-based methods tend to divide training data into different age groups and learn a transformation between groups. However, some can preserve personality but induce severe ghosting artifacts, others smooth out the ghosting effect but lose personality, while most relaxed the requirement of paired images over long time span, and the aging pattern can be learned between two adjacent age groups. Nonetheless, they still need paired samples over short time span.

In this paper, we investigate the age progression/regression problem from the perspective of generative modeling. The rapid development of generative adversarial networks (GANs) has shown impressive results in face image generation . In this paper, we assume that the face images lie on a high-dimensional manifold as shown in Fig.1. Given a query face, we could find the corresponding point (face) on the manifold. Stepping along the direction of age changing, we will obtain the face images of different ages while preserving personality. We propose a conditional adversarial autoencoder (CAAE)Bitbucket: https://bitbucket.org/aicip/face-aging-caae Github: https://zzutk.github.io/Face-Aging-CAAE network to learn the face manifold. By controlling the age attribute, it will be flexible to achieve age progression and regression at the same time. Because it is difficult to directly manipulate on the high-dimensional manifold, the face is first mapped to a latent vector through a convolutional encoder, and then the vector is projected to the face manifold conditional on age through a deconvolutional generator. Two adversarial networks are imposed on the encoder and generator, respectively, forcing to generate more photo-realistic faces.

The benefit of the proposed CAAE can be summarized from four aspects. First, the novel network architecture achieves both age progression and regression while generating photo-realistic face images. Second, we deviate from the popular group-based learning, thus not requiring paired samples in the training data or labeled face in the test data, making the proposed framework much more flexible and general. Third, the disentanglement of age and personality in the latent vector space helps preserving personality while avoiding the ghosting artifacts. Finally, CAAE is robust against variations in pose, expression, and occlusion.

Related Work

In recent years, the study on face age progression has been very popular, with approaches mainly falling into two categories, physical model-based and prototype-based. Physical model-based methods model the biological pattern and physical mechanisms of aging, e.g., the muscles , wrinkle , facial structure etc. through either parametric or non-parametric learning. However, in order to better model the subtle aging mechanism, it will require a large face dataset with long age span (e.g., from 0 to 80 years old) of each individual, which is very difficult to collect. In addition, physical modeling-based approaches are computationally expensive.

On the other hand, prototype-based approaches often divide faces into groups by age, e.g., the average face of each group, as its prototype. Then, the difference between prototypes from two age groups is considered the aging pattern. However, the aged face generated from averaged prototype may lose the personality (e.g., wrinkles). To preserve the personality, proposed a dictionary learning based method — age pattern of each age group is learned into the corresponding sub-dictionary. A given face will be decomposed into two parts: age pattern and personal pattern. The age pattern was transited to the target age pattern through the sub-dictionaries, and then the aged face is generated by synthesizing the personal pattern and target age pattern. However, this approach presents serious ghosting artifacts. The deep learning-based method represents the state-of-the-art, where RNN is applied on the coefficients of eigenfaces for age pattern transition. All prototype-based approaches perform the group-based learning which requires the true age of testing faces to localize the transition state which might not be convenient. In addition, these approaches only provide age progression from younger face to older ones. To achieve flexible bidirectional age changes, it may need to retrain the model inversely.

Face age regression, which predicts the rejuvenating results, is comparatively more challenging. Most age regression works so far are physical model-based, where the textures are simply removed based on the learned transformation over facial surfaces. Therefore, they cannot achieve photo-realistic results for baby face predictions.

2 Generative Adversarial Network

Generating realistically appealing images is still challenging and has not achieved much success until the rapid advancement of the generative adversarial network (GAN). The original GAN work introduced a novel framework for training generative models. It simultaneously trains two models: 1) the generative model GG captures the distribution of training samples and learns to generate new samples imitating the training, and 2) the discriminative model DD discriminates the generated samples from the training. GG and DD compete with each other using a min-max game as Eq. 1, where zz denotes a vector randomly sampled from certain distribution p(z)p(\mathbf{z}) (e.g., Gaussian or uniform), and the data distribution is pdata(x)p_{data}(\mathbf{x}), i.e., the training data x∼pdata(x)x\sim p_{data}(\mathbf{x}).

The two parts, GG and DD, are trained alternatively.

One of the biggest issues of GAN is that the training process is unstable, and the generated images are often noisy and incomprehensible. During the last two years, several approaches have been proposed to improve the original GAN from different perspectives. For example, DCGAN adopted deconvolutional and convolutional neural networks to implement GG and DD, respectively. It also provided empirical instruction on how to build a stable GAN, e.g., replacing the pooling by strides convolution and using batch normalization. CGAN modified GAN from unsupervised learning into semi-supervised learning by feeding the conditional variable (e.g., the class label) into the data. The low resolution of the generated image is another common drawback of GAN. extended GAN into sequential or pyramid GANs to handle this problem, where the image is generated step by step, and each step utilizes the information from the previous step to further improve the image quality. Some GAN-related works have shown visually impressive results of randomly drawing face images . However, GAN generates images from random noise, thus the output image cannot be controlled. This is undesirable in age progression and regression, where we have to ensure the output face looks like the same person as queried.

Traversing on the Manifold

We assume the face images lie on a high-dimensional manifold, on which traversing along certain direction could achieve age progression/regression while preserving the personality. This assumption will be demonstrated experimentally in Sec. 4.2. However, modeling the high-dimensional manifold is complicated, and it is difficult to directly manipulate (traversing) on the manifold. Therefore, we will learn a mapping between the manifold and a lower-dimensional space, referred to as the latent space, which is easier to manipulate. As illustrated in Fig. 2, faces x1x_{1} and x2x_{2} are mapped to the latent space by EE (i.e., an encoder), which extracts the personal features z1z_{1} and z2z_{2}, respectively. Concatenating with the age labels l1l_{1} and l2l_{2}, two points are generated in the latent space, namely [z1,l1][z_{1},l_{1}] and [z2,l2][z_{2},l_{2}]. Note that the personality zz and age ll are disentangled in the latent space, thus we could simply modify age while preserving the personality. Starting from the red rectangular point [z2,l2][z_{2},l_{2}] (corresponding to x2x_{2}) and evenly stepping bidirectionally along the age axis (as shown by the solid red arrows), we could obtain a series of new points (red circle points). Through another mapping GG (i.e. a generator), those points are mapped to the manifold M\mathcal{M} – generating a series of face images, which will present the age progression/regression of x2x_{2}. By the same token, the green points and arrows demonstrate the age progressing/regression of x1x_{1} based on the learned manifold and the mappings. If we move the point along the dotted arrow in the latent space, both personality and age will be changed as reflected on M\mathcal{M}. We will learn the mappings EE and GG to ensure the generated faces lie on the manifold, which indicates that the generated faces are realistic and plausible for a given age.

Approach

In this section, we first present the pipeline of the proposed conditional adversarial autoencoder (CAAE) network (Sec. 4.1) that learns the face manifold (Sec. 4.2). The CAAE incorporates two discriminator networks, which are the key to generating more realistic faces. Sections 4.3 and 4.4 demonstrate their effectiveness, respectively. Finally, Section 4.5 discusses the difference of the proposed CAAE from other generative models.

2 Objective Function

Simultaneously, the uniform distribution is imposed on zz through DzD_{z} – the discriminator on zz. We denote the distribution of the training data as pdata(x)p_{data}(\mathbf{x}), then the distribution of zz is q(z∣x)q(\mathbf{z}|\mathbf{x}). Assuming p(z)p(\mathbf{z}) is a prior distribution, and z∗∼p(z)z^{*}\sim p(\mathbf{z}) denotes the random sampling process from p(z)p(\mathbf{z}). A min-max objective function can be used to train EE and DzD_{z},

By the same token, the discriminator on face image, DimgD_{img}, and GG with condition ll can be trained by

where TV(⋅)TV(\cdot) denotes the total variation which is effective in removing the ghosting artifacts. The coefficients λ\lambda and γ\gamma balance the smoothness and high resolution.

Note that the age label is resized and concatenated to the first convolutional layer of DimgD_{img} to make it discriminative on both age and human face. Sequentially updating the network by Eqs. 2, 3, and 4, we could finally learn the manifold M\mathcal{M} as illustrated in Fig. 4.

3 Discriminator on 𝒛𝒛\boldsymbol{z}

The discriminator on zz, denoted by DzD_{z}, imposes a prior distribution (e.g., uniform distribution) on zz. Specifically, DzD_{z} aims to discriminate the zz generated by encoder EE. Simultaneously, EE will be trained to generate zz that could fool DzD_{z}. Such adversarial process forces the distribution of the generated zz to gradually approach the prior. We use uniform distribution as the prior, forcing zz to evenly populate the latent space with no apparent “holes”. As shown in Fig. 5, the generated zz’s (depicted by blue dots in a 2-D space) present uniform distribution under the regularization of DzD_{z}, while the distribution of zz exhibits a “hole” without the application of DzD_{z}.

Exhibition of the “hole” indicates that face images generated by interpolating between arbitrary zz’s may not lie on the face manifold – generating unrealistic faces. For example, given two faces x1x_{1} and x2x_{2} as shown in Fig. 5, we obtain the corresponding z1z_{1} and z2z_{2} by EE under the conditions with and without DzD_{z}, respectively. Interpolating between z1z_{1} and z2z_{2} (dotted arrows in Fig. 5), the generated faces are expected to show realistic and smooth morphing from x1x_{1} to x2x_{2} (bottom of Fig. 5). However, the morphing without DzD_{z} actually presents distorted (unrealistic) faces in the middle (indicated by dashed box), which corresponds to the interpolated zz’s passing through the “hole”.

4 Discriminator on Face Images

Inheriting the similar principle of GAN, the discriminator DimgD_{img} on face images forces the generator to yield more realistic faces. In addition, the age label is imposed on DimgD_{img} to make it discriminative against unnatural faces conditional on age. Although minimizing the distance between the input and output images as expressed in Eq. 2 forces the output face to be close to the real ones, Eq. 2 does not ensure the framework to generate plausible faces from those unsampled faces. For example, given a face that is unseen during training and a random age label, the pixel-wise loss could only make the framework generate a face close to the trained ones in a manner of interpolation, causing the generated face to be very blurred. The DimgD_{img} will discriminate the generated faces from real ones in aspects of reality, age, resolution, etc. Fig. 6 demonstrates the effect of DimgD_{img}.

Comparing the generated faces with and without DimgD_{img}, it is obvious that DimgD_{img} assists the framework to generate more realistic faces. The outputs without DimgD_{img} could also present aging but the effect is not as obvious as that with DimgD_{img} because DimgD_{img} enhances the texture especially for older faces.

5 Differences from Other Generative Networks

In this section, we comment on the similarity and difference of the proposed CAAE with other generative networks, including GAN , variational autoencoder (VAE) , and adversarial autoencoder (AAE) .

VAE vs. GAN: VAE uses a recognition network to predict the posterior distribution over the latent variables, while GAN uses an adversarial training procedure to directly shape the output distribution of the network via back-propagation . Because VAE follows an encoding-decoding scheme, we can directly compare the generated images to the inputs, which is not possible when using a GAN. A downside of VAE is that it uses mean squared error instead of an adversarial network in image generation, so it tends to produce more blurry images .

AAE vs. GAN and VAE: AAE can be treated as the combination of GAN and VAE, which maintains the autoencoder network like VAE but replaces the KL-divergence loss with an adversarial network like in GAN. Instead of generating images from random noise as in GAN, AAE utilizes the encoder part to learn the latent variables approximated on certain prior, making the style of generated images controllable. In addition, AAE better captures the data manifold compared to VAE.

CAAE vs. AAE: The proposed CAAE is more similar to AAE. The main difference from AAE is that the proposed CAAE imposes discriminators on the encoder and generator, respectively. The discriminator on encoder guarantees smooth transition in the latent space, and the discriminator on generator assists to generate photo-realistic face images. Therefore, CAAE would generate higher quality images than AAE as discussed in Sec. 4.4.

Experimental Evaluation

In the section, we will first clarify the process of data collection (Sec. 5.1) and implementation of the proposed CAAE (Sec. 5.2). Then, both qualitative and quantitative comparisons with prior works and ground truth are performed in Sec. 5.3. Finally, the tolerance to occlusion and variation in pose and expression is illustrated in Sec. 5.4 .

We first collect face images from the Morph dataset and the CACD dataset. The Morph dataset is the largest with multiple ages of each individual, including 55,000 images of 13,000 subjects from 16 to 77 years old. The CACD dataset contains 13,446 images of 2,000 subjects. Because both datasets have limited images from newborn or very old faces, we crawl images from Bing and Google search engines based on the keywords, e.g., baby, boy, teenager, 15 years old, etc. Because the proposed approach does not require multiple faces from the same subject, we simply randomly choose around 3,000 images from the Morph and CACD dataset and crawl 7,670 images from the website. The age and gender of the crawled faces are estimated based on the image caption or the result from age estimator . We divide the age into ten categories, i.e., 0–5, 6–10, 11–15, 16–20, 21–30, 31–40, 41–50, 51–60, 61–70, and 71–80. Therefore, we can use a one-hot vector of ten elements to indicate the age of each face during training. The final dataset consists of 10,670 face images with a uniform distribution on gender and age. We use the face detection algorithm with 68 landmarks to crop out and align the faces, making the training more attainable.

2 Implementation of CAAE

We construct the network according to Fig. 3 with kernel size of 5×55\times 5. The pixel values of the input images are normalized to $,andtheoutputof, and the output ofE(i.e.,(i.e.,z)isalsorestrictedto) is also restricted tobythehyperbolictangentactivationfunction.Then,thedesiredagelabel,theone−hotvector,isconcatenatedtoby the hyperbolic tangent activation function. Then, the desired age label, the one-hot vector, is concatenated toz,constructingtheinputof, constructing the input ofG.Tomakefairconcatenation,theelementsoflabelisalsoconfinedto<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><moseparator="true">,</mo><mi>w</mi><mi>h</mi><mi>e</mi><mi>r</mi><mi>e</mi><mo>−</mo><mn>1</mn><mi>c</mi><mi>o</mi><mi>r</mi><mi>r</mi><mi>e</mi><mi>s</mi><mi>p</mi><mi>o</mi><mi>n</mi><mi>d</mi><mi>s</mi><mi>t</mi><mi>o</mi><mn>0.</mn><mi>F</mi><mi>i</mi><mi>n</mi><mi>a</mi><mi>l</mi><mi>l</mi><mi>y</mi><moseparator="true">,</mo><mi>t</mi><mi>h</mi><mi>e</mi><mi>o</mi><mi>u</mi><mi>t</mi><mi>p</mi><mi>u</mi><mi>t</mi><mi>i</mi><mi>s</mi><mi>a</mi><mi>l</mi><mi>s</mi><mi>o</mi><mi>i</mi><mi>n</mi><mi>r</mi><mi>a</mi><mi>n</mi><mi>g</mi><mi>e</mi></mrow><annotationencoding="application/x−tex">,where−1correspondsto0.Finally,theoutputisalsoinrange</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.8889em;vertical−align:−0.1944em;"></span><spanclass="mpunct">,</span><spanclass="mspace"style="margin−right:0.1667em;"></span><spanclass="mordmathnormal"style="margin−right:0.0269em;">w</span><spanclass="mordmathnormal">h</span><spanclass="mordmathnormal"style="margin−right:0.0278em;">er</span><spanclass="mordmathnormal">e</span><spanclass="mspace"style="margin−right:0.2222em;"></span><spanclass="mbin">−</span><spanclass="mspace"style="margin−right:0.2222em;"></span></span><spanclass="base"><spanclass="strut"style="height:0.8889em;vertical−align:−0.1944em;"></span><spanclass="mord">1</span><spanclass="mordmathnormal"style="margin−right:0.0278em;">cor</span><spanclass="mordmathnormal"style="margin−right:0.0278em;">r</span><spanclass="mordmathnormal">es</span><spanclass="mordmathnormal">p</span><spanclass="mordmathnormal">o</span><spanclass="mordmathnormal">n</span><spanclass="mordmathnormal">d</span><spanclass="mordmathnormal">s</span><spanclass="mordmathnormal">t</span><spanclass="mordmathnormal">o</span><spanclass="mord">0.</span><spanclass="mordmathnormal"style="margin−right:0.1389em;">F</span><spanclass="mordmathnormal">ina</span><spanclass="mordmathnormal"style="margin−right:0.0197em;">l</span><spanclass="mordmathnormal"style="margin−right:0.0197em;">l</span><spanclass="mordmathnormal"style="margin−right:0.0359em;">y</span><spanclass="mpunct">,</span><spanclass="mspace"style="margin−right:0.1667em;"></span><spanclass="mordmathnormal">t</span><spanclass="mordmathnormal">h</span><spanclass="mordmathnormal">eo</span><spanclass="mordmathnormal">u</span><spanclass="mordmathnormal">tp</span><spanclass="mordmathnormal">u</span><spanclass="mordmathnormal">t</span><spanclass="mordmathnormal">i</span><spanclass="mordmathnormal">s</span><spanclass="mordmathnormal">a</span><spanclass="mordmathnormal"style="margin−right:0.0197em;">l</span><spanclass="mordmathnormal">so</span><spanclass="mordmathnormal">in</span><spanclass="mordmathnormal"style="margin−right:0.0278em;">r</span><spanclass="mordmathnormal">an</span><spanclass="mordmathnormal"style="margin−right:0.0359em;">g</span><spanclass="mordmathnormal">e</span></span></span></span></span>throughthehyperbolictangentfunction.Normalizingtheinputmaymakethetrainingprocessconvergefaster.Notethatwewillnotusethebatchnormalizationfor. To make fair concatenation, the elements of label is also confined to <span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo separator="true">,</mo><mi>w</mi><mi>h</mi><mi>e</mi><mi>r</mi><mi>e</mi><mo>−</mo><mn>1</mn><mi>c</mi><mi>o</mi><mi>r</mi><mi>r</mi><mi>e</mi><mi>s</mi><mi>p</mi><mi>o</mi><mi>n</mi><mi>d</mi><mi>s</mi><mi>t</mi><mi>o</mi><mn>0.</mn><mi>F</mi><mi>i</mi><mi>n</mi><mi>a</mi><mi>l</mi><mi>l</mi><mi>y</mi><mo separator="true">,</mo><mi>t</mi><mi>h</mi><mi>e</mi><mi>o</mi><mi>u</mi><mi>t</mi><mi>p</mi><mi>u</mi><mi>t</mi><mi>i</mi><mi>s</mi><mi>a</mi><mi>l</mi><mi>s</mi><mi>o</mi><mi>i</mi><mi>n</mi><mi>r</mi><mi>a</mi><mi>n</mi><mi>g</mi><mi>e</mi></mrow><annotation encoding="application/x-tex">, where -1 corresponds to 0. Finally, the output is also in range</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal" style="margin-right:0.0269em;">w</span><span class="mord mathnormal">h</span><span class="mord mathnormal" style="margin-right:0.0278em;">er</span><span class="mord mathnormal">e</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord">1</span><span class="mord mathnormal" style="margin-right:0.0278em;">cor</span><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="mord mathnormal">es</span><span class="mord mathnormal">p</span><span class="mord mathnormal">o</span><span class="mord mathnormal">n</span><span class="mord mathnormal">d</span><span class="mord mathnormal">s</span><span class="mord mathnormal">t</span><span class="mord mathnormal">o</span><span class="mord">0.</span><span class="mord mathnormal" style="margin-right:0.1389em;">F</span><span class="mord mathnormal">ina</span><span class="mord mathnormal" style="margin-right:0.0197em;">l</span><span class="mord mathnormal" style="margin-right:0.0197em;">l</span><span class="mord mathnormal" style="margin-right:0.0359em;">y</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal">t</span><span class="mord mathnormal">h</span><span class="mord mathnormal">eo</span><span class="mord mathnormal">u</span><span class="mord mathnormal">tp</span><span class="mord mathnormal">u</span><span class="mord mathnormal">t</span><span class="mord mathnormal">i</span><span class="mord mathnormal">s</span><span class="mord mathnormal">a</span><span class="mord mathnormal" style="margin-right:0.0197em;">l</span><span class="mord mathnormal">so</span><span class="mord mathnormal">in</span><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="mord mathnormal">an</span><span class="mord mathnormal" style="margin-right:0.0359em;">g</span><span class="mord mathnormal">e</span></span></span></span></span> through the hyperbolic tangent function. Normalizing the input may make the training process converge faster. Note that we will not use the batch normalization forEandandGbecauseitblurspersonalfeaturesandmakesoutputfacesdriftfarawayfrominputsintesting.However,thebatchnormalizationwillmaketheframeworkmorestableifitisappliedonbecause it blurs personal features and makes output faces drift far away from inputs in testing. However, the batch normalization will make the framework more stable if it is applied onD_{img}.Allintermediatelayersofeachblock(i.e.,. All intermediate layers of each block (i.e.,E,,G,,D_{z},and, andD_{img}$) use the ReLU activation function.

In training, λ=100\lambda=100, γ=10\gamma=10, and the four blocks are updated alternatively with a mini-batch size of 100 through the stochastic gradient descent solver, ADAM (α=0.0002\alpha=0.0002, β1=0.5\beta_{1}=0.5). Face and age pairs are fed to the network. After about 50 epochs, plausible generated faces can be obtained. During testing, only EE and GG are active. Given an input face without true age label, EE maps the image to zz. Concatenating an arbitrary age label to zz, GG will generate a photo-realistic face corresponding to the age and personality.

3 Qualitative and Quantitative Comparison

To evaluate that the proposed CAAE can generate more photo-realistic results, we compare ours with the ground truth and the best results from prior works , respectively. We choose FGNET as the testing dataset, which has 1002 images of 82 subjects aging from 0 to 69.

Comparison with ground truth: In order to verify whether the personality has been preserved by the proposed CAAE, we qualitatively and quantitatively compare the generated faces with the ground truth. The qualitative comparison is shown in Fig. 8, which shows appealing similarity. To quantitatively evaluate the performance, we pair the generated faces with the ground truth whose age gap is larger than 20 years. There are 856 pairs in total. We design a survey to compare the similarity where 63 volunteers participate. Each volunteer is presented with three images, an original image X, a generated image A, and the corresponding ground truth image B under the same group. They are asked whether the generated image A looks similar to the ground truth B; or not sure. We ask the volunteers to randomly choose 45 questions and leave the rest blank. We receive 3208 votes in total, with 48.38% indicating that the generated image A is the same person as the ground truth, 29.58% indicating they are not, and 22.04% not sure. The voting results demonstrate that we can effectively generate photo-realistic faces under different ages while preserving their personality.

Comparison with prior work: We compare the performance of our method with some prior works , for face age progression and Face Transformer (FT) for face age regression. To demonstrate the advantages of CAAE, we use the same input images collected from those prior works and perform long age span progression. To compare with prior works, we cite their results as shown in Fig. 7. We also compare with age regression works using the FT demo as shown in Fig. 9. Our results obviously show higher fidelity, demonstrating the capability of CAAE in achieving smooth face aging and rejuvenation. CAAE better preserves the personality even with a long age span. In addition, our results provide richer texture (e.g., wrinkle for old faces), making old faces look more realistic. Another survey is conducted to statistically evaluate the performance as compared with prior works, where for each testing image, the volunteer is to select the better result from CAAE or prior works, or hard to tell. We collect 235 paired images of 79 subjects from previous works . We receive 47 responses and 1508 votes in total with 52.77% indicating CAAE is better, 28.99% indicating the prior work is better, and 18.24% indicating they are equal. This result further verifies the superior performance of the proposed CAAE.

4 Tolerance to Pose, Expression, and Occlusion

As mentioned above, the input images have large variation in pose, expression, and occlusion. To demonstrate the robustness of CAAE, we choose the faces with expression variation, non-frontal pose, and occlusion, respectively, as shown in Fig. 10. It is worth noting that the previous works often apply face normalization to alleviate from the variation of pose and expression but they may still suffer from the occlusion issue. In contrast, the proposed CAAE obtains the generated faces without the need to remove these variations, paving the way to robust performance in real applications.

Discussion and Future Works

In this paper, we proposed a novel conditional adversarial autoencoder (CAAE), which first achieves face age progression and regression in a holistic framework. We deviated from the conventional routine of group-based training by learning a manifold, making the aging progression/regression more flexible and manipulatable — from an arbitrary query face without knowing its true age, we can freely produce faces at different ages, while at the same time preserving the personality. We demonstrated that with two discriminators imposed on the generator and encoder, respectively, the framework generates more photo-realistic faces. Flexibility, effectiveness, and robustness of CAAE have been demonstrated through extensive evaluation.

The proposed framework has great potential to serve as a general framework for face-age related tasks. More specifically, we trained four sub-networks, i.e., EE, GG, DzD_{z}, and DimgD_{img}, but only EE and GG are utilized in the testing stage. The DimgD_{img} is trained conditional on age. Therefore, it is able to tell whether the given face corresponds to a certain age, which is exactly the task of age estimation. For the encoder EE, it maps faces to a latent vector (face feature), which preserves the personality regardless of age. Therefore, EE could be considered a candidate for cross-age recognition. The proposed framework could be easily applied to other image generation tasks, where the characteristics of the generated image can be controlled by the conditional label. In the future, we would extend current work to be a general framework, simultaneously achieving age progressing (EE and GG), cross-age recognition (EE), face morphing (GG), and age estimation (DimgD_{img}).

References