SharinGAN: Combining Synthetic and Real Data for Unsupervised Geometry Estimation

Koutilya PNVR, Hao Zhou, David Jacobs

Introduction

Understanding geometry from images is a fundamental problem in computer vision. It has many important applications. For instance, Monocular Depth Estimation (MDE) is important for synthetic object insertion in computer graphics , grasping in robotics and safety in self-driving cars. Face Normal Estimation can help in face image editing applications such as relighting . However, it is extremely hard to annotate real data for these regression tasks. Synthetic data and their ground truth labels, on the other hand, are easy to generate and are often used to compensate for the lack of labels in real data. Deep models trained on synthetic data, unfortunately, usually perform poorly on real data due to the domain gap between synthetic and real distributions. To deal with this problem, several research studies have proposed unsupervised domain adaptation methods to take advantage of synthetic data by mapping it into the real domain or vice versa, either at the feature level or image level. However, mapping examples from one domain to another domain itself is a challenging problem that can limit performance.

We observe that finding such a mapping solves an unnecessarily difficult problem. To train a regressor that applies to both real and synthetic domains, it is only necessary that we map both to a new representation that contains the task-relevant information present in both domains, in a common form. The mapping need not alter properties of the original domain that are irrelevant to the task since the regressor will learn to ignore them regardless.

To see this, we consider a simplified model of our problem. We suppose that real and synthetic images are formed by two components: domain agnostic (which has semantic information shared across synthetic and real, and is denoted as II) and domain specific. We further assume that domain specific information has two sub-components: domain specific information unrelated to the primary task (denoted as δs′\delta_{s}^{\prime} and δr′\delta_{r}^{\prime} for synthetic and real images respectively) and domain specific information related to the primary task (δs\delta_{s}, δr\delta_{r}). So real and synthetic images can be represented as: xr=f(I,δr,δr′)x_{r}=f(I,\delta_{r},\delta_{r}^{\prime}) and xs=f(I,δs,δs′)x_{s}=f(I,\delta_{s},\delta_{s}^{\prime}) respectively.

We believe the domain gap between {δs\{\delta_{s} and δr}\delta_{r}\} can affect the training of the primary network, which learns to expect information that is not always present. The domain gap between {δs′\{\delta_{s}^{\prime} and δr′}\delta_{r}^{\prime}\}, on the other hand, can be bypassed by the primary network since it does not hold information needed for the primary task. For example, in real face images, information such as the color and texture of the hair is unrelated to the task of estimating face normals but is discriminative enough to distinguish real from synthetic faces. This can be regarded as domain specific information unrelated to the primary task i.e., δr′\delta_{r}^{\prime}. On the other hand, shadows in the real and synthetic images, due to the limitations of the rendering engine, may have different appearances but may contain depth cues that are related to the primary task of MDE in both domains. The simplest strategy, then, for combining real and synthetic data is to map δs\delta_{s} and δr\delta_{r} to a shared representation, δsh\delta_{sh}, while not modifying δs′\delta^{\prime}_{s} and δr′\delta^{\prime}_{r} as shown in Figure 1.

Recent research studies show that a shared network for synthetic and real data can help reduce the discrepancy between images in different domains. For instance, achieved state-of-the-art results in face normal estimation by training a unified network for real and synthetic data. learned the joint distribution of multiple domain images by enforcing a weight-sharing constraint for different generative networks. Inspired by these research studies, we define a unified mapping function GG, which is called SharinGAN, to reduce the domain gap between real and synthetic images.

Different from existing research studies, our GG is trained so that minimum domain specific information is removed. This is achieved by pre-training GG as an auto-encoder on real and synthetic data, i.e., initializing GG as an identity function. Then GG is trained end-to-end with reconstruction loss in an adversarial framework, along with a network that solves the primary task, further pushing GG to map information relevant to the task to a shared domain.

As a result, a successfully trained GG will learn to reduce the domain gap existing in δs\delta_{s} and δr\delta_{r}, mapping them into a shared domain δsh\delta_{sh}. GG will leave II unchanged. δs′\delta_{s}^{\prime} and δr′\delta_{r}^{\prime} can be left relatively unchanged when it is difficult to map them to a common representation. Mathematically, G(xs)=f(I,δsh,δs′)G(x_{s})=f(I,\delta_{sh},\delta_{s}^{\prime}) and G(xr)=f(I,δsh,δr′)G(x_{r})=f(I,\delta_{sh},\delta_{r}^{\prime}). If successful, GG will map synthetic and real images to images that may look quite different to the eye, but the primary task network will extract the same information from both.

We apply our method to unsupervised monocular depth estimation using virtual KITTI (vKITTI) and KITTI as synthetic and real datasets respectively. Our method reduces the absolute error in the KITTI eigen test split and the test set of Make3D by 23.77%23.77\% and 6.45%6.45\% respectively compared with the state-of-the-art method . Additionally, our proposed method improves over SfSNet on face normal estimation. It yields an accuracy boost of nearly 4.3%4.3\% for normal prediction within 20∘20^{\circ} (Acc<20∘)(Acc<20^{\circ}) of ground truth on the Photoface dataset . Our code is available at https://github.com/koutilya40192/SharinGAN.

Related Work

Monocular Depth Estimation has long been an active area in computer vision. Because this problem is ill-posed, learning-based methods have predominated in recent years. Many early learning works applied Markov Random Fields (MRF) to infer the depth from a single image by modeling the relation between nearby regions . These methods, however, are time-consuming during inference and rely on manually defined features, which have limitations in performance.

More recent studies apply deep Convolutional Neural Networks (CNNs) to monocular depth estimation. Eigen et al. first proposed a multi-scale deep CNN for depth estimation. Following this work, proposed to apply CNNs to estimate depth, surface normal and semantic labels together. combined deep CNNs with a continuous CRF for monocular depth estimation. One major drawback of these supervised learning-based methods is the requirement for a huge amount of annotated data, which is hard to obtain in reality.

With the emergence of large scale, high-quality synthetic data , using synthetic data to train a depth estimator network for real data became popular . The biggest challenge for this task is the large domain gap between synthetic data and real data. proposed to first train a depth prediction network using synthetic data. A style transfer network is then trained to map real images to synthetic images in a cycle consistent manner . proposed to adapt the features of real images to the features of synthetic images by applying adversarial loss on latent features. A content congruent regularization is further proposed to avoid mode collapse. T2Net trained a network that translates synthetic data into real at the image level and further trained a task network in this translated domain. GASDA proposed to train the network by incorporating epipolar geometry constraints for real data along with the ground truth labels for synthetic data. All these methods try to align two domains by transferring one domain to another. Unlike these works, we propose a mapping function GG, also called SharinGAN, to just align the domain specific information that affects the primary task, resulting in a minimum change in the images in both domains. We show that this makes learning the primary task network much easier and can help it focus on the useful information.

Self-supervised learning is another way to avoid collecting ground truth labels for monocular depth estimation. Such methods need monocular videos , stereo pairs , or both for training. Our proposed method is complementary to these self-supervised methods, it does not require this additional data, but can use it when available.

Face Geometry Estimation is a sub-problem of inverse face rendering which is the key for many applications such as face image editing. Conventional face geometry estimation methods are usually based on 3D Morphable Models (3DMM) . Recent studies demonstrate the effectiveness of deep CNNs for solving this problem . Thanks to the 3DMM, generating synthetic face images with ground truth geometry is easy. make use of synthetic face images with ground truth shape to help train a network for predicting face shape using real images. Most of these works initially pre-train the network with synthetic data and then fine-tune it with a mix of real and synthetic data, either using no supervision or weak supervision, overlooking the domain gap between real and synthetic face images. In this work, we show that by reducing the domain gap between real and synthetic data using our proposed method, face geometry can be better estimated.

Domain Adaptation using GANs There are many works that use a GAN framework to perform domain adaptation by mapping one domain into another via a supervised translation. However, most of these show performance on just toy datasets in a classification setting. We attempt to map both synthetic and real domains into a new shared domain that is learned during training and use this to solve complex problems of unsupervised geometry estimation. Moreover, we apply adversarial loss at the image level for our regression task, in contrast to some of the above previous works where domain invariant feature engineering sufficed for classification tasks.

Method

To compensate for the lack of annotations for real data and to train a primary task network on easily available synthetic data, we propose SharinGAN to reduce the domain gap between synthetic and real. We aim to train a primary task network on a shared domain created by SharinGAN, which learns the mapping function G:xr↦xrshG:x_{r}\mapsto x_{r}^{sh} and G:xs↦xsshG:x_{s}\mapsto x_{s}^{sh}, where xk=f(I,δk,δk′);  xksh=f(I,δsh,δk′);  k∈{r,s}x_{k}=f(I,\delta_{k},\delta_{k}^{\prime});\thickspace x_{k}^{sh}=f(I,\delta_{sh},\delta_{k}^{\prime});\thickspace k\in\{r,s\} as shown in Figure 1. GG allows the primary task network to train on a shared space that holds the information needed to do the primary task, making the network more applicable to real data during testing.

To achieve this, an adversarial loss is used to find the shared information, δsh\delta_{sh}. This is done by minimizing the discrepancy in the distributions of xrshx_{r}^{sh} and xsshx_{s}^{sh}. But at the same time, to preserve the domain agnostic information (shared semantic information II), we use reconstruction loss. Now, without a loss from the primary task network, GG might change the images so that they don’t match the labels. To prevent that, we additionally use a primary task loss for both real and synthetic examples to guide the generator. It is important to note that both the translations from synthetic to real and vice versa are equally crucial for this symmetric setup to find a shared space. To facilitate that, we use a form of weak supervision we call virtual supervision. Some possible virtual supervisions include a prior on the input data or a constraint that can narrow the solution space for the primary task network (details discussed in 3.2.2). For synthetic examples, we use the known labels.

Adversarial, Reconstruction and Primary task losses together train the generator and primary task network to align the domain specific information {δs,δr}\{\delta_{s},\delta_{r}\} in both the domains into a shared space δsh\delta_{sh}, preserving everything else.

In this work, we propose to train a generative network which is called SharinGAN, to reduce the domain gap between real and synthetic data so as to help to train the primary network. Figure 2 shows the framework of our proposed method. It contains a generative network GG, a discriminator on image-level DD that embodies the SharinGAN module and a task network TT to perform the primary task. The generative network GG takes either a synthetic image xsx_{s} or real image xrx_{r} as input and transforms it to xsshx_{s}^{sh} or xrshx_{r}^{sh} in an attempt to fool DD. Different from existing works that transfer images in one domain to another , our generative network GG tries to map the domain specific parts δs\delta_{s} and δr\delta_{r} of synthetic and real images to a shared space δsh\delta_{sh}, leaving δs′\delta_{s}^{\prime} and δr′\delta_{r}^{\prime} unchanged. As a result, our transformed synthetic and real images (xsshx_{s}^{sh} and xrshx_{r}^{sh}) have fewer differences from xsx_{s} and xrx_{r}. Our task network TT then takes the transformed images xsshx_{s}^{sh} and xrshx_{r}^{sh} as input and predicts the geometry. The generative network GG and task network TT are trained together in an end-to-end manner.

2 Losses

In this section, we describe the losses we use for the generative and task networks.

We design a single generative network GG for synthetic and real data since sharing weights can help align distributions of different domains . Moreover, existing research studies such as also demonstrate that a unified framework works reasonably well on synthetic and real images. In order to map δs\delta_{s} and δr\delta_{r} to a shared space δsh\delta_{sh}, we apply adversarial loss at the image level. More specifically, we use the Wasserstein discriminator that uses the Earth-Mover’s distance to minimize the discrepancy between the distributions for synthetic and real examples {G(xs),G(xr)}\{G(x_{s}),G(x_{r})\}, i.e.:

DD is a discriminator and GeG_{e} is the encoder part of the generator. Following , to overcome the problem of vanishing or exploding gradients due to the weight clipping proposed in , a gradient penalty term is added for training the discriminator:

Our overall adversarial loss is then defined as:

where λ\lambda is chosen to be 1010 while training the discriminator and while training the generator.

Without any constraints, the adversarial loss may learn to remove all domain specific parts δ\delta and δ′\delta^{\prime} or even some of the domain agnostic part II in order to fool the discriminator. This may lead to loss of geometric information, which can degrade the performance of the primary task network TT. To avoid this, we propose to use the self-regularization loss similar to to force the transformed image to keep as much information as possible:

2.2 Losses for the Task Network

The task network takes transformed synthetic or real images as input and predicts geometric information. Since the ground truth labels for synthetic data are available, we apply a supervised loss using these ground truth labels. For real images, domain specific losses or regularizations are applied as a form of virtual supervision for training according to the task. We apply our proposed SharinGAN to two tasks: monocular depth estimation (MDE) and face normal estimation (FNE). For MDE, we use the combination of depth smoothness and geometric consistency losses used in GASDA as the virtual supervision. For FNE however, for virtual supervision we use the pseudo supervision used in SfSNet . We use the term \sayvirtual supervision to summarize these two losses as a kind of weak supervision on the real examples.

Monocular Depth Estimation. To make use of ground truth labels for synthetic data, we apply L1L_{1} loss for predicted synthetic depth images:

where y^s\hat{y}_{s} is the predicted synthetic depth map and ys∗y_{s}^{*} is its corresponding ground truth. Following , we apply smoothness loss on depth LDSL_{DS} to encourage it to be consistent with local homogeneous regions. Geometric consistency loss LGCL_{GC} is applied so that the task network can learn the physical geometric structure through epipolar constraints. LDSL_{DS} and LGCL_{GC} are defined as:

y^r\hat{y}_{r} represents the predicted depth for the real image and ∇\nabla represents the first derivative. xrx_{r} is the left image in the KITTI dataset . xrr′x_{rr}^{\prime} is the inverse warped image from the right counterpart of xrx_{r} based on the predicted depth y^r\hat{y}_{r}. The KITTI dataset provides the camera focal length and the baseline distance between the cameras. Similar to , we set η\eta as 0.85 and μ\mu as 0.15 in our experiments. The overall loss for the task network is defined as:

where β1=0.01,β2=β3=100.\beta_{1}=0.01,\beta_{2}=\beta 3=100.

Face Normal Estimation. SfSnet currently achieves the best performance on face normal estimation. We thus follow its setup for face normal estimation and apply “SfS-supervision” for both synthetic and real images during training.

where LreconL_{recon}, LNL_{N} and LAL_{A} are L1L_{1} losses on the reconstructed image, normal and albedo, whereas LlightL_{light} is the L2 loss over the 27 dimensional spherical harmonic coefficients. The supervision for real images is from the “pseudo labels”, obtained by applying a pre-trained task network on real images. Please refer to for more details.

3 Overall loss

The overall loss used to train our geometry estimation pipeline is then defined as:

where (α1,α2,α3)=(1,10,1)(\alpha_{1},\alpha_{2},\alpha_{3})=(1,10,1) for monocular depth estimation task and (α1,α2,α3)=(1,10,0.1)(\alpha_{1},\alpha_{2},\alpha_{3})=(1,10,0.1) for face normal estimation task.

Experiments

We apply our proposed SharinGAN to monocular depth estimation and face normal estimation. We discuss the details of the experiments in this section.

Datasets Following , we use vKITTI and KITTI as synthetic and real datasets to train our network. vKITTI contains 21,26021,260 image-depth pairs, which are all used for training. KITTI provides 42,38242,382 stereo pairs, among which, 22,60022,600 images are used for training and 888888 are used for validation as suggested by .

Implementation details We use a generator GG and a primary task network TT, whose architectures are identical to . We pre-train the generative network GG on both synthetic and real data using reconstruction loss LrL_{r}. This results in an identity mapping that can help GG to keep as much of the input image’s geometry information as possible. Our task network is pre-trained using synthetic data with supervision. GG and TT are then trained end to end using Equation 10 for 150,000 iterations with a batch size of 2, by using an Adam optimizer with a learning rate of 1e−51e-5. The best model is selected based on the validation set of KITTI.

Results Table 1 shows the quantitative results on the eigen test split of the KITTI dataset for different methods on the MDE task. The proposed method outperforms the previous unsupervised domain adaptation methods for MDE on almost all the metrics. Especially, compared with , we reduce the absolute error by 19.7%19.7\% and 21.0%21.0\% on 80m cap and 50m cap settings respectively. Moreover, the performance of our method is much closer to the methods in a supervised setting , which was trained on the real KITTI dataset with ground truth depth labels. Figure 4 visually compares the predicted depth map from the proposed method with . We show three typical examples: near distance, medium distance, and far distance. It shows that our proposed method performs much better for predicting depth at details. For instance, our predicted depth map can better preserve the shape of the car (Figure 4 (a) and (c)) and the structure of the tree and the building behind it (Figure 4 (b)). This shows the advantage of our proposed SharinGAN compared with . learns to transfer real images to the synthetic domain and vice versa, which solves a much harder problem compared with SharinGAN, which removes a minimum of domain specific information. As a result, the quality of the transformation for may not be as good as the proposed method. Moreover, the unsupervised transformation cannot guarantee to keep the geometry information unchanged.

To understand how our generative network GG works, we show some examples of synthetic and real images, their transformed versions, and the difference images in Figure 6. This shows that GG mainly operates on edges. Since depth maps are mostly discontinuous at edges, they provide important cues for the geometry of the scene. On the other hand, due to the difference between the geometry and material of objects around the edges, the rendering algorithm may find it hard to render realistic edges compared with other parts of the scene. As a result, most of the domain specific information related to geometry lies in the edges, on which SharinGAN correctly focuses.

To demonstrate the generalization ability of the proposed method, we test our trained model on Make3D . Note that we do not fine-tune our model using the data from Make3D. Table 2 shows the quantitative results of our method, which outperforms existing state-of-the-art methods by a large margin.

Moreover, the performance of SharinGAN is more comparable to the supervised methods. We further visually compare the proposed method with GASDA in Figure 7. It is clear that the proposed depth map captures more details in the input images, reflecting more accurate depth prediction.

2 Face Normal Estimation

Datasets We use the synthetic data provided by and CelebA as real data to train the SharinGAN for face normal estimation similar to . Our trained model is then evaluated on the Photoface dataset .

Implementation details We use the RBDN network as our generator and SfSNet as the primary task network. Similar to before, we pre-train the Generator on both synthetic and real data using reconstruction loss and pre-train the primary task network on just synthetic data in a supervised manner. Then, we train GG and TT end-to-end using the overall loss (10) for 120,000 iterations. We use a batch size of 16 and a learning rate of 1e−41e-4. The best model is selected based on the validation set of Photoface.

Results Table 4 shows the quantitative performance of the estimated surface normals by our method on the test split of the Photoface dataset. With the proposed SharinGAN module, we were able to significantly improve over SfSNet on all the metrics. In particular, we were able to significantly reduce the mean angular error metric by roughly 1.5∘.

Additionally, Figure 8 depicts the qualitative comparison of our method with SfSNet on the test split of Photoface. Both SfSNet and our pipeline are not finetuned on this dataset, and yet we were able to generalize better compared to SfSNet. This demonstrates the generalization capacity of the proposed SharinGAN to unseen data in training.

Ablation studies

We carried out our ablation study using the KITTI and Make3D datasets on monocular depth estimation. We study the role of the SharinGAN module by removing it and training a primary network on the original synthetic and real data using (8). We observe that the performance drops significantly as shown in Table 3 and Table 5. This shows the importance of the SharinGAN module that helps train the primary task network efficiently.

To demonstrate the role of reconstruction loss, we remove it and train our whole pipeline α1Ladv+α3LT\alpha_{1}L_{adv}+\alpha_{3}L_{T}. We show the results on the testset of KITTI in the second row of Table 3 and on the testset of Make3D in the second row of Table 5. For both the testsets, we can see the performance drop compared to our full model. Although the drop is smaller in the case of KITTI, it can be seen that the drop is significant for Make3D dataset that is unseen during training. This signifies the importance of reconstruction loss to generalize well to a domain not seen during training.

Conclusion

Our primary motivation is to simplify the process of combining synthetic and real images in training. Prior approaches often pick one domain and try to map images into it from the other domain. Instead, we train a generator to map all images into a new, shared domain. In doing this, we note that in the new domain, the images need not be indistinguishable to the human eye, only to the network that performs the primary task. The primary network will learn to ignore extraneous, domain-specific information that is retained in the shared domain.

To achieve this, we propose a simple network architecture that rests on our new SharinGAN, which maps both real and synthetic images to a shared domain. The resulting images retain domain-specific details that do not prevent the primary network from effectively combining training data from both domains. We demonstrate this by achieving significant improvements over state-of-the-art approaches in two important applications, surface normal estimation for faces, and monocular depth estimation for outdoor scenes. Finally, our ablation studies demonstrate the significance of the proposed SharinGAN in effectively combining synthetic and real data.

References

More Implementation details

The discriminator architecture we used for this work is: {CBR(n,3,1),CBR(2∗n,3,2)}n={32,64,128,256}\{CBR(n,3,1),CBR(2*n,3,2)\}_{n=\{32,64,128,256\}}, {CBR(512,3,1),CBR(512,3,2)}Ksets{\{CBR(512,3,1),CBR(512,3,2)\}}_{Ksets}, {FcBR(1024)\{FcBR(1024), FcBR(512)FcBR(512), Fc(1)}Fc(1)\}, where, CBR(out channels, kernel size, stride) = Conv + BatchNorm2d + ReLU and FcBR(out nodes) = Fully conncected + BatchNorm1D + ReLU and Fc is a fully connected layer. For face normal estimation, we do not use batchnorm layers in the discriminator. We use the value K=2K=2 for MDE and K=1K=1 for FNE.

Face Normal Estimation We update the generator 3 times for each update of the discriminator, which in turn is updated 5 times internally as per . The generator learns from a new batch each time, while the discriminator trains on a single batch for 5 times.

Experiments

Monocular Depth Estimation We provide more qualitative results on the test set of the Make3D dataset . Figure 10 further demonstrates the generalization ability of our method compared to .

Face Normal Estimation Figure 11 depicts the qualitative results on the CelebA and Synthetic datasets. The translated images corresponding to synthetic and real images look similar in contrast to the MDE task (Figure 4 of the paper). We suppose that for the task of MDE, regions such as edges are domain specific, and yet hold primary task related information such as depth cues, which is why SharinGAN modifies such regions. However, for the task of FNE, we additionally predict albedo, lighting, shading and a reconstructed image along with estimating normals. This means that the primary network needs a lot of shared information across domains for good generalization to real data. Thus the SharinGAN module seems to bring everything into a shared space, making the translated images {xrsh,xssh}\{x_{r}^{sh},x_{s}^{sh}\} look visually similar.

Figure 9 depicts additional qualitative results of the predicted face normals for the test set of the Photoface dataset .

Lighting Estimation The primary network estimates not only face normals but also lighting. We also evaluate this. Following a similar evaluation protocol as that of , Table 6 summarizes the light classification accuracy on the MultiPIE dataset . Since we do not have the exact cropped dataset that used, we used our own cropping and resizing on the original MultiPIE data: centercrop 300x300 and resize to 128x128. For a fair comparison, we used the same dataset to re-evaluate the lighting performance for and reported the results in Table 6. Our method not only outperforms on the face normal estimation, but also on lighting estimation.