Multi-Channel Attention Selection GAN with Cascaded Semantic Guidance for Cross-View Image Translation

Hao Tang, Dan Xu, Nicu Sebe, Yanzhi Wang, Jason J. Corso, Yan Yan

Introduction

Cross-view image translation is a task that aims at synthesizing new images from one viewpoint to another. It has been gaining a lot interest especially from computer vision and virtual reality communities, and has been widely investigated in recent years . Earlier works studied this problem using encoder-decoder Convolutional Neural Networks (CNNs) by involving viewpoint codes in the bottle-neck representations for city scene synthesis and 3D object translation . There also exist some works exploring Generative Adversarial Networks (GAN) for similar tasks . However, these existing works consider an application scenario in which the objects and the scenes have a large degree of overlapping in appearances and views.

Different from previous works, in this paper, we focus on a more challenging setting in which fields of views have little or even no overlap, leading to significantly distinct structures and appearance distributions for the input source and the output target views, as illustrated in Fig. 1. To tackle this challenging problem, Regmi and Borji recently proposed a conditional GAN model which jointly learns the generation in both the image domain and the corresponding semantic domain, and the semantic predictions are further utilized to supervise the image generation. Although this approach performed an interesting exploration, we observe unsatisfactory aspects mainly in the generated scene structure and details, which are due to different reasons. First, since it is always costly to obtain manually annotated semantic labels, the label maps are usually produced from pretrained semantic models from other large-scale segmentation datasets, leading to insufficiently accurate predictions for all the pixels, and thus misguiding the image generation. Second, we argue that the translation with a single phase generation network is not able to capture the complex scene structural relationships between the two views. Third, a three-channel generation space may not be suitable enough for learning a good mapping for this complex synthesis problem. Given these problems, could we enlarge the generation space and learn an automatic selection mechanism to synthesize more fine-grained generation results?

Based on these observations, in this paper, we propose a novel Multi-Channel Attention Selection Generative Adversarial Network (SelectionGAN), which contains two generation stages. The overall framework of the proposed SelectionGAN is shown in Fig. 2. In this first stage, we learn a cycled image-semantic generation sub-network, which accepts a pair consisting of an image and the target semantic map, and generates images for the other view, which further fed into a semantic generation network to reconstruct the input semantic maps. This cycled generation adds more strong supervision between the image and semantic domains, facilitating the optimization of the network.

The coarse outputs from the first generation network, including the input image, together with the deep feature maps from the last layer, are input into the second stage networks. Several intermediate outputs are produced, and simultaneously we learn a set of multi-channel attention maps with the same number as the intermediate generations. These attention maps are used to spatially select from the intermediate generations, and are combined to synthesize a final output. Finally, to overcome the inaccurate semantic label issue, the multi-channel attention maps are further used to generate uncertainty maps to guide the reconstruction loss. Through extensive experimental evaluations, we demonstrate that SelectionGAN produces remarkably better results than the baselines such as Pix2pix , Zhai et al. , X-Fork and X-Seq . Moreover, we establish state-of-the-art results on three different datasets for the arbitrary cross-view image synthesis task.

Overall, the contributions of this paper are as follows:

A novel multi-channel attention selection GAN framework (SelectionGAN) for the cross-view image translation task is presented. It explores cascaded semantic guidance with a coarse-to-fine inference, and aims at producing a more detailed synthesis from richer and more diverse multiple intermediate generations.

A novel multi-channel attention selection module is proposed, which is utilized to attentively select interested intermediate generations and is able to significantly boost the quality of the final output. The multi-channel attention module also effectively learns uncertainty maps to guide the pixel loss for more robust optimization.

Extensive experiments clearly demonstrate the effectiveness of the proposed SelectionGAN, and show state-of-the-art results on two public benchmarks, i.e. Dayton and CVUSA . Meanwhile, we also create a larger-scale cross-view synthesis benchmark using the data from Ego2Top , and present results of multiple baseline models for the research community.

Related Work

Generative Adversarial Networks (GANs) have shown the capability of generating better high-quality images , compared to existing methods such as Restricted Boltzmann Machines and Deep Boltzmann Machines . A vanilla GAN model has two important components, i.e. a generator GG and a discriminator DD. The goal of GG is to generate photo-realistic images from a noise vector, while DD is trying to distinguish between a real image and the image generated by GG. Although it is successfully used in generating images of high visual fidelity , there are still some challenges, i.e. how to generate images in a controlled setting. To generate domain-specific images, Conditional GAN (CGAN) has been proposed. CGAN usually combines a vanilla GAN and some external information, such as class labels or tags , text descriptions , human pose and reference images .

Image-to-Image Translation frameworks adopt input-output data to learn a parametric mapping between inputs and outputs. For example, Isola et al. propose Pix2pix, which is a supervised model and uses a CGAN to learn a translation function from input to output image domains. Zhu et al. introduce CycleGAN, which targets unpaired image translation using the cycle-consistency loss. To further improve the generation performance, the attention mechanism has been recently investigated in image translation, such as . However, to the best of our knowledge, our model is the first attempt to incorporate a multi-channel attention selection module within a GAN framework for image-to-image translation task.

Learning Viewpoint Transformations. Most existing works on viewpoint transformation have been conducted to synthesize novel views of the same object, such as cars, chairs and tables . Another group of works explore the cross-view scene image generation, such as . However, these works focus on the scenario in which the objects and the scenes have a large degree of overlapping in both appearances and views. Recently, several works started investigating image translation problems with drastically different views and generating a novel scene from a given arbitrary one. This is a more challenging task since different views have little or no overlap. To tackle this problem, Zhai et al. try to generate panoramic ground-level images from aerial images of the same location by using a convolutional neural network. Krishna and Ali propose a X-Fork and a X-Seq GAN-based structure to address the aerial to street view image translation task using an extra semantic segmentation map. However, these methods are not able to generate satisfactory results due to the drastic difference between source and target views and their model design. To overcome these issues, we aim at a more effective network design, and propose a novel multi-channel attention selection GAN, which allows to automatically select from multiple diverse and rich intermediate generations and thus significantly improves the generation quality.

Multi-Channel Attention Selection GAN

In this section we present the details of the proposed multi-channel attention selection GAN. An illustration of the overall network structure is depicted in Fig. 2. In the first stage, we present a cascade semantic-guided generation sub-network, which utilizes the images from one view and conditional semantic maps from another view as inputs, and reconstruct images in another view. These images are further input into a semantic generator to recover the input semantic map forming a generation cycle. In the second stage, the coarse synthesis and the deep features from the first stage are combined, and then are passed to the proposed multi-channel attention selection module, which aims at producing more fine-grained synthesis from a larger generation space and also at generating uncertainty maps to guide multiple optimization losses.

Semantic-guided Generation. Cross-view synthesis is a challenging task, especially when the two views have little overlapping as in our study case, which apparently leads to ambiguity issues in the generation process. To alleviate this problem, we use semantic maps as conditional guidance. Since it is always costly to obtain annotated semantic maps, following we generate the maps using segmentation deep models pretrained from large-scale scene parsing datasets such as Cityscapes . However, uses semantic maps only in the reconstruction loss to guide the generation of semantics, which actually provides a weak guidance. Different from theirs, we apply the semantic maps not only in the output loss but also as part of the network’s input. Specifically, as shown in Fig. 2, we concatenate the input image IaI_{a} from the source view and the semantic map SgS_{g} from a target view, and input them into the image generator GiG_{i} and synthesize the target view image Ig′I_{g}^{{}^{\prime}} as Ig′=Gi(Ia,Sg)I_{g}^{{}^{\prime}}{=}G_{i}(I_{a},S_{g}). In this way, the ground-truth semantic maps provide stronger supervision to guide the cross-view translation in the deep network.

Semantic-guided Cycle. Regmi and Borji observed that the simultaneous generation of both the images and the semantic maps improves the generation performance. Along the same line, we propose a cycled semantic generation network to benefit more the semantic information in learning. The conditional semantic map SgS_{g} together with the input image IaI_{a} are input into the image generator GiG_{i}, and produce the synthesized image Ig′I_{g}^{{}^{\prime}}. Then Ig′I_{g}^{{}^{\prime}} is further fed into the semantic generator GsG_{s} which reconstructs a new semantic map Sg′S_{g}^{{}^{\prime}}. We can formalize the process as Sg′=Gs(Ig′)=Gs(Gi(Ia,Sg))S_{g}^{{}^{\prime}}{=}G_{s}(I_{g}^{{}^{\prime}}){=}G_{s}(G_{i}(I_{a},S_{g})). Then the optimization objective is to make Sg′S_{g}^{{}^{\prime}} as close as possible to SgS_{g}, which naturally forms a semantic generation cycle, i.e. [Ia,Sg]→GiIg′→GsSg′≈Sg[I_{a},S_{g}]\stackrel{{\scriptstyle G_{i}}}{{\rightarrow}}I_{g}^{{}^{\prime}}\stackrel{{\scriptstyle G_{s}}}{{\rightarrow}}S_{g}^{{}^{\prime}}{\approx}S_{g}. The two generators are explicitly connected by the ground-truth semantic maps, which in this way provide extra constraints on the generators to learn better the semantic structure consistency.

Cascade Generation. Due to the complexity of the task, after the first stage, we observe that the image generator GiG_{i} outputs a coarse synthesis, which yields blurred scene details and high pixel-level dis-similarity with the target-view images. This inspires us to explore a coarse-to-fine generation strategy in order to boost the synthesis performance based on the coarse predictions. Cascade models have been used in several other computer vision tasks such as object detection and semantic segmentation , and have shown great effectiveness. In this paper, we introduce the cascade strategy to deal with the complex cross-view translation problem. In both stages we have a basic cycled semantic guided generation sub-network, while in the second stage, we propose a novel multi-channel attention selection module to better utilize the coarse outputs from the first stage and produce fine-grained final outputs. We observed significant improvement by using the proposed cascade strategy, illustrated in the experimental part.

2 Multi-Channel Attention Selection

An overview of the proposed multi-channel attention selection module GaG_{a} is shown in Fig. 3. The module consists of a multi-scale spatial pooling and a multi-channel attention selection component.

Multi-Channel Attention Selection. In previous cross-view image synthesis works, the image is generated only in a three-channel RGB space. We argue that this is not enough for the complex translation problem we are dealing with, and thus we explore using a larger generation space to have a richer synthesis via constructing multiple intermediate generations. Accordingly, we design a multi-channel attention mechanism to automatically perform spatial and temporal selection from the generations to synthesize a fine-grained final output.

where Ig′′I_{g}^{{}^{\prime\prime}} represents the final synthesized generation selected from the multiple diverse results, and the symbol ⊕\oplus denotes the element-wise addition. We also generate a final semantic map in the second stage as in the first stage, i.e. Sg′′=Gs(Ig′′)S_{g}^{{}^{\prime\prime}}{=}G_{s}(I_{g}^{{}^{\prime\prime}}). Due to the same purpose of the two semantic generators, we use a single GsG_{s} twice by sharing the parameters in both stages to reduce the network capacity.

Uncertainty-guided Pixel Loss. As we discussed in the introduction, the semantic maps obtained from the pretrained model are not accurate for all the pixels, which leads to a wrong guidance during training. To tackle this issue, we propose the generated attention maps to learn uncertainty maps to control the optimization loss. The uncertainty learning has been investigated in for multi-task learning, and here we introduce it for solving the noisy semantic label problem. Assume that we have KK different loss maps which need a guidance. The multiple generated attention maps are first concatenated and passed to a convolution layer with KK filters {Wui}i=1K\{W_{u}^{i}\}_{i=1}^{K} to produce a set of KK uncertainty maps. The reason of using the attention maps to generate uncertainty maps is that the attention maps directly affect the final generation leading to a close connection with the loss. Let Lpi\mathcal{L}_{p}^{i} denote a pixel-level loss map and UiU_{i} denote the ii-th uncertainty map, we have:

where σ(⋅)\sigma(\cdot) is a Sigmoid function for pixel-level normalization. The uncertainty map is automatically learned and acts as a weighting scheme to control the optimization loss.

Parameter-Sharing Discriminator. We extend the vanilla discriminator in to a parameter-sharing structure. In the first stage, this structure takes the real image IaI_{a} and the generated image Ig′I_{g}^{{}^{\prime}} or the ground-truth image IgI_{g} as input. The discriminator DD learns to tell whether a pair of images from different domains is associated with each other or not. In the second stage, it accepts the real image IaI_{a} and the generated image Ig′′I_{g}^{{}^{\prime\prime}} or the real image IgI_{g} as input. This pairwise input encourages DD to discriminate the diversity of image structure and capture the local-aware information.

3 Overall Optimization Objective

Adversarial Loss. In the first stage, the adversarial loss of DD for distinguishing synthesized image pairs [Ia,Ig′][I_{a},I_{g}^{{}^{\prime}}] from real image pairs [Ia,Ig][I_{a},I_{g}] is formulated as follows,

In the second stage, the adversarial loss of DD for distinguishing synthesized image pairs [Ia,Ig′′][I_{a},I_{g}^{{}^{\prime\prime}}] from real image pairs [Ia,Ig][I_{a},I_{g}] is formulated as follows:

Both losses aim to preserve the local structure information and produce visually pleasing synthesized images. Thus, the adversarial loss of the proposed SelectionGAN is the sum of Eq. (6) and (7),

Overall Loss. The total optimization loss is a weighted sum of the above losses. Generators GiG_{i}, GsG_{s}, attention selection network GaG_{a} and discriminator DD are trained in an end-to-end fashion optimizing the following min-max function,

where Lpi\mathcal{L}_{p}^{i} uses the L1 reconstruction to separately calculate the pixel loss between the generated images Ig′I_{g}^{{}^{\prime}}, Sg′S_{g}^{{}^{\prime}}, Ig′′I_{g}^{{}^{\prime\prime}} and Sg′′S_{g}^{{}^{\prime\prime}} and the corresponding real images. Ltv\mathcal{L}_{tv} is the total variation regularization on the final synthesized image Ig′′I_{g}^{{}^{\prime\prime}}. λi\lambda_{i} and λtv\lambda_{tv} are the trade-off parameters to control the relative importance of different objectives. The training is performed by solving the min-max optimization problem.

4 Implementation Details

Network Architecture. For a fair comparison, we employ U-Net as our generator architectures GiG_{i} and GsG_{s}. U-Net is a network with skip connections between a down-sampling encoder and an up-sampling decoder. Such architecture comprehensively retains contextual and textural information, which is crucial for removing artifacts and padding textures. Since our focus is on the cross-view image generation task, GiG_{i} is more important than GsG_{s}. Thus we use a deeper network for GiG_{i} and a shallow network for GsG_{s}. Specifically, the filters in first convolutional layer of GiG_{i} and GsG_{s} are 64 and 4, respectively. For the network GaG_{a}, the kernel size of convolutions for generating the intermediate images and attention maps are 3×33{\times}3 and 1×11{\times}1, respectively. We adopt PatchGAN for the discriminator DD.

Training Details. Following , we use RefineNet and to generate segmentation maps on Dayton and Ego2Top datasets as training data, respectively. We follow the optimization method in to optimize the proposed SelectionGAN, i.e. one gradient descent step on discriminator and generators alternately. We first train GiG_{i}, GsG_{s}, GaG_{a} with DD fixed, and then train DD with GiG_{i}, GsG_{s}, GaG_{a} fixed. The proposed SelectionGAN is trained and optimized in an end-to-end fashion. We employ Adam with momentum terms β1=0.5\beta_{1}{=}0.5 and β2=0.999\beta_{2}{=}0.999 as our solver. The initial learning rate for Adam is 0.0002. The network initialization strategy is Xavier , weights are initialized from a Gaussian distribution with standard deviation 0.2 and mean 0.

Experiments

Datasets. We perform the experiments on three different datasets: (i) For the Dayton dataset , following the same setting of , we select 76,048 images and create a train/test split of 55,000/21,048 pairs. The images in the original dataset have 354×354354{\times}354 resolution. We resize them to 256×256256{\times}256; (ii) The CVUSA dataset consists of 35,532/8,884 image pairs in train/test split. Following , the aerial images are center-cropped to 224×224224{\times}224 and resized to 256×256256{\times}256. For the ground level images and corresponding segmentation maps, we take the first quarter of both and resize them to 256×256256{\times}256; (iii) The Ego2Top dataset is more challenging and contains different indoor and outdoor conditions. Each case contains one top-view video and several egocentric videos captured by the people visible in the top-view camera. This dataset has more than 230,000 frames. For training data, we randomly select 386,357 pairs and each pair is composed of two images of the same scene but different viewpoints. We randomly select 25,600 pairs for evaluation.

Parameter Settings. For a fair comparison, we adopt the same training setup as in . All images are scaled to 256×256256{\times}256, and we enabled image flipping and random crops for data augmentation. Similar to , the low resolution (64×6464{\times}64) experiments on Dayton dataset are carried out for 100 epochs with batch size of 16, whereas the high resolution (256×256256{\times}256) experiments for this dataset are trained for 35 epochs with batch size of 4. For the CVUSA dataset, we follow the same setup as in , and train our network for 30 epochs with batch size of 4. For the Ego2Top dataset, all models are trained with 10 epochs using batch size 8. In our experiment, we set λtv\lambda_{tv}=1e−61e{-}6, λ1=100\lambda_{1}{=}100, λ2=1\lambda_{2}{=}1, λ3=200\lambda_{3}{=}200 and λ4=2\lambda_{4}{=}2 in Eq. (9), and λ=4\lambda{=}4 in Eq. (8). The number of attention channels NN in Eq. (3) is set to 10. The proposed SelectionGAN is implemented in PyTorch. We perform our experiments on Nvidia GeForce GTX 1080 Ti GPU with 11GB memory to accelerate both training and inference.

Evaluation Protocol. Similar to , we employ Inception Score, top-k prediction accuracy and KL score for the quantitative analysis. These metrics evaluate the generated images from a high-level feature space. We also employ pixel-level similarity metrics to evaluate our method, i.e. Structural-Similarity (SSIM), Peak Signal-to-Noise Ratio (PSNR) and Sharpness Difference (SD).

2 Experimental Results

Ablation Analysis. The results of ablation study are shown in Table 4. We observe that Baseline B is better than baseline A since SgS_{g} contains more structural information than IaI_{a}. By comparison Baseline A with Baseline C, the semantic-guided generation improves SSIM, PSNR and SD by 8.19, 3.1771 and 0.3205, respectively, which confirms the importance of the conditional semantic information; By using the proposed cycled semantic generation, Baseline D further improves over C, meaning that the proposed semantic cycle structure indeed utilizes the semantic information in a more effective way, confirming our design motivation; Baseline E outperforms D showing the importance of using the uncertainty maps to guide the pixel loss map which contains an inaccurate reconstruction loss due to the wrong semantic labels produced from the pretrained segmentation model; Baseline F significantly outperforms E with around 4.67 points gain on the SSIM metric, clearly demonstrating the effectiveness of the proposed multi-channel attention selection scheme; We can also observe from Table 4 that, by adding the proposed multi-scale spatial pool scheme and the TV regularization, the overall performance is further boosted. Finally, we demonstrate the advantage of the proposed two-stage strategy over the one-stage method. Several examples are shown in Fig. 5. It is obvious that the coarse-to-fine generation model is able to generate sharper results and contains more details than the one-stage model.

State-of-the-art Comparisons. We compare our SelectionGAN with four recently proposed state-of-the-art methods, which are Pix2pix , Zhai et al. , X-Fork and X-Seq . The comparison results are shown in Tables 1, 2, 3, and 5. We can observe the significant improvement of SelectionGAN in these tables. SelectionGAN consistently outperforms Pix2pix, Zhai et al., X-Fork and X-Seq on all the metrics except for Inception Score. In some cases in Table 3 we achieve a slightly lower performance as compared with X-Seq. However, we generate much more photo-realistic results than X-Seq as shown in Fig. 4 and 6.

Qualitative Evaluation. The qualitative results in higher resolution on Dayton and CVUSA datasets are shown in Fig. 4 and 6. It can be seen that our method generates more clear details on objects/scenes such as road, tress, clouds, car than the other comparison methods in the generated ground level images. For the generated aerial images, we can observe that grass, trees and house roofs are well rendered compared to others. Moreover, the results generated by our method are closer to the ground truths in layout and structure, such as the results in a2g direction in Fig. 4 and 6.

Arbitrary Cross-View Image Translation. Since Dayton and CVUSA datasets only contain two views in one scene, i.e. aerial and ground views. We further use the Ego2Top dataset to conduct the arbitrary cross-view image translation experiments. The quantitative and qualitative results are shown in Table 5 and Fig. 7, respectively. Given an image and some novel semantic maps, SelectionGAN is able to generate the same scene but with different viewpoints.

Conclusion

We propose the Multi-Channel Attention Selection GAN (SelectionGAN) to address a novel image synthesizing task by conditioning on a reference image and a target semantic map. In particular, we adopt a cascade strategy to divide the generation procedure into two stages. Stage I aims to capture the semantic structure of the scene and Stage II focus on more appearance details via the proposed multi-channel attention selection module. We also propose an uncertainty map-guided pixel loss to solve the inaccurate semantic labels issue for better optimization. Extensive experimental results on three public datasets demonstrate that our method obtains much better results than the state-of-the-art.

Acknowledgements: This research was partially supported by National Institute of Standards and Technology Grant 60NANB17D191 (YY, JC), Army Research Office W911NF-15-1-0354 (JC) and gift donation from Cisco Inc (YY).

References

Influence of the Number of Attention Channels N𝑁N

We investigate the influence of the number of attention channels NN in Equation 3 in the main paper. Results are shown in Table 6. We observe that the performance tends to be stable after N=10N=10. Thus, taking both performance and training speed into consideration, we have set N=10N=10 in all our experiments.

Coarse-to-Fine Generation

We provide more comparison results of coarse-to-fine generation in Table 7 and Figures 8, 9 and 10. We observe that our two-stage method generate much visually better results than the one-stage model, which further confirms our motivations.

Visualization of Uncertainty Map

In Figures 8, 9, 10 and 11, we show some samples of the generated uncertainty maps. We can see that the generated uncertainty maps learn the layout and structure of the target images. Note that most textured regions are similar in our generation images, while the junction/edge of different regions is uncertain, and thus the model learns to highlight these parts.

Arbitrary Cross-View Image Translation

We also conducted the arbitrary cross-view image translation experiments on Ego2Top dataset. As we can see from Figure 11, given an image and some novel semantic maps, SelectionGAN is able to generate the same scene but with different viewpoints in both outdoor and indoor environments.

Generated Segmentation Maps

Since the proposed SelectionGAN can generate segmentation maps, we also compare it with X-Fork and X-Seq on Dayton dataset. Following , we compute per-class accuracies and mean IOU for the most common classes in this dataset: “vegetation”, “road”, “building” and “sky” in ground segmentation maps. Results are shown in Table 8. We can see that the proposed SelectionGAN achieves better results than X-Fork and X-Seq on both metrics.

State-of-the-art Comparisons

In Figures 12, 13, 14, 15 and 16, we show more image generation results on Dayton, CVUSA and Ego2Top datasets compared with the state-of-the-art methods i.e., Pix2pix , X-Fork and X-Seq . For Figures 12, 13, 14, 15, we reproduced the results of Pix2pix , X-Fork and X-Seq using the pre-trained models provided by the authorshttps://github.com/kregmi/cross-view-image-synthesis. As we can see from all these figures, the proposed SelectionGAN achieves significantly visually better results than the competing methods.