Pluralistic Image Completion

Chuanxia Zheng, Tat-Jen Cham, Jianfei Cai

Introduction

Image completion is a highly subjective process. Supposing you were shown the various images with missing regions in fig. 1, what would you imagine to be occupying these holes? Bertalmio et al. related how expert conservators would inpaint damaged art by: 1) imagining the semantic content to be filled based on the overall scene; 2) ensuring structural continuity between the masked and unmasked regions; and 3) filling in visually realistic content for missing regions. Nonetheless, each expert will independently end up creating substantially different details, even if they may universally agree on high-level semantics, such as general placement of eyes on a damaged portrait.

Based on this observation, our main goal is thus to generate multiple and diverse plausible results when presented with a masked image — in this paper we refer to this task as pluralistic image completion (depicted in fig. 1). This is as opposed to approaches that attempt to generate only a single “guess” for missing parts.

Early image completion works focus only on steps 2 and 3 above, by assuming that gaps should be filled with similar content to that of the background. Although these approaches produced high-quality texture-consistent images, they cannot capture global semantics and hallucinate new content for large holes. More recently, some learning-based image completion methods were proposed that infer semantic content (as in step 1). These works treated completion as a conditional generation problem, where the input-to-output mapping is one-to-many. However, these prior works are limited to generate only one “optimal” result, and do not have the capacity to generate a variety of semantically meaningful results.

To obtain a diverse set of results, some methods utilize conditional variational auto-encoders (CVAE) , a conditional extension of VAE , which explicitly code a distribution that can be sampled. However, specifically for an image completion scenario, the standard single-path formulation usually leads to grossly underestimating variances. This is because when the condition label is itself a partial image, the number of instances in the training data that match each label is typically only one. Hence the estimated conditional distributions tend to have very limited variation since they were trained to reconstruct the single ground truth. This is further elaborated on in section 3.1.

An important insight we will use is that partial images, as a superset of full images, may also be considered as generated from a latent space with smooth prior distributions. This provides a mechanism for alleviating the problem of having scarce samples per conditional partial image. To do so, we introduce a new image completion network with two parallel but linked training pipelines. The first pipeline is a VAE-based reconstructive path that not only utilizes the full instance ground truth (i.e. both the visible partial image, as well as its complement — the hidden partial image), but also imposes smooth priors for the latent space of complement regions. The second pipeline is a generative path that predicts the latent prior distribution for the missing regions conditioned on the visible pixels, from which can be sampled to generate diverse results. The training process for the latter path does not attempt to steer the output towards reconstructing the instance-specific hidden pixels at all, instead allowing the reasonableness of results be driven by an auxiliary discriminator network . This leads to substantially great variability in content generation. We also introduce an enhanced short+long term attention layer that significantly increases the quality of our results.

We compared our method with existing state-of-the-art approaches on multiple datasets. Not only can higher-quality completion results be generated using our approach, it also presents multiple diverse solutions.

A probabilistically principled framework for image completion that is able to maintain much higher sample diversity as compared to existing methods;

A new network structure with two parallel training paths, which trades off between reconstructing the original training data (with loss of diversity) and maintaining the variance of the conditional distribution;

A novel self-attention layer that exploits short+long term context information to ensure appearance consistency in the image domain, in a manner superior to purely using GANs; and

We demonstrate that our method is able to complete the same mask with multiple plausible results that have substantial diversity, such as those shown in figure 1.

Related Work

Existing work on image completion either uses information from within the input image , or information from a large image dataset . Most approaches will generate only one result per masked image.

Intra-Image Completion Traditional intra-image completion, such as diffusion-based methods and patch-based methods , assume image holes share similar content to visible regions; thus they would directly match, copy and realign the background patches to complete the holes. These methods perform well for background completion, e.g. for object removal, but cannot hallucinate unique content not present in the input images.

Inter-Image Completion To generate semantically new content, inter-image completion borrows information from a large dataset. Hays and Efros presented an image completion method using millions of images, in which the image most similar to the masked input is retrieved, and corresponding regions are transferred. However, this requires a high contextual match, which is not always available. Recently, learning-based approaches were proposed. Initial works focused on small and thin holes. Context encoders (CE) handled 64×\times64-sized holes using GANs . This was followed by several CNN-based methods, which included combining global and local discriminators as adversarial loss , identifying closest features in the latent space of masked images , utilizing semantic labels to guide the completion network , introducing additional face parsing loss for face completion , and designing particular convolutions to address irregular holes . A common drawback of these methods is that they often create distorted structures and blurry textures inconsistent with the visible regions, especially for large holes.

Combined Intra- and Inter-Image Completion To overcome the above problems, Yang et al. proposed multi-scale neural patch synthesis, which generates high-frequency details by copying patches from mid-layer features. However, this optimization is computational costly. More recently, several works exploited spatial attention to get high-frequency details. Yu et al. presented a contextual attention layer to copy similar features from visible regions to the holes. Yan et al. and Song et al. proposed PatchMatch-like ideas on feature domain. However, these methods identify similar features by comparing features of holes and features of visible regions, which is somewhat contradictory as feature transfer is unnecessary when two features are very similar, but when needed the features are too different to be matched easily. Furthermore, distant information is not used for new content that differs from visible regions. Our model will solve this problem by extending self-attention to harness abundant context.

Image Generation Image generation has progressed significantly using methods such as VAE and GANs . These have been applied to conditional image generation tasks, such as image translation , synthetic to realistic , future prediction , and 3D models . Perhaps most relevant are conditional VAEs (CVAE) and CVAE-GAN , but these were not specially targeted for image completion. CVAE-based methods are most useful when the conditional labels are few and discrete, and there are sufficient training instances per label. Some recent work utilizing these in image translation can produce diverse output , but in such situations the condition-to-sample mappings are more local (e.g. pixel-to-pixel), and only change the visual appearance. This is untrue for image completion, where the conditional label is itself the masked image, with only one training instance of the original holes. In , different outputs were obtained for face completion by specifying facial attributes (e.g. smile), but this method is very domain specific, requiring targeted attributes.

Approach

Suppose we have an image, originally Ig\mathbf{I}_{g}, but degraded by a number of missing pixels to become Im\mathbf{I}_{m} (the masked partial image) comprising the observed / visible pixels. We also define Ic\mathbf{I}_{c} as its complement partial image comprising the ground truth hidden pixels. Classical image completion methods attempt to reconstruct the ground truth unmasked image Ig\mathbf{I}_{g} in a deterministic fashion from Im\mathbf{I}_{m} (see fig. 2 “Deterministic”). This results in only a single solution. In contrast, our goal is to sample from p(Ic∣Im)p(\mathbf{I}_{c}|\mathbf{I}_{m}).

In order to have a distribution to sample from, a current approach is to employ the CVAE which estimates a parametric distribution over a latent space, from which sampling is possible (see fig. 2 “CVAE”). This involves a variational lower bound of the conditional log-likelihood of observing the training instances:

where zc\mathbf{z}_{c} is the latent vector, qψ(⋅∣⋅)q_{\psi}(\cdot|\cdot) the posterior importance sampling function, pϕ(⋅∣⋅)p_{\phi}(\cdot|\cdot) the conditional prior, pθ(⋅∣⋅)p_{\theta}(\cdot|\cdot) the likelihood, with ψ\psi, ϕ\phi and θ\theta being the deep network parameters of their corresponding functions. This lower bound is maximized w.r.t. all parameters.

A possible way to diversify the output is to simply not incentivize the output to reconstruct the instance-specific Ig\mathbf{I}_{g} during training, only needing it to fit in with the training set distribution as deemed by an learned adversarial discriminator (see fig. 2 “Instance Blind”). However, this approach is unstable, especially for large and complex scenes .

Latent Priors of Holes In our approach, we require that missing partial images, as a superset of full images, to also arise from a latent space distribution, with a smooth prior of p(zc)p(\mathbf{z}_{c}). The variational lower bound is:

where in the prior is set as p(zc)=N(0,I)p(\mathbf{z}_{c})=\mathcal{N}(\mathbf{0},\mathbf{I}). However, we can be more discerning when it comes to partial images since they have different numbers of pixels. A missing partial image zc\mathbf{z}_{c} with more pixels (larger holes) should have greater latent prior variance than a missing partial image zc\mathbf{z}_{c} with fewer pixels (smaller holes). Hence we generalize the prior p(zc)=Nm(0,σ2(n)I)p(\mathbf{z}_{c})=\mathcal{N}_{m}(\mathbf{0},\sigma^{2}(n)\mathbf{I}) to adapt to the number of pixels nn.

Next, we combine the latent priors into the conditional lower bound of (3.1). This can be done by assuming zc\mathbf{z}_{c} is much more closely related to Ic\mathbf{I}_{c} than to Im\mathbf{I}_{m}, so qψ(zc∣Ic,Im)<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>≈</mo></mrow><annotationencoding="application/x−tex">≈</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.4831em;"></span><spanclass="mrel">≈</span></span></span></span></span>qψ(zc∣Ic)q_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{c},\mathbf{I}_{m})<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>≈</mo></mrow><annotation encoding="application/x-tex">\approx</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4831em;"></span><span class="mrel">≈</span></span></span></span></span>q_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{c}). Updating (3.1):

However, unlike in (3.1), notice that qψ(zc∣Ic)q_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{c}) is no longer freely learned during training, but is tied to its presence in (3.1). Intuitively, the learning of qψ(zc∣Ic)q_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{c}) is regularized by the prior p(zc)p(\mathbf{z}_{c}) in (3.1), while the learning of the conditional prior pϕ(zc∣Im)p_{\phi}(\mathbf{z}_{c}|\mathbf{I}_{m}) is in turn regularized by qψ(zc∣Ic)q_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{c}) in (3.1).

One issue with (3.1) is that the sampling is taken from qψ(zc∣Ic)q_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{c}) during training, but is not available during testing, whereupon sampling must come from pϕ(zc∣Im)p_{\phi}(\mathbf{z}_{c}|\mathbf{I}_{m}) which may not be adequately learned for this role. In order to mitigate this problem, we modify (3.1) to have a blend of formulations with and without importance sampling. So, with simplified notation:

Our overall training objective may then be expressed as jointly maximizing the lower bounds in (3.1) and (3.1), with the likelihood in (3.1) unified to that in (3.1) as pθ(Ic∣zc)≅pθr(Ic∣zc,Im)p_{\theta}(\mathbf{I}_{c}|\mathbf{z}_{c})\cong p^{r}_{\theta}(\mathbf{I}_{c}|\mathbf{z}_{c},\mathbf{I}_{m}). See the supplemental section B.2.

2 Dual Pipeline Network Structure

This formulation is implemented as our dual pipeline framework, shown in fig. 3. It consists of two paths: the upper reconstructive path uses information from the whole image, i.e. Ig\mathbf{I}_{g}={Ic,Im}\{\mathbf{I}_{c},\mathbf{I}_{m}\}, while the lower generative path only uses information from visible regions Im\mathbf{I}_{m}. Both representation and generation networks share identical weights. Specifically:

For the upper reconstructive path, the complement partial image Ic\mathbf{I}_{c} is used to infer the importance function qψ(⋅∣Ic)q_{\psi}(\cdot|\mathbf{I}_{c})=Nψ(⋅)\mathcal{N}_{\psi}(\cdot) during training. The sampled latent vector zc\mathbf{z}_{c} thus contains information of the missing regions, while the conditional feature fm\mathbf{f}_{m} encodes the information of the visible regions. Since there is sufficient information, the loss function in this path is geared towards reconstructing the original image Ig\mathbf{I}_{g}.

For the lower generative path, which is also the test path, the latent distribution of the holes Ic\mathbf{I}_{c} is inferred based only on the visible Im\mathbf{I}_{m}. This would be significantly less accurate than the inference in the upper path. Thus the reconstruction loss is only targeted at the visible regions Im\mathbf{I}_{m} (via fm\mathbf{f}_{m}).

In addition, we also utilize adversarial learning networks on both paths, which ideally ensure that the full synthesized data fit in with the training set distribution, and empirically leads to higher quality images.

3 Training Loss

Various terms in (3.1) and (3.1) may be more conventionally expressed as loss functions. Jointly maximizing the lower bounds is then minimizing a total loss L\mathcal{L}, which consists of three groups of component losses:

where the LKL\mathcal{L}_{\text{KL}} group regularizes consistency between pairs of distributions in terms of KL divergences, the Lapp\mathcal{L}_{\text{app}} group encourages appearance matching fidelity, and while the Lad\mathcal{L}_{\text{ad}} group forces sampled images to fit in with the training set distribution. Each of the groups has a separate term for the reconstructive and generative paths.

The typical interpretation of the KL divergence term in a VAE is that it regularizes the learned importance sampling function qψ(⋅∣Ic)q_{\psi}(\cdot|\mathbf{I}_{c}) to a fixed latent prior p(zc)p(\mathbf{z}_{c}). Defining as Gaussians, we get:

For the generative path, the appropriate interpretation is reversed: the learned conditional prior pϕ(⋅∣Im)p_{\phi}(\cdot|\mathbf{I}_{m}), also a Gaussian, is regularized to qψ(⋅∣Ic)q_{\psi}(\cdot|\mathbf{I}_{c}).

Note that the conditional prior only uses Im\mathbf{I}_{m}, while the importance function has access to the hidden Ic\mathbf{I}_{c}.

The likelihood term pθr(Ic∣zc,Im)p^{r}_{\theta}(\mathbf{I}_{c}|\mathbf{z}_{c},\mathbf{I}_{m}) may be interpreted as probabilistically encouraging appearance matching to the hidden Ic\mathbf{I}_{c}. However, our framework also auto-encodes the visible Im\mathbf{I}_{m} deterministically, and the loss function needs to cater for this reconstruction. As such, the per-instance loss here is:

where Irec(i)I_{\text{rec}}^{(i)}=G(zc,fm)G(z_{c},f_{m}) and Ig(i)I_{g}^{(i)} are the reconstructed and original full images respectively. In contrast, for the generative path we ignore instance-specific appearance matching for Ic\mathbf{I}_{c}, and only focus on reconstructing Im\mathbf{I}_{m} (via fm\mathbf{f}_{m}):

where fD1(⋅)f_{D_{1}}(\cdot) is the feature output of the final layer of D1D_{1}. This encourages the original and reconstructed features in the discriminator to be close together. Conversely, the adversarial loss in the generative path for the generator is:

This is based on the generator loss in LSGAN , which performs better than the original GAN loss in our scenario. The discriminator loss for both D1D_{1} and D2D_{2} is also based on LSGAN.

4 Short+Long Term Attention

Extending beyond the Self-Attention GAN , we propose not only to use the self-attention map within a decoder layer to harness distant spatial context, but also to further capture feature-feature context between encoder and decoder layers. Our key novel insight is: doing so would allow the network a choice of attending to the finer-grained features in the encoder or the more semantically generative features in the decoder, depending on circumstances.

Our proposed structure is shown in fig. 4. We first calculate the self-attention map from the features fd\mathbf{f}_{d} of a decoder middle layer, using the attention score of:

NN is the number of pixels, Q(fd)Q(\mathbf{f}_{d})=Wqfd\mathbf{W}_{q}\mathbf{f}_{d}, and Wq\mathbf{W}_{q} is a 1x1 convolution filter. This leads to the short-term intra-layer attention feature (self-attention in fig. 4) and the output yd\mathbf{y}_{d}:

where, following , we use a scale parameter γd\gamma_{d} to balance the weights between cd\mathbf{c}_{d} and fd\mathbf{f}_{d}. The initial value of γd\gamma_{d} is set to zero. In addition, for attending to features fe\mathbf{f}_{e} from an encoder layer, we have a long-term inter-layer attention feature (contextual flow in fig. 4) and the output ye\mathbf{y}_{e}:

As before, a scale parameter γe\gamma_{e} is used to combine the encoder feature fe\mathbf{f}_{e} and the attention feature ce\mathbf{c}_{e}. However, unlike the decoder feature fd\mathbf{f}_{d} which has information for generating a full image, the encoder feature fe\mathbf{f}_{e} only represents visible parts Im\mathbf{I}_{m}. Hence, a binary mask M\mathbf{M} (holes=0) is used. Finally, both the short and long term attention features are aggregated and fed into further decoder layers.

Experimental Results

We evaluated our proposed model on four datasets including Paris , CelebA-HQ , Places2 , and ImageNet using the original training and test splits for those datasets. Since our model can generate multiple outputs, we sampled 5050 images for each masked image, and chose the top 10 results based on the discriminator scores. We trained our models for both regular and irregular holes. For brevity, we refer to our method as PICNet. We provide PyTorch implementations and interactive demo.

Our generator and discriminator networks are inspired by SA-GAN , but with several important modifications, including the short+long term attention layer. Furthermore, inspired by the growing-GAN , multi-scale output is applied to make the training faster.

The image completion network, implemented in Pytorch v0.4.0, contains 6M trainable parameters. During optimization, the weights of different losses are set to αKL=αrec\alpha_{\text{KL}}=\alpha_{\text{rec}}=20, αad\alpha_{\text{ad}}=1. We used Orthogonal Initialization and the Adam solver . All networks were trained from scratch, with a fixed learning rate of λ\lambda=10-410^{\text{-4}}. Details are in the supplemental section D.

2 Comparison with Existing Work

Quantitative evaluation is hard for the pluralistic image completion task, as our goal is to get diverse but reasonable solutions for one masked image. The original image is only one solution of many, and comparisons should not be made based on just this image.

Qualitative Comparisons First, we show the results in fig. 5 on the Paris dataset . For fair comparison among learning-based methods, we only compared with those trained on this dataset. PatchMatch worked by copying similar patches from visible regions and obtained good results on this dataset with repetitive structures. Context Encoder (CE) generated reasonable structures with blurry textures. Shift-Net made improvements by feature copying. Compared to these, our model not only generated more natural images, but also with multiple solutions, e.g. different numbers of windows and varying door sizes.

Next, we evaluated our methods on CelebA-HQ face dataset, with fig. 6 showing examples with large regular holes to highlight the diversity of our output. Context Attention (CA) generated reasonable completion for many cases, but for each masked input they were only able to generate a single result; furthermore, on some occasions, the single solution may be poor. Our model produced various plausible results by sampling from the latent space conditional prior.

Finally, we report the performance on the more challenging ImageNet dataset by comparing to the previous PatchMatch , CE , GL and CA . Different from the CE and GL models that were trained on the 100100k subset of training images of ImageNet, our model is directly trained on original ImageNet training dataset with all images resized to 256×256256\times 256. Visual results on a variety of objects from the validation set are shown in fig. 7. Our model was able to infer the content quite effectively.

3 Ablation Study

We investigated the influence of using our two-path training structure in comparison to other variants such as the CVAE and “instance blind” structures in fig. 2. We trained the three models using common parameters. As shown in fig. 9, for the CVAE, even after sampling from the latent prior distribution, the outputs were almost identical, as the conditional prior learned is narrowly centered at the maximum latent likelihood solution. As for “instance blind”, if reconstruction loss was used only on visible pixels, the training may become unstable. If we used reconstruction loss on the full generated image, there is also little variation as the framework has likely learned to ignore the sampling and predicted a deterministic outcome purely from Im\mathbf{I}_{m}.

We also trained and tested BicycleGAN for center masks. As is obvious in fig. 8, BicycleGAN is not directly suitable, leading to poor results or minimal variation.

Diversity Measure We computed diversity scores using the LPIPS metric reported in . The average score is calculated between 50K pairs generated from a sampling of 1K center-masked images. Iout\mathbf{I}_{out} and Iout(m)\mathbf{I}_{out(m)} are the full output and mask-region output, respectively. While obtained relatively higher diversity scores (still lower than ours), most of their generated images look unnatural (fig. 8).

Short+Long Term Attention vs Contextual Attention We visualized our attention maps as in . To compare to the contextual attention (CA) layer , we retrained CA on the Paris dataset via the authors’ code, and used their publicly released face model. The CA attention maps are presented in their color-directional format. As shown in fig. 10, our short+long term attention layer borrowed features from different positions with varying attention weights, rather than directly copying similar features from just one visible position. For the building scene, CA’s results were of similar high quality to ours, due to the repeated structures present. However for a face with a large mask, CA was unable to borrow features for the hidden content (e.g. mouth, eyes) from visible regions, with poor output. Our attention map is able to utilize both decoder features (which do not have masked parts) and encoder features as appropriate.

Conclusion

We proposed a novel dual pipeline training architecture for pluralistic image completion. Unlike existing methods, our framework can generate multiple diverse solutions with plausible content for a single masked input. The experimental results demonstrate this prior-conditional lower bound coupling is significant for conditional image generation. We also introduced an enhanced short+long term attention layer which improves realism. Experiments on a variety of datasets showed that our multiple solutions were diverse and of high-quality, especially for large holes.

This research is supported by the BeingTogether Centre, a collaboration between Nanyang Technological University (NTU) Singapore and University of North Carolina (UNC) at Chapel Hill. The BeingTogether Centre is supported by the National Research Foundation, Prime Minister’s Office, Singapore under its International Research Centres in Singapore Funding Initiative. This research was also conducted in collaboration with Singapore Telecommunications Limited and partially supported by the Singapore Government through the Industry Alignment Fund ‐- Industry Collaboration Projects Grant.

References

Appendix A A Additional Examples

We first show our results on center hole completion, in relation to those from other methods trained on corresponding datasets. As for random irregular and regular holes, we simply present our results so that readers may appreciate the multiple diverse results we can get with differently sized and shaped holes. Finally, we show the interesting application on face editing.

A.2 Additional Results on Random and Irregular Hole Completion

A.3 Additional Results on Free-Form Mask Using Our Interactive Demo

A.4 Video for Additional Results

Besides this document, we also included two video clips of additional results as part of the supplemental material. The first video, shows free-from mask results on various datasets. The second video consists of four parts to show multiple examples of center hole completion, random hole completion, comparison results with different training strategies and face editing of my self-portraits.

Appendix B B Mathematical Derivation and Analysis

Here we elaborate on the difficulties encountered when using the classical CVAE formulation for pluralistic image completion, expanding on the shorter description in section 3.1.

The broad CVAE framework of Sohn et al. is a straightforward conditioning of the classical VAE. Using the notation in our main paper, a latent variable zc\mathbf{z}_{c} is assumed to stochastically generate the hidden partial image Ic\mathbf{I}_{c}. When conditioned on the visible partial image Im\mathbf{I}_{m}, we get the conditional probability:

The variance of the Monte Carlo estimate can be reduced by importance sampling to get

Taking logs and apply Jensen’s inequality leads to

The variational lower bound V\mathcal{V} totaled over all training data is jointly maximized w.r.t. the network parameters θ\theta, ϕ\phi and ψ\psi in attempting to maximize the total log likelihood of the observed training instances.

B.1.2 Single Instance Per Conditioning Label

As is typically the case for image completion, there is only one training instance of Ic\mathbf{I}_{c} for each unique Im\mathbf{I}_{m}. This means that for the function qψ(zc∣Ic,Im)q_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{c},\mathbf{I}_{m}), Ic\mathbf{I}_{c} can simply be learnt into the network as a hardcoded dependency of the input Im\mathbf{I}_{m}, so qψ(zc∣Ic,Im)≅q^ψ(zc∣Im)q_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{c},\mathbf{I}_{m})\cong\hat{q}_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{m}). Assuming that the network for pϕ(zc∣Im)p_{\phi}(\mathbf{z}_{c}|\mathbf{I}_{m}) has similar or higher modeling power and there are no other explicit constraints imposed on it, then in training pϕ(zc∣Im)→q^ψ(zc∣Im)p_{\phi}(\mathbf{z}_{c}|\mathbf{I}_{m})\rightarrow\hat{q}_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{m}), and the KL divergence in (B.3) goes to zero.

In this situation of zero KL divergence, we can rewrite the variational lower bound and replace q^ψ(zc∣Im)\hat{q}_{\psi}(\mathbf{z}_{c}|\mathbf{I}_{m}) with pϕ(zc∣Im)p_{\phi}(\mathbf{z}_{c}|\mathbf{I}_{m}) without loss of generality, as

B.1.3 Unconstrained Learning of the Conditional Prior

We can analyze howV\mathcal{V} can be maximized, by using Jensen’s inequality again (reversing earlier use)

By further applying Hölder’s inequality (i.e. ∥fg∥1≤∥f∥p∥g∥q\left\|fg\right\|_{1}\leq\left\|f\right\|_{p}\left\|g\right\|_{q} for \nicefrac1p+\nicefrac1q=1\nicefrac{{1}}{{p}}+\nicefrac{{1}}{{q}}=1), we get

Assuming that there is a unique global maximum for log⁡pϕ(zc∣Im)\log p_{\phi}(\mathbf{z}_{c}|\mathbf{I}_{m}), the bound achieves equality when the conditional prior becomes a Dirac delta function centered at the maximum latent likelihood point

Intuitively, subject to the vagaries of stochastic gradient descent, the network for pϕ(zc∣Im)p_{\phi}(\mathbf{z}_{c}|\mathbf{I}_{m}) without further constraints will learn a narrow delta-like function that sifts out maximum latent likelihood value of log⁡pθ(Ic∣zc,Im)\log p_{\theta}(\mathbf{I}_{c}|\mathbf{z}_{c},\mathbf{I}_{m}).

As mentioned in section 3.1, although this narrow conditional prior may be helpful in estimating a single solution for Ic\mathbf{I}_{c} given Im\mathbf{I}_{m} during testing during testing, this is poor for sampling a diversity of solutions. In our framework, the (unconditional) latent priors are imposed for the partial images themselves, which prevent this delta function degeneracy.

B.1.4 CVAE with Fixed Prior

An alternative CVAE variant assumes that conditional prior is independent of the Im\mathbf{I}_{m} and fixed, so p(zc∣Im)≅p(zc)p(\mathbf{z}_{c}|\mathbf{I}_{m})\cong p(\mathbf{z}_{c}), where p(zc)p(\mathbf{z}_{c}) is a fixed distribution (e.g. standard normal). This means

Now we can consider the case for a fixed Im=Im∗\mathbf{I}_{m}=\mathbf{I}^{*}_{m}, and rewrite (B.8) as

Doing so makes it obvious we can then derive the standard (unconditional) VAE formulation from here. Thus an appropriate interpretation of this CVAE variant is that it uses Im\mathbf{I}_{m} as a “switch” parameeter to choose between different VAE models that are trained for the specific conditions.

Once again, this is fine if there are multiple training instances per conditional label. However, in the image completion problem, there is only one Ic\mathbf{I}_{c} per unique Im\mathbf{I}_{m}, so the condition-specific VAE model will simply ignore the sampling “noise” and learn to predict the single instance of Ic\mathbf{I}_{c} from Im\mathbf{I}_{m} directly, i.e. p(Ic∣zc,Im)≈p(Ic∣Im)p(\mathbf{I}_{c}|\mathbf{z}_{c},\mathbf{I}_{m})\approx p(\mathbf{I}_{c}|\mathbf{I}_{m}), which incidentally achieves equality for the variational lower bound. This results in negligible variation of output despite now sampling from p(zc)=N(0,1)p(\mathbf{z}_{c})=\mathcal{N}(0,1).

Our framework resolves this in part by defining all (unconditional) partial images of Ic\mathbf{I}_{c} as sharing a common latent space with adaptive priors, with the likelihood parameters learned as an unconditional VAE, and further coupling on the conditional portion (i.e. the generative path) to get a more distinct but regularized estimate for p(zc∣Im)p(\mathbf{z}_{c}|\mathbf{I}_{m}).

B.2 Joint Maximization of Unconditional and Conditional Variational Lower Bounds

The overall training loss function (5) used in our framework has a direct link to jointly maximizing the unconditional and unconditional variational lower bounds, respectively expressed by (3.1) and (3.1). Using simplified notation, we rewrite these bounds respectively as:

To clarify, B1\mathcal{B}_{1} is the lower bound related to the unconditional log likelihood of observing Ic\mathbf{I}_{c}, while B2\mathcal{B}_{2} relates to the log likelihood of observing Ic\mathbf{I}_{c} conditioned on Im\mathbf{I}_{m}. The expression of B2\mathcal{B}_{2} reflects a blend of conditional likelihood formulations with and without the use of importance sampling, which are matched to different likelihood models, as explained in section 3.1. Note that the (1−λ)(1-\lambda) coefficient from (3.1) is left out here for simplicity, but there is no loss of generality since we can ignore a constant factor of the true lower bound if we are simply maximizing it.

We can then define a combined objective function as our maximization goal

To understand the relation between B\mathcal{B} in (B.11) and L\mathcal{L} in (5), we consider the equivalence of:

For the generative path that involves sampling from the conditional prior pϕ(zc∣Im)p_{\phi}(\mathbf{z}_{c}|\mathbf{I}_{m}), we have the generative log likelihood formulation as

Appendix C C Architectural Details

Our pluralistic image completion network (PICNet) architecture is inspired by SA-GAN and BigGAN, but features several important modifications that enable us to train for this image-conditional generation task. We first replace the batch normalization with instance normalization in the generation network (ResBlock up in Fig. C.7), and remove the batch normalization in our other networks, (i.e. the representation, inference and discriminator networks comprising ResBlock start and ResBlock in Fig. C.7), because different holes will affect the means and variances in each batch. ResBlock down is similar to ResBlock, in which we add the average pooling layer after Conv3×33\times 3 and Conv1×11\times 1.

The Infer1 network only consists of one Residual Block, for self-inferring the latent distribution of the ground truth Ic\mathbf{I}_{c} (treated as known in the reconstructive path), while the Infer2 network consists of seven Residual Blocks, which are applied to predict the latent distribution of Ic\mathbf{I}_{c} (treated as unknown in the generative path) based on the visible pixels Im\mathbf{I}_{m}.

Appendix D D Experimental Details

Our network is implemented in Pytorch v0.4.0, and employs the architectures of Appendix C. To reduce memory cost, we restrained the feature channel width to 4 ⋅ ch4~{}\cdot~{}ch and selected ch=32ch=32. We experimented with different channels with largest being 16 ⋅ ch=102416~{}\cdot~{}ch=1024, but found that the improvement was not obvious. In addition, we applied the self-attention layer of the discriminator and the short+long term attention layer of the generator on a 32×3232\times 32 feature size. Spectral Normalization is used in all networks. All networks are initialized with Orthogonal Initialization and trained from scratch with a fixed learning rate of λ=10−4\lambda=10^{-4}. We used the Adam optimizer with β1=0\beta_{1}=0 and β2=0.999\beta_{2}=0.999.

The final weights we used were αKL=αapp\alpha_{\text{KL}}=\alpha_{\text{app}}=20, αad\alpha_{\text{ad}}=1. The KL loss and appearance matching loss weights come from the variational lower bound. Since the appearance matching loss is used in four output scales, the final weight for the KL loss is αKL=αKL×Nscale\alpha_{\text{KL}}=\alpha_{\text{KL}}\times N_{\text{scale}}, where NscaleN_{\text{scale}} is the number of output scales. We also tried different values of αKL\alpha_{\text{KL}} and αapp\alpha_{\text{app}}, and found that the bigger the KL loss weight, the greater the diversity of the generated Ic′\mathbf{I}_{c}^{{}^{\prime}}, but it was also harder to retain the appearance consistency of the generated Ic′\mathbf{I}_{c}^{{}^{\prime}} to the visible region Im\mathbf{I}_{m}. The values of αapp\alpha_{\text{app}} and αad\alpha_{\text{ad}} were obtained from α\alpha-GAN. We experimented with the number of DD steps per GG step (varying it from 1 to 5), and found that one DD step per GG step gave the best results. When αapp\alpha_{\text{app}} is smaller than 1, we can use two or four DD steps per GG step, but the full generated Ig′\mathbf{I}_{g}^{{}^{\prime}} does not reconstruct the original conditional visible regions Im\mathbf{I}_{m} well. When αapp\alpha_{\text{app}} is larger than 100, we needed two or four GG steps per DD step, if not the discriminator loss will become zero and the generated Ic′\mathbf{I}_{c}^{{}^{\prime}} will be blurry.

We trained each model on a single GPU, with a batch size of 20 on a GTX 1080TI (11GB) and 32 on a NVIDIA V100 (16GB). Training models for centered holes of Paris and CelebA-HQ takes roughly 3 days, while for ImageNet and Places2 it takes roughly 2 weeks. On the other hand, training models for random irregular and un-centered holes takes about twice the time compared to models for centered holes. Moreover, since the prior distribution of random holes p(z)=Nm(0,σ2(n)I)p(\mathbf{z})=\mathcal{N}_{m}(\mathbf{0},\sigma^{2}(n)\mathbf{I}) is changed with the number of pixels in each hole nn, the training loss may sometimes change abruptly due to the KL loss component.