RealFill: Reference-Driven Generation for Authentic Image Completion
Luming Tang, Nataniel Ruiz, Qinghao Chu, Yuanzhen Li, Aleksander Holynski, David E. Jacobs, Bharath Hariharan, Yael Pritch, Neal Wadhwa, Kfir Aberman, Michael Rubinstein
Introduction
Photographs capture frozen moments in time corresponding to ephemeral and invaluable experiences in our lives, but can sometimes fail to do these memories justice. In many cases, no single shot may have captured the perfect angle, framing, timing, and composition, and unfortunately, just as the experiences themselves cannot be revisited, these elements of the captured images are similarly unalterable. We show one such example in Fig. 2: imagine having taken a nearly perfect photo of your daughter dancing on stage, but her unique and intricate crown is only barely cut out of the frame. Of course, there are many other pictures from the performance that showcase her crown, but they all fail to capture that precise special moment: her pose mid-dance, her facial expression, and the perfect lighting. Given your memories of this event and this collection of imperfect photos, you can certainly imagine the missing parts of this perfect shot, but actually creating a complete version of this image, e.g., to share with family and friends, is a much harder task.
In this paper, we focus on this problem, which we call Authentic Image Completion. Given a few reference images (up to five) and one target image that captures roughly the same scene (but in a different arrangement or appearance), we aim to fill missing regions of the target image with high-quality image content that is faithful to the originally captured scene. Note that for the sake of practical benefit, we focus particularly on the more challenging, unconstrained setting in which the target and reference images may have very different viewpoints, environmental conditions, camera apertures, image styles, or even moving objects.
Approaches to solve variants of this problem have been proposed using classical geometry-based pipelines that rely on correspondence matching, depth estimation, and 3D transformations, followed by patch fusion and image harmonization. These methods tend to encounter catastrophic failure when the scene’s structure cannot be accurately estimated, e.g., when the scene geometry is too complex or contains dynamic objects. On the other hand, recent generative models , and in particular diffusion models , have demonstrated strong performance on the tasks of image inpainting and outpainting . These methods, however, struggle to recover the genuine scene structure and fine details, since they are only guided by text prompts, and therefore lack a mechanism for utilizing reference image content.
To this end, we present a simple yet effective reference-driven image completion framework called RealFill. For a given scene, we first create a personalized generative model by fine-tuning a pre-trained inpainting diffusion model on the reference and target images. This fine-tuning process is designed such that the adapted model not only maintains a good image prior, but also learns the contents, lighting, and style of the scene in the input images. We then use this fine-tuned model to fill the missing regions in the target image through a standard diffusion sampling process. Given the stochastic nature of generative inference, we also propose Correspondence-Based Seed Selection, to automatically select a small set of high-quality generations by exploiting a special property of our completion task: the fact that there should exist true correspondence between our generated content and our reference images. Specifically, we filter out samples that have too few keypoint correspondences with our reference images, a filtering process that greatly limits the need for human intervention in selecting high-quality model outputs.
As shown in Fig 1, 3, and 4, RealFill is able to very effectively inpaint or outpaint a target image with its genuine scene content. Most importantly, our method is able to handle large differences between reference and target images, e.g., viewpoint, lighting, aperture, style or dynamic deformations — differences which are very difficult for previous geometry-based approaches. Existing benchmarks for image completion mainly focus on small inpainting tasks and minimal changes between reference and target images. In order to quantitatively evaluate the aforementioned challenging use-case, we collect a dataset containing 10 inpainting and 23 outpainting examples along with corresponding ground-truth, and show that RealFill outperforms baselines by a large margin across multiple image similarity metrics.
In summary, in our work, we propose the following contributions:
We define a new problem, named Authentic Image Completion, where, given a set of reference images and a target image with missing regions, we seek to complete those missing regions with content faithful to the scene observed in the reference images. In essence, the goal is to complete the target image with what “should have been there” rather than what “could have been there”, as is often the case in typical generative inpainting.
We introduce RealFill, a method that aims to solve this problem by finetuning a diffusion-based text-to-image inpainting model on reference and target images. This model is sampled with Correspondence-Based Seed Selection to filter ouptuts with low fidelity to the reference images. RealFill is the first method that expands the expressive power of generative inpainting models by conditioning the process on more than text (i.e., by adding reference images).
We propose RealBench, a dataset for quantitative and qualitative evaluation of authentic image completion, composed of 33 scenes spanning both inpainting and outpainting tasks.
Related Work
Adapting Pre-trained Diffusion Models. Diffusion models have demonstrated superior performance in text-to-image (T2I) generation . Recent works take advantage of this useful pre-trained image prior by fine-tuning these models, either for added controllability, personalization, or for specialized tasks. Personalization methods propose to fine-tune the T2I model on a few choice images in order to achieve subject-driven generation, allowing for arbitrary text-driven generation of a given object or style. Other techniques instead fine-tune a T2I model to add new conditioning signals, either for image editing or more controllable image generation . This same approach has been shown to be useful for specialized tasks such as adding camera viewpoint conditioning, e.g., to aid in text-to-3D generation or converting a T2I model into a generative video model. Our method shows that a pre-trained T2I inpainting diffusion model can be adapted to perform reference-driven image completion.
Image Completion. An enduring challenge in computer vision, image completion aims to fill missing parts of an image with plausible content. This task is interchangeably referred to as inpainting or outpainting depending on the characteristics of the missing region. Traditional approaches to this problem rely on handcrafted heuristics while more recent deep learning based methods instead directly train end-to-end neural networks that take original image and mask as inputs and generate the completed image. Given the challenging nature of this problem , many works propose to leverage the image prior from a pre-trained generative model for this task. Built upon powerful T2I diffusion models, recent diffusion-based solutions demonstrate strong text-driven image completion capabilities. However, due to their sole dependence on a text prompt (which has limited descriptive power), generated image content can often be hard to control, resulting in tedious prompt tuning, especially when a particular or otherwise true scene content is desired. This is one of the main issues we aim to tackle in our work.
Reference Based Image Inpainting. Existing work for reference-based inpainting or outpainting usually make use of carefully tuned pipelines containing many individual components like depth and pose estimation, image warping, and harmonization. Each of these modules usually tackles a moderately challenging problem itself and the resulting prediction error can, and often does, propagate and accumulate through the pipeline. This can lead to catastrophic failure especially in challenging cases with complex scene geometry, changes in appearance, or scene deformation. Paint-by-Example propose to fine-tune a latent diffusion model such that the generation is conditioned on both a reference and target image. However, the image conditioning is based on a CLIP embedding of a single reference image, and is therefore only able to capture high-level semantic information of the reference object. In contrast, our method is the first to demonstrate multiple reference image-driven inpainting and outpainting that is both visually compelling and faithful to the original scene, even in cases where there are large appearance changes between reference and target images.
Method
Given a set of casually captured reference images (up to five), our goal is to complete (i.e., either outpaint or inpaint) a target image of roughly the same scene. The output image is expected to not only be plausible and photorealistic, but to also be faithful to the reference images — recovering content and scene detail that was present in the actual scene. In essence, we want to achieve authentic image completion, where we generate what “should have been there” instead of what “could have been there”. We purposefully pose this as a broad and challenging problem with few constraints on the inputs. For example, the images could be taken from very different viewpoints with unknown camera poses. They could also have different lighting conditions or styles, and the scene could potentially be non-static and have significantly varying layout across images.
In this section, we first provide background knowledge on diffusion models and subject-driven generation (Sec. 3.2). Then, we formally define the problem of authentic image completion (Sec. 3.3). Finally, we present RealFill, our method to perform reference-driven image completion with a pre-trained diffusion image prior (Sec. 3.4).
2 Preliminaries
Diffusion models are generative models that aim to transform a Gaussian distribution into an arbitrary target data distribution. During training, different magnitudes of Gaussian noise are added to a clean data point to obtain noisy :
where the noise , and define a fixed noise schedule with larger corresponding to more noise. Then, a neural network is trained to predict the noise using the following loss function:
where the generation is conditioned on some signal , e.g., a language prompt for a text-to-image model, or a masked image for an inpainting model. During inference, starting from , is used to iteratively remove noise from to get a less noisy , eventually leading to a sample from the target data distribution.
3 Problem Setup
4 RealFill
This task is challenging for both geometry-based and reconstruction-based approaches because there are barely any geometric constraints between and , there are only a few images available as inputs, and the reference images may have different styles, lighting conditions, and subject poses from the target. One alternative is to use a controllable inpainting or outpainting methods, however, these methods are either prompt-based or single-image object-driven , which makes them hard to use for recovering complex scene-level structure and details.
To this end, we propose to first fine-tune a pre-trained generative model by injecting knowledge of the scene (from a set of reference images), such that the model is aware of the contents of the scene when generating , conditioned on and .
Training. Starting from a state-of-the-art T2I diffusion inpainting model , we inject LoRA weights and fine-tune it on both and with randomly generated binary masks . The loss function is
where , is a fixed language prompt, denotes the element-wise product and therefore is the masked clean image. For , the loss is only calculated on the existing region, i.e., where ’s entry equals 0. Specifically, we use the open-sourced Stable Diffusion v2 inpainting model and inject LoRA layers into its text encoder and U-Net for fine-tuning. Following , we fix to be a sentence containing a rare token, i.e., “a photo of [V]”. For each training example, similar to , we generate multiple random rectangles and take either their union or the complement of the union to get the final random mask . Our fine-tuning pipeline is illustrated in Fig. 2.
Inference. After training, we use the DDPM sampler to generate an image , conditioning the model on , and . However, similar to the observation in , we notice that the existing region in is distorted in . To resolve this, we first blur the mask , then use it to alpha composite and , leading to the final with full recovery on the existing area and a smooth transition at the boundary of the generated region.
Correspondence-Based Seed Selection. The diffusion inference process is stochastic, i.e., the same input conditioning images may produce any number of generated images depending on the input seed to the sampling process. This stochasticity often results in variance in the quality of generated results, often requiring human intervention to select high-quality samples. While there exists work in identifying good samples from a collection of generated outputs , this remains an open problem. Nevertheless, our proposed problem of authentic image completion is a special case of this more general problem statement. In particular, the reference images provide a grounding signal for the true content of the scene, and can be used to help identify high-quality outputs. Specifically, we find that the number of image feature correspondences between and can be used as a metric to roughly quantify whether the result is faithful to the reference images. We propose Correspondence-Based Seed Selection, a process that consists of generating a batch of outputs, i.e., , extracting a set of correspondences (using LoFTR , for example) between and the filled region of each , (i.e., where ’s entry equals 1), and finally ranking the generated results by the number of matched keypoints. This allows us to automatically filter generations to a small set of high-quality results. Compared to traditional seed selection approaches in other domains, our proposed method greatly alleviates the need for human intervention in selecting best samples.
Experiments
In Fig. 3 and 4, we show that RealFill is able to convincingly outpaint and inpaint image content that is faithful to the reference images. Notably, it is able to handle dramatic differences in camera pose, lighting, defocus blur, image style and even subject pose. This is because RealFill has both a good image prior (from the pre-trained diffusion model) and knowledge of the scene (from fine-tuning on the input images). Thus, it is able to inherit knowledge about the contents of the scene, but generate content that fits seamlessly into the target image.
2 Comparisons
Evaluation Dataset. Existing benchmarks for reference-driven image completion primarily focus on inpainting small regions, and assume at most very minor changes between the reference and target images. To better evaluate our target use-case, we create our own dataset, RealBench. RealBench consists of 33 scenes (23 outpainting and 10 inpainting), where each scene has a set of reference images , a target image to fill, a binary mask indicating the missing region and the ground-truth result . The number of reference images in each scene varies from 1 to 5. The dataset contains diverse, challenging scenarios with significant variations between the reference and target images, such as changes in viewpoint, defocus blur, lighting, style and subject pose.
Evaluation Metrics. We use multiple metrics to evaluate the quality and fidelity of our model outputs. We compare the generated images with the ground-truth target image at multiple levels of image similarity, including PSNR, SSIM, and LPIPS for low-level, DreamSim for mid-level, and DINO and CLIP for high-level.
For low-level metrics, we only calculate a loss on the filled-in region, i.e., where is 1. For high-level image similarity, we use the cosine distance between the full image embeddings from CLIP and DINO. For mid-level similarity, we use the full image embedding using DreamSim which is designed to emphasize differences in image layouts, object poses, and semantic contents.
Baseline Approaches. We compare to two baselines: the exemplar-based image inpainting method Paint-by-Example and the popular prompt-based image filling approach Stable Diffusion Inpainting . Since Paint-by-Example only uses one reference image during generation, we randomly pick a reference image for each run of this baseline. Choosing an appropriate prompt for Stable Diffusion Inpainting is a necessary component of getting a high quality result. So, for a fair comparison, instead of using a generic prompt like “a beautiful photo”, we manually write a long prompt that describes the scene in detail. For example, the prompt for the first row of Fig. 5 is “two men sitting together with a child in the middle, the man on the left is playing guitar, the man on right is wearing a birthday hat with some stickers on it. There is a blue decorator hanging on the wall”.
Implementation Details of RealFill. For each scene, we fine-tune the inpainting diffusion model for 2,000 iterations with a batch size of 16 on a single NVIDIA A100 GPU with LoRA rank 8. With a probability of 0.1, we randomly dropout prompt , mask and LoRA layers independently during training. The learning rate is set to 2e-4 for the U-Net and 4e-5 for the text encoder. Note that these hyper-parameters could be further tuned for each scene to get better performance, e.g., some scenes converge more quickly may overfit if trained for too long. However, for the sake of fair comparison, we use a constant set of hyper-parameters for all results shown in the paper.
Quantitative Comparison. We quantitatively compare our method with the baseline methods. For each method, we report average metrics across all target images , where each image’s metric is itself computed from an average of 64 stochastically generated samples.. In Tab. 1, we report these aggregate metrics and find that RealFill outperforms all baselines by a large margin across all levels of similarity.
Qualitative Comparison. In Fig. 5, we present a visual comparison between RealFill and the baselines. We also show the ground-truth and input images for each example. In order to better highlight the regions which are being generated, we overlay a semi-transparent white mask on the ground truth and output images, covering the known regions of the target image. RealFill not only generates high-quality images, but also more faithfully reproduces the scene than the baseline methods. Paint-by-Example relies on the CLIP embedding of the reference images as the condition. This poses a challenge when dealing with complex scenes or attempting to restore object details, since CLIP embeddings only capture high-level semantic information. The generated results from Stable Diffusion Inpainting are plausible on their own. However, because natural language is limited in conveying complex visual information, they often exhibit substantial deviations from the original scenes depicted in the reference images.
Correspondence-Based Seed Selection. We evaluate the effect of our proposed correspondence-based seed selection described in Sec. 3.4. To measure the correlation between our seed selection mechanism and high-quality results, we rank RealFill’s outputs according to the number of matched keypoints, and then filter out a certain percent of the lowest-ranked samples. We then average the evaluation metrics only across the remaining samples. We find that higher filtering rates like 75% greatly improve the quantitative metrics, when compared to unfiltered results (Tab. 2). In Fig 6, we show multiple RealFill outputs with the corresponding number of matched keypoints. These demonstrate a clear trend, where fewer matches usually indicate lower-quality results.
Discussion
Image Stitching. One straight-forward approach is to utilize the correspondence between reference and target images and stitch them together. However, we find that this does not yield acceptable results most of the time, even using commercial image stitching software, particularly when there are dramatic viewpoint changes, lighting changes, or moving objects. Taking the two scenes in Fig. 7 as example, multiple commercial software solutions produce no output, asserting that the reference and target images do not have sufficient correspondences. On the contrary, RealFill recovers these scenes both faithfully and realistically.
Vanilla DreamBooth. Instead of adapting an inpainting model, another alternative is to fine-tune a standard Stable Diffusion model on the reference images, i.e., vanilla DreamBooth, then use the fine-tuned T2I model to inpaint the target image , as implemented in the popular Diffusers library Diffusers’ Stable Diffusion inpainting pipeline code.. However, because this model is never trained with a masked prediction objective, it performs much worse compared to RealFill, as shown in Fig. 8.
2 What makes RealFill work?
In order to explore why our proposed method leads to strong results, especially on complex scenes, we make the following two hypotheses:
RealFill relates multiple elements in a scene. If we make the conditioning image a blank canvas during inference, i.e., all entries of equal 1, we can see in Fig. 9 that the fine-tuned model is able to generate multiple scene variants with different structures, e.g., removing the foreground or background object, or manipulating the object layouts. This suggest that RealFill is able to relate the elements inside the scene in a compositional way.
RealFill captures correspondences among input images. Even if the reference and target images do not depict the same scene, the fine-tuned model is still able to fuse the corresponding contents of the reference images into the target area seamlessly, as shown in Fig. 10. This suggests that RealFill is able to capture and utilize real or invented correspondences between reference and target images to do generation. Previous works also found similar emergent correspondence inside pre-trained Stable Diffusion models.
3 Limitations
Because RealFill needs to go through a gradient-based fine-tuning process on input images, it is relatively slow and far from real time. Empirically, we also find that, when the viewpoint change between reference and target images is dramatic, RealFill fails to recover the 3D scene faithfully, especially when there’s only a single reference image. For example, as seen in the top row of Fig. 11, the reference image is captured from a side view while the target is from a center view. Although the RealFill output looks plausible at first glance, the pose of the husky is different from the reference, e.g., the left paw should be on the gap between the cushions. Lastly, because RealFill mainly relies on the image prior inherited from the base pre-trained model, it also fails to handle cases where that are challenging for the base model. For instance, Stable Diffusion is known to be less effective when it comes to generating fine image details, such as text, human faces, or body parts. As shown in the bottom row of Fig. 11 where the store sign is wrongly spelled, this is also true for RealFill.
Societal Impact
This research aims to create a tool that can help users express their creativity and improve the quality of their personal photographs through image generation. However, advanced image generation methods can have complex impacts on society. Our proposed method inherits some of the concerns that are associated with this class of technology, such as the potential to alter sensitive personal characteristics. The open source pre-trained model that we use in our work, Stable Diffusion, exhibits some of these concerns. However, we have not found any evidence that our method is more likely to produce biased or harmful content than previous work. Despite these findings, it is important to continue investigating the potential risks of image generation technology. Future research should focus on developing methods to mitigate bias and harmful content, and to ensure that image generation tools are used in a responsible manner.
Conclusion
In this work, we introduce the problem of Authentic Image Completion, where given a few reference images, we intend to complete some missing regions of a target image with the content that “should have been there” — rather that “what could have been there”. To tackle this problem, we proposed a simple yet effective approach called RealFill, which first fine-tunes a T2I inpainting diffusion model on the reference and target images, and then uses the adapted model to fill the missing regions. We show that RealFill produces high-quality image completions that are faithful to the content in the reference images, even when there are large differences between reference and target images such as viewpoint, aperture, lighting, image style and object position, pose and articulation.
Acknowledgements. We would like to thank Rundi Wu, Qianqian Wang, Viraj Shah, Ethan Weber, Zhengqi Li, Kyle Genova, Boyang Deng, Maya Goldenberg, Noah Snavely, Ben Poole, Ben Mildenhall, Alex Rav-Acha, Pratul Srinivasan, Dor Verbin and Jon Barron for their valuable discussion and feedbacks, and thank Zeya Peng, Rundi Wu, Shan Nan for their contribution to the evaluation dataset. A special thanks to Jason Baldridge, Kihyuk Sohn, Kathy Meier-Hellstern, and Nicole Brichtova for their feedback and support for the project.