Zero-1-to-3: Zero-shot One Image to 3D Object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, Carl Vondrick
Introduction
From just a single camera view, humans are often able to imagine an object’s 3D shape and appearance. This ability is important for everyday tasks, such as object manipulation and navigation in complex environments , but is also key for visual creativity, such as painting . While this ability can be partially explained by reliance on geometric priors like symmetry, we seem to be able to generalize to much more challenging objects that break physical and geometric constraints with ease. In fact, we can predict the 3D shape of objects that do not (or even cannot) exist in the physical world (see third column in Figure 1). To achieve this degree of generalization, humans rely on prior knowledge accumulated through a lifetime of visual exploration.
In contrast, most existing approaches for 3D image reconstruction operate in a closed-world setting due to their reliance on expensive 3D annotations (e.g. CAD models) or category-specific priors . Very recently, several methods have made major strides in the direction of open-world 3D reconstruction by pre-training on large-scale, diverse datasets such as CO3D . However, these approaches often still require geometry-related information for training, such as stereo views or camera poses. As a result, the scale and diversity of the data they use remain insignificant compared to the recent Internet-scale text-image collections that enable the success of large diffusion models . It has been shown that Internet-scale pre-training endows these models with rich semantic priors, but the extent to which they capture geometric information remains largely unexplored.
In this paper, we demonstrate that we are able to learn control mechanisms that manipulate the camera viewpoint in large-scale diffusion models, such as Stable Diffusion , in order to perform zero-shot novel view synthesis and 3D shape reconstruction. Given a single RGB image, both of these tasks are severely under-constrained. However, due to the scale of training data available to modern generative models (over 5 billion images), diffusion models are state-of-the-art representations for the natural image distribution, with support that covers a vast number of objects from many viewpoints. Although they are trained on 2D monocular images without any camera correspondences, we can fine-tune the model to learn controls for relative camera rotation and translation during the generation process. These controls allow us to encode arbitrary images that are decoded to a different camera viewpoint of our choosing. Figure 1 shows several examples of our results.
The primary contribution of this paper is to demonstrate that large diffusion models have learned rich 3D priors about the visual world, even though they are only trained on 2D images. We also demonstrate state-of-the-art results for novel view synthesis and state-of-the-art results for zero-shot 3D reconstruction of objects, both from a single RGB image. We begin by briefly reviewing related work in Section 2. In Section 3, we describe our approach to learn controls for camera extrinsics by fine-tuning large diffusion models. Finally, in Section 4, we present several quantitative and qualitative experiments to evaluate zero-shot view synthesis and 3D reconstruction of geometry and appearance from a single image. We will release all code and models as well as an online demo.
Related Work
3D generative models. Recent advancements in generative image architectures combined with large scale image-text datasets have made it possible to synthesize high-fidelity of diverse scenes and objects . In particular, diffusion models have shown to be very effective at learning scalable image generators using a denoising objective . However, scaling them to the 3D domain would require large amounts of expensive annotated 3D data. Instead, recent approaches rely on transferring pre-trained large-scale 2D diffusion models to 3D without using any ground truth 3D data. Neural Radiance Fields or NeRFs have emerged as a powerful representation, thanks to their ability to encode scenes with high fidelity. Typically, NeRF is used for single-scene reconstruction, where many posed images covering the entire scene are provided. The task is then to predict novel views from unobserved angles. DreamFields has shown that NeRF is a more versatile tool that can also be used as the main component in a 3D generative system. Various follow-up works substitute CLIP for a distillation loss from a 2D diffusion model that is repurposed to generate high-fidelity 3D objects and scenes from text inputs.
Our work explores an unconventional approach to novel-view synthesis, modeling it as a viewpoint-conditioned image-to-image translation task with diffusion models. The learned model can also be combined with 3D distillation to reconstruct 3D shape from a single image. Prior work adopted a similar pipeline but did not demonstrate zero-shot generalization capability. Concurrent approaches proposed similar techniques to perform image-to-3D generation using language-guided priors and textual inversion . In comparison, our method learns control of viewpoints through a synthetic dataset and demonstrates zero-shot generalization to in-the-wild images.
Single-view object reconstruction. Reconstructing 3D objects from a single view is a highly challenging problem that requires strong priors. One line of work builds priors from relying on collections of 3D primitives represented as meshes , voxels , or point clouds , and use image encoders for conditioning. These models are constrained by the variety of the used 3D data collection and show poor generalization capabilities due to the global nature of this type of conditioning. Moreover, they require an additional pose estimation step to ensure alignment between the estimated shape and the input. On the other hand, locally conditioned models aim to use local image features directly for scene reconstruction and show greater cross-domain generalization capabilities, though are generally limited to close-by view reconstructions. Recently, MCC learns a general-purpose representation for 3D reconstruction from RGB-D views and is trained on large-scale dataset of object-centric videos.
In our work, we demonstrate that rich geometric information can be extracted directly from a pre-trained Stable Diffusion model, alleviating the need for additional depth information.
Method
where we denote as the synthesized image. We want our estimated to be perceptually similar to the true but unobserved novel view .
Novel view synthesis from monocular RGB image is severely under-constrained. Our approach will capitalize on large diffusion models, such as Stable Diffusion, in order to perform this task, since they show extraordinary zero-shot abilities when generating diverse images from text descriptions. Due to the scale of their training data , pre-trained diffusion models are state-of-the-art representations for the natural image distribution today.
However, there are two challenges that we must overcome to create . Firstly, although large-scale generative models are trained on a large variety of objects in different viewpoints, the representations do not explicitly encode the correspondences between viewpoints. Secondly, generative models inherit viewpoint biases reflected on the Internet. As shown in Figure 2, Stable Diffusion tends to generate images of forward-facing chairs in canonical poses. These two problems greatly hinder the ability to extract 3D knowledge from large-scale diffusion models.
Since diffusion models have been trained on internet-scale data, their support of the natural image distribution likely covers most viewpoints for most objects, but these viewpoints cannot be controlled in the pre-trained models. Once we are able to teach the model a mechanism to control the camera extrinsics with which a photo is captured, then we unlock the ability to perform novel view synthesis.
To this end, given a dataset of paired images and their relative camera extrinsics , our approach, shown in Figure 3, fine-tunes a pre-trained diffusion model in order to learn controls over the camera parameters without destroying the rest of the representation. Following , we use a latent diffusion architecture with an encoder , a denoiser U-Net , and a decoder . At the diffusion time step , let be the embedding of the input view and relative camera extrinsics. We then solve for the following objective to fine-tune the model:
After the model is trained, the inference model can generate an image by performing iterative denoising from a Gaussian noise image conditioned on .
The main result of this paper is that fine-tuning pre-trained diffusion models in this way enables them to learn a generic mechanism for controlling the camera viewpoints, which extrapolates outside of the objects seen in the fine-tuning dataset. In other words, this fine-tuning allows controls to be “bolted on” and the diffusion model can retain the ability to generate photorealistic images, except now with control of viewpoints. This compositionality establishes zero-shot capabilities in the model, where the final model can synthesize new views for object classes that lack 3D assets and never appear in the fine-tuning set.
2 View-Conditioned Diffusion
3D reconstruction from a single image requires both low-level perception (depth, shading, texture, etc.) and high-level understanding (type, function, structure, etc.). Therefore, we adopt a hybrid conditioning mechanism. On one stream, a CLIP embedding of the input image is concatenated with to form a “posed CLIP” embedding . We apply cross-attention to condition the denoising U-Net, which provides high-level semantic information of the input image. On the other stream, the input image is channel-concatenated with the image being denoised, assisting the model in keeping the identity and details of the object being synthesized. To be able to apply classifier-free guidance , we follow a similar mechanism proposed in , setting the input image and the posed CLIP embedding to a null vector randomly, and scaling the conditional information during inference.
3 3D Reconstruction
In many applications, synthesizing novel views of an object is not enough. A full 3D reconstruction capturing both the appearance and geometry of an object is desired. We adopt a recently open-sourced framework, Score Jacobian Chaining (SJC) , to optimize a 3D representation with priors from text-to-image diffusion models. However, due to the probabilistic nature of diffusion models, gradient updates are highly stochastic. A crucial technique used in SJC, inspired by DreamFusion , is to set the classifier-free guidance value to be significantly higher than usual. This methodology decreases the diversity of each sample but improves the fidelity of the reconstruction.
As shown in Figure 4, similarly to SJC, we randomly sample viewpoints and perform volumetric rendering. We then perturb the resulting images with Gaussian noise , and denoise them by applying the U-Net conditioned on the input image , posed CLIP embedding , and timestep , in order to approximate the score toward the non-noisy input :
where is the PAAS score introduced by .
In addition, we optimize the input view with an MSE loss. To further regularize the NeRF representation, we also apply a depth smoothness loss to every sampled viewpoint, and a near-view consistency loss to regularize the change in appearance between nearby views.
4 Dataset
We use the recently released Objaverse dataset for fine-tuning, which is a large-scale open-source dataset containing 800K+ 3D models created by 100K+ artists. While it has no explicit class labels like ShapeNet , Objaverse embodies a large diversity of high-quality 3D models with rich geometry, many of them with fine-grained details and material properties. For each object in the dataset, we randomly sample 12 camera extrinsics matrices pointing at the center of the object and render 12 views with a ray-tracing engine. At training time, two views can be sampled for each object to form an image pair . The corresponding relative viewpoint transformation that defines the mapping between both perspectives can easily be derived from the two extrinsic matrices.
Experiments
We assess our model’s performance on zero-shot novel view synthesis and 3D reconstruction. As confirmed by the authors of Objaverse, the datasets and images we used in this paper are outside of the Objaverse dataset, and can thus be considered zero-shot results. We quantitatively compare our model to the state-of-the-art on synthetic objects and scenes with different levels of complexity. We also report qualitative results using diverse in-the-wild images, ranging from pictures we took of daily objects to paintings.
We describe two closely related tasks that take single-view RGB images as input, and we apply them zero-shot.
Novel view synthesis. Novel view synthesis is a long-standing 3D problem in computer vision that requires a model to learn the depth, texture, and shape of an object implicitly. The extremely limited input information of only a single view requires a novel view synthesis method to leverage prior knowledge. Recent popular methods have relied on optimizing implicit neural fields with CLIP consistency objectives from randomly sampled views . Our approach for view-conditional image generation is orthogonal, however, because we invert the order of 3D reconstruction and novel view synthesis, while still retaining the identity of the object depicted in the input image. This way, the aleatoric uncertainty due to self-occlusion can be modeled by a probabilistic generative model when rotating around objects, and the semantic and geometric priors learned by large diffusion models can be leveraged effectively.
3D Reconstruction. We can also adapt a stochastic 3D reconstruction framework such as SJC or DreamFusion to create a most likely 3D representation. We parameterize this as a voxel radiance field , and subsequently extract a mesh by performing marching cubes on the density field. The application of our view-conditioned diffusion model for 3D reconstruction provides a viable path to channel the rich 2D appearance priors learned by our diffusion model toward 3D geometry.
2 Baselines
To be consistent with the scope of our method, we compare only to methods that operate in a zero-shot setting and use single-view RGB images as input.
For novel view synthesis, we compare against several state-of-the-art, single-image algorithms. In particular, we benchmark DietNeRF , which regularizes NeRF with a CLIP image-to-image consistency loss across viewpoints. In addition, we compare with Image Variations (IV) , which is a Stable Diffusion model fine-tuned to be conditioned on images instead of text prompts and could be seen as a semantic nearest-neighbor search engine with Stable Diffusion. Finally, we adapted SJC , a diffusion-based text-to-3D model where the original text-conditioned diffusion model is replaced with an image-conditioned diffusion model, which we termed SJC-I.
For 3D reconstruction, we use two state-of-the-art, single-view algorithms as baselines: (1) Multiview Compressive Coding (MCC) , which is a neural field-based approach that completes RGB-D observations into a 3D representation, as well as (2) Point-E , which is a diffusion model over colorized point clouds. MCC is trained on CO3Dv2 , while Point-E is notably trained on a significantly bigger OpenAI’s internal 3D dataset. We also compare against SJC-I.
Since MCC requires depth input, we use MiDaS off-the-shelf for depth estimation. We convert the obtained relative disparity map to an absolute pseudo-metric depth map by assuming standard scale and shift values that look reasonable across the entire test set.
3 Benchmarks and Metrics
We evaluate both tasks on Google Scanned Objects (GSO) , which is a dataset of high-quality scanned household items, as well as RTMV , which consists of complex scenes, each composed of 20 random objects. In all experiments, the respective ground truth 3D models are used for evaluating 3D reconstruction.
For novel view synthesis, we numerically evaluate our method and baselines extensively with four metrics covering different aspects of image similarity: PSNR, SSIM , LPIPS , and FID . For 3D reconstruction, we measure Chamfer Distance and volumetric IoU.
4 Novel View Synthesis Results
We show the numerical results in Tables 1 and 2. Figure 5 shows that our method, as compared to all baselines on GSO, is able to generate highly photorealistic images that are closely consistent with the ground truth. Such a trend can also be found on RTMV in Figure 6, even though the scenes are out-of-distribution compared to the Objaverse dataset. Among our baselines, we observed that Point-E tends to achieve much better results than other baselines, maintaining impressive zero-shot generalizability. However, the small size of the generated point clouds greatly limits the applicability of Point-E for novel view synthesis.
In Figure 7, we further demonstrate the generalization performance of our model to objects with challenging geometry and texture as well as its ability to synthesize high-fidelity viewpoints while maintaining the object type, identity and low-level details.
Diversity across samples. Novel view synthesis from a single image is a severely under-constrained task, which makes diffusion models a particularly apt choice of architecture compared to NeRF in terms of capturing the underlying uncertainty. Because input images are 2D, they always depict only a partial view of the object, leaving many parts unobserved. Figure 8 exemplifies the diversity of plausible, high-quality images sampled from novel viewpoints.
5 3D Reconstruction Results
We show numerical results in Tables 3 and 4. Figure 9 qualitatively shows our method reconstructs high-fidelity 3D meshes that are consistent with the ground truth. MCC tends to give a good estimation of surfaces that are visible from the input view, but often fails to correctly infer the geometry at the back of the object.
SJC-I is also frequently unable to reconstruct a meaningful geometry. On the other hand, Point-E has an impressive zero-shot generalization ability, and is able to predict a reasonable estimate of object geometry. However, it generates non-uniform sparse point clouds of only 4,096 points, which sometimes leads to holes in the reconstructed surfaces (according to their provided mesh conversion method). Therefore, it obtains a good CD score but falls short of the volumetric IoU. Our method leverages the learned multi-view priors from our view-conditioned diffusion model and combines them with the advantages of a NeRF-style representation. Both factors provide improvements in terms of CD and volumetric IoU over prior works, as indicated by Tables 3 and 4.
6 Text to Image to 3D
In addition to in-the-wild images, we also tested our method on images generated by txt2img models such as Dall-E-2 . As shown in Figure 10, our model is able to generate novel views of these images while preserving the identity of the objects. We believe this could be very useful in many text-to-3D generation applications.
Discussion
In this work, we have proposed a novel approach, Zero-1-to-3, for zero-shot, single-image novel-view synthesis and 3D reconstruction. Our method capitalizes on the Stable Diffusion model, which is pre-trained on internet-scaled data and captures rich semantic and geometric priors. To extract this information, we have fine-tuned the model on synthetic data to learn control over the camera viewpoint. The resulting method demonstrated state-of-the-art results on several benchmarks due to its ability to leverage strong object shape priors learned by Stable Diffusion.
From objects to scenes. Our approach is trained on a dataset of single objects on a plain background. While we have demonstrated a strong degree of generalizations to scenes with several objects on RTMV dataset, the quality still degrades compared to the in-distribution samples from GSO. Generalization to scenes with complex backgrounds thus remains an important challenge for our method.
From scenes to videos. Being able to reason about geometry of dynamic scenes from a single view would open novel research directions, such as understanding occlusions and dynamic object manipulation. A few approaches for diffusion-based video generation have been proposed recently , and extending them to 3D would be key to opening up these opportunities.
Combining graphics pipelines with Stable Diffusion. In this paper, we demonstrate a framework to extract 3D knowledge of objects from Stable Diffusion. A powerful natural image generative model like Stable Diffusion contains other implicit knowledge about lighting, shading, texture, etc. Future work can explore similar mechanisms to perform traditional graphics tasks, such as scene relighting.
Acknowledgements: We would like to thank Changxi Zheng and Samir Gadre for their helpful feedback. We would also like to thank the authors of SJC , NeRDi , SparseFusion , and Objaverse for their helpful discussions. This research is based on work partially supported by the Toyota Research Institute, the DARPA MCS program under Federal Agreement No. N660011924032, and the NSF NRI Award #1925157.
References
Appendix A Coordinate System & Camera Model
We use a spherical coordinate system to represent camera locations and their relative transformations. As shown in Figure 11, assuming the center of the object is the origin of the coordinate system, we can use , , and to represent the polar angle, azimuth angle, and radius (distance away from the center) respectively. For the creation of the dataset, we normalize all assets to be contained inside the XYZ unit cube . Then, we sample camera viewpoints such that , uniformly cover the unit sphere, and is sampled uniformly in the interval . During training, when two images from different viewpoints are sampled, let their camera locations be and . We denote their relative camera transformation as . Since the camera is always pointed at the center of the coordinate system, the extrinsics matrices are uniquely defined by the location of the camera in a spherical coordinate system. We assume the horizontal field of view of the camera to be , and follow a pinhole camera model.
Due to the incontinuity of the azimuth angle, we encode it with . Subsequently, at both training and inference time, four values representing the relative camera viewpoint change, are fed to the model, along with an input image, in order to generate the novel view.
Appendix B Dataset Creation
We use Blender to render training images of the finetuning dataset. The specific rendering code is inherited from a publicly released repositoryhttps://github.com/allenai/objaverse-rendering by authors of Objaverse . For each object in Objaverse, we randomly sample 12 views and use the Cycles engine in Blender with 128 samples per ray along with a denoising step to render each image. We render all images in 512512 resolution and pad transparent backgrounds with white color. We also apply randomized area lighting. In total, we rendered a dataset of around 10M images for finetuning.
Appendix C Finetuning Stable Diffusion
We use the rendered dataset to finetune a pretrained Stable Diffusion model for performing novel view synthesis. Since the original Stable Diffusion network is not conditioned on multimodal text embeddings, the original Stable Diffusion architecture needs to be tweaked and finetuned to be able to take conditional information from an image. This is done in , and we use their released checkpoints. To further adapt the model to accept conditional information from an image along with a relative camera pose, we concatenate the image CLIP embedding (dimension 768) and the pose vector (dimension 4) and initialize another fully-connected layer (772 768) to ensure compatibility with the diffusion model architecture. The learning rate of this layer is scaled up to be 10 larger than the other layers. The rest of the network architecture is kept the same as the original Stable Diffusion.
We use AdamW with a learning rate of for training. First, we attempted a batch size of 192 while maintaining the original resolution (image dimension , latent dimension ) for training. However, we discovered that this led to a slower convergence rate and higher variance across batches. Because the original Stable Diffusion training procedure used a batch size of 3072, we subsequently reduce the image size to (and thus the corresponding latent dimension to ), in order to be able to increase the batch size to 1536. This increase in batch size has led to better training stability and a significantly improved convergence rate. We finetuned our model on an 8A100-80GB machine for 7 days.
C.2 Inference Details
To generate a novel view, Zero-1-to-3 takes only 2 seconds on an RTX A6000 GPU. Note that in prior works, typically a NeRF is trained in order to render novel views, which takes significantly longer. In comparison, our approach inverts the order of 3D reconstruction and novel view synthesis, causing the novel view synthesis process to be fast and contain diversity under uncertainty. Since this paper addresses the problem of a single image to a 3D object, when an in-the-wild image is used during inference, we apply an off-the-shelf background removal tool to every image before using it as input to Zero-1-to-3.
Appendix D 3D Reconstruction
Different from the original Score Jacobian Chaining (SJC) implementation, we removed the “emptiness loss” and “center loss”. To regularize the VoxelRF representation, we differentiably render a depth map, and apply a smoothness loss to the depth map. This is based on the prior knowledge that the geometry of an object typically contains less high-frequency information than its texture. It is particularly helpful in removing holes in the object representation. We also apply a near-view consistency loss to regularize the difference between an image rendered from one view and another image rendered from a nearby randomly sampled view. We found this to be very helpful in improving the cross-view consistency of an object’s texture. All implementation details can be found in the code that is submitted as part of the appendix. Running a full 3D reconstruction on an image takes around 30 minutes on an RTX A6000 GPU.
We extract the 3D mesh from the VoxelRF representation as follows. We first query the density grids at resolution . Then we smooth the density grids using a mean filter of size , followed by an erosion operator of size . Finally, we run marching cubes on the resulting density grids. Let denote the average value of the density grids. For the GSO dataset, we use a density threshold of . For the RTMV dataset, we use a density threshold of .
Evaluation.
The ground truth 3D shape and the predicted 3D shape are first normalized within the unit cube. To compute the chamfer distance (CD), we randomly sample 2000 points. For Point-E and MCC, we sample from their predicted point clouds directly. For our method and SJC-I, we sample points from the reconstructed 3D mesh. We compute the volumetric IoU at resolution . For our method, Point-E and SJC-I, we vocalize the reconstructed 3D surface meshes using marching cubes. For MCC, we directly voxelize the predicted dense point clouds by occupancy.
Appendix E Baselines
To be consistent with the scope of our method, we compare only to methods that (1) operate in a zero-shot setting, (2) use single-view RGB images as input, and (3) have official reference implementations available online that can be adapted in a reasonable timeframe. In the following sections, we describe the implementation details of our baselines.
We use the official implementation located on GitHubhttps://github.com/ajayjain/DietNeRF, which, at the time of writing, has code for low-view NeRF optimization from scratch with a joint MSE and consistency loss, though provides no functionality related to finetuning PixelNeRF. For fairness, we use the same hyperparameters as the experiments performed with the NeRF synthetic dataset in . For the evaluation of novel view synthesis, we render the resulting NeRF from the designated camera poses in the test set.
E.2 Point-E
We use the official implementation and pretrained models located on GitHubhttps://github.com/openai/point-e. We keep all the hyperparameters and follow their demo example to do 3D reconstruction from single input image. The prediction is already normalized, so we do not need to perform any rescaling to match the ground truth. For surface mesh extraction, we use their default method with a grid size of 128.
E.3 MCC
We use the official implementation located on GitHubhttps://github.com/facebookresearch/MCC. Since this approach requires a colorized point cloud as input rather than an RGB image, we first apply an online off-the-self foreground segmentation method as well as a state-of-the-art depth estimation method for preprocessing. For fairness, we keep all hyperparameters the same as the zero-shot, in-the-wild experiments described in . For the evaluation of 3D reconstruction, we normalize the prediction, rotate it according to camera extrinsics, and compare it with the 3D ground truth.