Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion Prior
Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, Dong Chen
Introduction
Given a single image as in Figure 1, how would the object portrayed in the image look like from a different perspective? Humans possess an innate ability to effortlessly imagine 3D geometry and hallucinate the appearance of novel views with a glance at the picture based on their prior knowledge about the world. In this work, we aim to achieve a similar goal: creating high-fidelity 3D content from a real or artificially generated single image. This will open up new avenues for artistic expression and creativity, such as bringing 3D effects to the fantasy images created by the cutting-edge 2D generative models like Stable Diffusion . By offering a more accessible and automated way to create visually stunning 3D content, we hope to engage a broader audience with the world of 3D modeling with ease.
The creation of 3D objects from a single image presents a significant challenge due to the limited information that can be inferred from a single viewpoint. One categories of works aim to produce 3D photo effect in the manner of image-based rendering or single-view 3D reconstruction with neural rendering . However, these methods often struggle with reconstructing fine geometry and fall short of rendering in large views. Another line of research projects the input image into the latent space of the pretrained 3D-aware generative networks. Despite their impressive performance, existing 3D generative networks mainly model objects from a specific class and are therefore incapable of handling general 3D objects. In our case, we aim for general 3D creation from an arbitrary image, yet constructing a sufficiently large and diverse dataset for estimating the novel views or building a powerful 3D foundation model for general objects remains insurmountable.
Unlike the scarcity of 3D models, images are much more readily available, and recent advancements in diffusion models have sparked a revolution in 2D image generation . Interestingly, we observed that well-trained image diffusion models can generate images under various views, which implies that they have already incorporated 3D knowledge. This has motivated us to explore the possibility of cultivating prior knowledge in a 2D diffusion model to reconstruct 3D objects. With diffusion prior, we propose Make-It-3D, a two-stage 3D content creation method that can generate a high-fidelity 3D object with superior quality from only one image.
In the first stage, we leverage diffusion prior to optimize a neural radiance field (NeRF) by applying score distillation sampling (SDS) , and constrain this optimization with reference-view supervision. Different from prior text-to-3D works , we focus on image-based 3D creation so that we need to prioritize the faithfulness to the reference image. However, we observed that while 3D models generated with SDS match text prompts well, they often fail to align faithfully with reference images since textual descriptions do not capture all object details. To address this issue, we go beyond SDS by simultaneously maximizing the image-level similarity between the reference and the novel view rendering denoised by a diffusion model. Also, as images inherently capture more geometry-related information than textual descriptions, we can thus incorporate the depth of the reference image as an extra geometry prior to alleviate the shape ambiguity of NeRF optimization.
While the first stage generates a coarse model with plausible geometry, its appearance often deviates from the quality of the reference, exhibiting over-smooth textures and saturated colors . This has limited its overall realism, and it is imperative to further bridge the gap between coarse model and reference image. As texture is more critical than geometry for human perception in the context of high-quality rendering, we choose to prioritize texture enhancement in the second stage, while inheriting the geometry from the first stage. We refine the model by leveraging the availability of ground-truth textures for regions that are observable in the reference image. To achieve this, we export the coarse NeRF model to textured point clouds and project reference textures onto their corresponding areas in the point clouds. We then utilize diffusion prior to enhance the texture of the remaining points by jointly optimizing the point feature and a point cloud renderer, resulting in a clearly improved texture of the generated 3D model.
With diffusion prior as multi-view supervision, our approach can be applied to general objects without being limited to specific categories. To evaluate the method, we create a benchmark consisting of 400 images including both real images and generated images from 2D diffusion. We evaluate the proposed method on public DTU dataset and our benchmark, and extensive experiments show a clear improvement over previous works. Furthermore, our method enables a range of applications beyond image-to-3D creation such as texture editing and high-quality text-to-3D creation. Our main contributions are summarized as:
We propose Make-It-3D, a framework to create a high-fidelity 3D object from a single image, using a 2D diffusion model as 3D-aware prior. It does not require multi-view images for training and can be applied to any input image, whether it is real or generated.
With a two-stage creation scheme, Make-It-3D represents the first work to achieve high-fidelity 3D creation for general objects. The resulting 3D models exhibit detailed geometry and realistic textures that accurately conform to the reference images.
Beyond image-to-3D creation, our method enables multiple applications such as high-quality text-to-3D creation and texture editing.
Related Work
Novel view synthesis from a few images. Early attempts usually require dense observations of a scene from uniformly sampled poses. Recent emergence of implicit representations significantly advances the synthesis quality of novel views, whereas they tend to find a degenerate solution when given only very few input views. To enable novel view from sparse input views, a growing body of works hence turn to extra prior knowledge as additional regularizations. PixelNeRF predicts a continuous neural representation conditioned on the input images rather than only leveraging input views for supervision. DietNeRF penalizes a semantic consistency loss by minimizing distance between CLIP features of different views. With recent rapid progress of diffusion models, 3DiM introduces a pose-conditional diffusion model that generates a novel view conditioned on a source view and a target pose. RenderDiffusion presents a diffusion model for 3D generation that incorporates a triplane rendering mode into the denoiser.
Single-image 3D photography. Synthesizing novel views from a single image is quite challenging as it is a highly ill-posed problem, requiring precise geometry estimation and disocclusion of both geometry and texture. Increasing effort has been dedicated to this problem , many of which can only handle specific types. Among them, a number of methods rely on layered representations such as layered depth images and multi-plane images (MPIs) . For example, predicts MPIs for view synthesis from single image without requiring ground truth 3D. generates a 3D photo from a given RGB-D input through layered depth image with inpainted color and depth. Yet such a solution is limited by the number of planes and sensitive to discontinuities. Subsequent efforts generalize MPIs to continuous 3D representations such as NeRF and latent 3D point cloud .
Lift 2D pretrained model to 3D. With the emergence of recent advances in modeling natural image manifold, how to exploit such powerful 2D pretrained model to recover 3D object structure has received considerable research interest. attempts to reconstruct the 3D shape using pretrained 2D GANs. Subsequently, some works explore zero-shot text-guided 3D content creation utilizing the guidance from CLIP . Recent efforts such as DreamFusion , Magic3D and Score Jacobian Chaining explore text-to-3D generation by exploiting a score distillation sampling (SDS) loss derived from a 2D text-to-image diffusion model instead, showing impressive results. LatentNeRF proposes to use a shape prior to guide and assist the 3D generation directly in the latent space of the diffusion model. Prior works NeuralLift-360 and NeRDi also leverage the generative prior for 3D reconstruction from a single view. Yet the reconstructed 3D model has limited quality and is poorly aligned with the input image. In contrast, we propose a two-stage 3D synthesis framework with a relaxed SDS loss, yielding high-quality 3D representation faithful to the given input image.
Method
Generating novel views for general scenes or objects from only a single image is inherently challenging due to the difficulty of inferring both geometry and missing texture. We therefore tackle this challenge by cultivating the dark knowledge of pretrained 2D diffusion models. Specifically, given an input image , we first hallucinate its underlying 3D representation, neural radiance field (NeRF), whose rendering appears as a plausible sample to a pretrained denoising diffusion model, and we constrain this optimization process with the texture and depth supervision at the reference view. To further improve the rendering realism, we keep the learned geometry and enhance the textures with the reference image. As such, in the second stage, we lift the input image to textured point clouds and focus on refining the color of the points occluded in the reference view. We leverage prior knowledge of the text-to-image generative model and the text-image contrastive model for both stages. In this way, we achieve a faithful 3D representation of the input image with restored high-fidelity texture and geometry. The proposed two-stage 3D learning framework is illustrated in Figure 2. We will subsequently brief the preliminaries and then detail our method.
Recent findings show that pretrained 2D generative models offer rich 3D geometry knowledge for their 2D generation samples. Notably, DreamFusion uses a text-to-image diffusion model to guide the optimization of the 3D representation. Let be the rendered image at the given viewpoint , where is the differentiable rendering function for the 3D representation parameterized by and is amenable to choice. DreamFusion optimizes the neural radiance field such that its multi-view renderings look like high-quality samples from a frozen diffusion model.
Specifically, a diffusion model introduces a random amount of noise to the rendered image at different timestep , i.e., , where ; and define a noise schedule whose log signal-to-noise ratio linearly decreases with the timestep . A pretrained text-conditioned diffusion model is trained to reverse this noising process given the text embedding . To optimize the 3D representation parameters to render images as close as good generation samples, a score distillation sampling (SDS) loss is introduced to push rendered images toward higher density region conditioned on the text embedding. Specifically, computes the difference of predicted noise and the added noise as per-pixel gradient which is used to update the scene parameters, i.e.,
where is a weight function of different noise levels. It can be proved that this loss essentially measures the similarity between the image and the text prompt. The diffusion model acts as a critic and the gradient of will not be back-propagated through the diffusion network, resulting in efficient computation. As training proceeds, the NeRF parameters are updated during which the 3D object gradually reveals its texture and geometry. In practice, it is found that using a diffusion model with a strong classifier-free guidance strength leads to higher-quality 3D samples.
While DreamFusion uses the Imagen to reverse the noising process at the pixel level, we use the publicly available Stable Diffusion that models the latent space of the VQ-VAE with an encoder and a decoder . Hence, the used diffusion model digests the latent and the reconstructed latent can be mapped to image space through .
2 Coarse Stage: Single-view 3D Reconstruction
As the first stage, we reconstruct a coarse NeRF from the single reference image with the diffusion prior constraining the novel views. Our optimization is expected to meet the following requirements simultaneously: 1) the optimized 3D representation should closely resemble the rendering appearance of the input observation at the reference view; 2) the novel view renderings should demonstrate consistent semantics with the input and appear as plausible as possible; 3) the generated 3D model should exhibit compelling geometry. In view of these, we randomly sample the camera poses around the reference view and enforce constraints upon the rendered images for both the reference view and unseen views.
Reference view per-pixel loss. To encourage consistent appearance with the input image, we penalize the pixel-wise difference between the rendering and the input image at the reference view :
Here we apply the foreground matting mask to segment out the foreground as we empirically find that this eases the geometry reconstruction, which conforms to .
Diffusion prior. Optimizing with the aforementioned losses can be unstable and may lead to implausible results, due to the ill-posed nature of the problem. In order to encourage semantically plausible results, additional constraints are needed on the novel view rendering. To tackle this challenge, we resort to the diffusion prior. Prior works on text-to-3D applied to leverage text-conditioned diffusion models as 3D-aware prior. To utilize in our case, we use an image captioning model , to generate a detailed text description for the reference image. With the text prompt , we can perform the SDS on the latent space of Stable Diffusion,
where the noisy latent is obtained form a novel view rendering by Stable Diffusion encoder.
However, as discussed before, essentially measures the similarity between the image and the given text prompt. While can generate 3D models that are faithful to the text prompt, they do not align perfectly with the reference image (see baseline in Figure 3), since text prompts cannot capture all object details. We go beyond this by a diffusion CLIP loss, denoted as , that additionally enforces the generated model to match the reference image:
where is a CLIP image encoder . Rather than directly measuring CLIP loss on the rendered images , we encode the novel view rendering to noisy latent and then denoise it to a clean image with 2D diffusion. By imposing the similarity loss on denoised images sampled from diffusion models, we encourage the rendering to align with the reference image, while resembling high-quality samples from a frozen diffusion.
In detail, we do not optimize and at the same time. We use at small timesteps and switch to at large timesteps. More details and analysis are in the Supplement. Combining and , our diffusion prior ensures that the resulting 3D model appears visually appealing and plausible while also conforming to the given image (see Figure 3).
Depth prior. Nonetheless, even if the rendered image appears meaningful to the diffusion model, there still exists shape ambiguity that brings about issues such as sunken faces, over-flat geometry or depth ambiguity (see Figure 3). We mitigate these by leveraging depth prior learned from abundant external images and directly enforcing the supervision in 3D. To be specific, we utilize an off-the-shelf single-view depth estimator to estimate the depth for the input image. While the estimated depth may not accurately characterize the geometric detail, it suffices to ensure plausible geometry and resolve most of the ambiguity. To account for the inaccuracy and the scale mismatch in , akin to , we regularize the negative Pearson correlation between the estimated depth and the depth modeled by NeRF at the reference viewpoint, i.e.,
where denotes the covariance, computes the standard deviation. With this regularization, the NeRF depth estimation is encouraged to be linearly correlated with the depth prior.
Overall training. The overall loss can be formulated as a combination of , , and .To stabilize the optimization process, we adopt a progressive training strategy, where we start with a narrow range of views near the reference view and gradually expand the range during training. With progressive training, we can achieve a reconstruction of an object, as shown in Figure 4.
3 Refine Stage: Neural Texture Enhancement
After the coarse stage, we obtained a 3D model with plausible geometry, but it often displays coarse textures that can bottleneck the overall quality in Figure 6. Further refinement is thus desired for high-fidelity 3D models. Given that humans are more discerning when it comes to texture quality than geometry, we prioritize texture enhancement while preserving the geometry of the coarse model.
Our key insight for texture enhancement is that for a novel view, certain pixels can be observable in both the novel and reference views. Consequently, we can exploit this overlap to project the high-quality texture of the reference image onto the corresponding areas of the 3D representation. We then focus on enhancing the textures of regions that are occluded in the reference view.
While NeRF is a suitable representation in the coarse stage as it can handle topological changes continuously, projecting the reference image onto it is challenging. We thus opt to export the neural radiance field to an explicit representation, specifically point clouds. Compared to the noisy mesh exported by marching cube, point clouds offer a cleaner and more straightforward projection.
Textured point cloud building. A naive attempt to build point clouds is to render multi-view RGBD images from NeRF and lift them to textured points in 3D space. However, we found this simple method leads to noisy point clouds due to the conflict among different views: a 3D point may possess different RGB colors in NeRF rendering under different views . We thus propose an iterative strategy to build clean point clouds from multi-view observations.
As in Figure 5, we first build point clouds from the reference view according to the rendered depth and alpha mask of NeRF,
where and are the extrinsic and intrinsic matrices of the camera, and denotes depth-to-point projection. These points are visible under the reference view and thus colorized with ground-truth textures. For the projection of the remaining views , it is important to avoid introducing points that overlap with existing points but have conflicting colors. To this end, we project the existing points to the novel view to yield a mask indicating the presence of existing points. With this mask as guidance, we only lift those points that have not been observed yet, as shown in Figure 5. These invisible points are then initialized with coarse textures from NeRF rendering and integrated into the dense point clouds.
Deferred point cloud rendering. So far, we have built a set of textured point clouds . Though already have high-fidelity textures projected from the reference image, the other points that are occluded in the reference view still suffer smooth textures from the coarse NeRF, as shown in Figure 6. To enhance the texture, we optimize the texture of the other points and constrain novel-view rendering with diffusion prior. Specifically, we optimize a 19-dimensional descriptor for each point, whose first three dimensions are initialized with the initial RGB colors. To avoid noisy colors and bleeding artifacts , we adopt a multi-scale deferred rendering scheme. In particular, given a novel view , we rasterize the point cloud V for times to obtain feature maps with varying sizes of , where . These feature maps are then concatenated and rendered into an image I using a U-Net renderer that is jointly optimized:
where is a differentiable point rasterizer. The objective of the texture enhancement process is similar to that of the geometry creation discussed in Sec. 3.2, but we additionally include a regularization term that penalizes large differences between the optimized texture and the initial texture.
Experiments
NeRF rendering. We use the multi-scale hash encoding from Instant-NGP to implement the NeRF representation in the coarse optimization stage, which enables neural rendering at a computational cost. Similar to Instant-NGP, we maintain an occupancy grid to enable efficient ray sampling by skipping empty space. Additionally, we adopt several shading augmentations on the rendered images, such as Lambertian and normal shading, akin to .
Point cloud rendering. For deferred rendering, we use a 2D U-Net architecture with gated convolutions . The dimension of the point descriptor is 19, where the first 3 dimensions are initialized RGB colors and the remaining dimensions are randomly initialized. We also set a learnable descriptor for the background.
Camera setting. Following the camera sampling method used in , we randomly sample novel views with a 75 probability and sample the pre-defined reference view with a 25 probability. We also randomly enlarge the FOV when rendering with NeRF, following .
Score distillation sampling. We randomly sample from 200 to 600, and set as a uniform weighting depending on the timestep. We also use classifier-free guidance with a guidance weight . Our method aims to align the created 3D model with the input image, and we use a guidance weight .
Training speed. We use Adam with a learning rate of 0.001 for both stages. The coarse stage is trained for 5,000 iterations at a rendering resolution of 100100. The refine stage then takes another 5,000 iterations at a rendering resolution of 800800. The entire training process takes approximately 2 hours on a single Tesla 32GB V100 GPU.
Test Benchmark. To the best of our knowledge, we are the first method focusing on high-fidelity 3D creation from an arbitrary image. So we build a test benchmark consisting of 400 images, comprising both real images and images generated by Stable Diffusion . Each image in the benchmark is accompanied by a foreground alpha mask, an estimated depth map, and a text prompt. The text prompts for real images are obtained from an image caption model . We will make this test benchmark publicly available.
2 Comparisons with the State of the Arts
Baselines. We compare our method with five representative baselines. 1) DietNeRF , a few-shot NeRF. We train it with three input views. 2) SinNeRF , a single-view NeRF method. 3) DreamFusion . As it is originally conditioned on text prompts, we also modify it with image reconstruction loss at the reference view, referred as DreamFusion+ for fair comparison. 4) Point-E , point cloud generation conditioned on image. 5) 3D-Photo , depth-based image warping and inpainting method.
Qualitative comparison. We first compare our method with 3D generation baselines, where DreamFusion and DreamFusion+ leverage 2D diffusion as 3D prior and PointE is a 3D diffusion model. As shown in Figure 7, their generated models fail to align faithfully with the reference image and suffer smooth textures. In contrast, our method produces high-fidelity 3D models with fine geometry and realistic textures. Figure 8 shows additional comparison on novel view synthesis. SinNeRF and DietNeRF encounter difficulties in reconstructing complex objects due to the lack of multi-view supervision. 3D-Photo fails to reconstruct underlying geometry and produces visible artifacts in large views. In comparison, our method achieves remarkably faithful geometry and visually pleasing textures under novel views.
Quantitative comparison. A compelling generated 3D model should closely resemble the input image at the reference view, and demonstrate consistent semantics with the reference under novel views. We evaluate these two aspects using the following metrics: 1) LPIPS , which assesses the reconstruction quality at the reference view, 2) contextual distance , which measures pixel-level similarity between novel-view rendering and the reference, and 3) CLIP score , which evaluates the semantic similarity between the novel view and the reference. As shown in Table 1 and Table 2, our approach substantially outperforms baselines in terms of both reference-view and novel-view quality.
Applications
Real scene modeling. As shown in Figure 9, Make-It-3D can successfully convert a single photo of a complex scene to a 3D model, such as buildings and landscapes. This empowers users to model a scene with ease, which could be difficult for some traditional 3D modeling techniques.
High-quality text-to-3D generation with diversity. Prior arts often produce models with limited diversity and excessively smooth textures. To perform high-quality text-to-3D creation, we first convert the text prompt to a reference image using 2D diffusion, and proceed with our image-based 3D creation method. As shown in Figure 10, Make-It-3D is capable of producing diverse examples from a text prompt that exhibit stunning quality.
3D-aware texture modification. Make-It-3D enables view-consistent texture editing by manipulating the reference image in the refine stage while freezing the geometry. Figure 11 shows that we can add a tattoo and apply stylization to the generated 3D model.
Conclusions
We introduce Make-It-3D, a novel two-stage method for creating high-fidelity 3D content from one single image. Leveraging diffusion prior as 3D-aware supervision, the generated 3D models exhibit faithful geometry and realistic textures with the diffusion CLIP loss and textured point cloud enhancement. Make-It-3D is applicable to general objects, empowering versatile fascinating applications. We believe our method takes a big step in extending the success of 2D content creation to 3D, providing users with a fresh 3D creation experience.
References
Appendix
Appendix A Broad Impact
We have presented Make-It-3D, a novel approach to create novel views from a single image of general genre. Make-It-3D first hallucinates the 3D geometry by the usage of depth prior at the frontal view and the geometry prior of a pretrained diffusion model to ensure plausibility at novel views. Motivated by the fact that human eyes are more sensitive to texture over geometry, we thus reuse the coarse 3D geometry estimated from the implicit representation as well as the texture from the reference image, and specifically “inpaints” the texture of explicit 3D representation at occluded regions, ultimately producing compelling novel view renderings with highly-detailed texture.
Our primary aim is to advance the research of generative modeling from 2D to 3D. Without relying on 3D training data that is hardly accessible in scale, this work tackles the 3D synthesis problem by lifting 2D generated images to 3D. This way essentially builds on the assumption that a diffusion model not only generates 2D observations but also implicitly contains rich 3D understanding of the scene. Thus, using our technique, one can generate a 3D scene that can be immersively viewed by merely using a 2D diffusion model. Compared to DreamFusion and Magic3D, our work produces more diverse 3D synthesis results with significantly improved realism. On top of creatively generated images, this work also performs well on real images with complicated structures.
We hope this work opens the door towards high-quality 3D synthesis and inspires more following works along this way. While we have demonstrated the ability to synthesize novel views in 360 degree, it is still non-trivial to produce holistically plausible 3D objects when viewed from large viewpoints. Moreover, while this work aims for 3D synthesis from a single image, the same pipeline is applicable to the few-shot scenario where a few multi-view images can be obtained. In addition, it would be fruitful to generalize the proposed technique to augment the quality of 4D synthesis. We will release the code to facilitate the research in this emerging area.
Appendix B Additional Implementation Details
Scene representation and rendering. We use the explicit-implicit representation from Instant-NGP to implement the NeRF representation in the coarse optimization stage, where we choose 16-level hash encoding of size and dimension 32, with a 3-layer MLP with 64 hidden units to decode the density and color for each spatial location. During volumetric rendering, we sample 96 points for each ray, including 64 points for uniform sampling and 32 for importance sampling. We initialize the density field as a Gaussian sphere, which leads to faster convergence and more stable training. Specifically, we initialize the density as , where we set density bias and ; denotes the distance between the ray point and the scene center.
Camera setting. Following the camera sampling method used in , we randomly sample camera distance from 0.8 to 1.2, and the field-of-view (FOV) from 40 to 80 degrees. We find that randomly sampling FOV is instrumental to mitigate the artifacts that arise in large rendering view angles.
Augmentation and Regularization. To encourage the network to focus more on the foreground and avoid adversarial samples that hack the pretrained diffusion model, we train NeRF with a random background augmentation. Specifically, during training, we randomly jitters the background color of both the reference alpha image and NeRF rendering. During inference, we render the scene with a white background. Furthermore, following , we use three types of geometric regularization including sparsity, opacity and smoothness.
B.2 Refine stage
Background regularization. To handle pixels without corresponding point cloud projection, we assign a learnable descriptor as the background. During texture enhancement optimization, we additionally add a regularization to encourage the scene to be rendered with a white background according to the binary occupancy mask mentioned above.
Deferred neural rendering. For deferred rendering of the point clouds, we use a 2D U-Net architecture with gated convolutions . It contains 3 down- and up-sampling layers to integrate multi-scale feature maps and output the final RGB image.
Appendix C Additional Ablation Study and Analysis
As mentioned in Sec 3.1, in the coarse stage, we use the diffusion prior by applying score distillation sampling (SDS) scheme on novel view renderings. It can successfully encourage the generated scene to match the conditioned text prompt. However, as an image-based 3D content creation model, we need to prioritize the faithfulness between created 3D and the reference image. Although we add pixel-wise constrain under the reference view for optimization, SDS provides a strong geometric prior and enforces the optimized scene to be a plausible result according to the text condition. Constraints under a single view can be limited. Thus the created results may not be rigorously aligned with the reference image (See Figure 12).
Therefore, we need to relax the strong geometric guidance provided by SDS and add more image-level constraints under multi-views. We achieve this goal by simultaneously maximizing the image-level similarity between the reference image and the novel view renderings denoised by the diffusion model, named as a diffusion CLIP loss . Compared with introducing this constraint directly on novel view renderings, the CLIP-D encourages the pretrained diffusion model to provide better guidance to generate more faithful 3D content with the reference image.
In view of this, we conduct several experiments to study the effect of SDS and CLIP-D loss during optimization, which is shown in Figure 12. Results show that using only SDS generates high-quality and plausible geometry, but the optimized 3D does not align with the image. On the contrary, using only CLIP-D preserves the appearance of the reference image, but fails to generate good geometry. A simple solution is to combine the two losses, but this does not fully address the non-alignment issue. To achieve a balance between geometric quality and appearance alignment, we introduce an optimization strategy by setting a threshold of sampling steps. Specifically, we optimize CLIP-D loss at small timesteps and optimize SDS at large steps. We conduct several qualitative and qualitative studies on different threshold settings, which are shown in Figure 12 and Table 3. During training, we randomly sample noise step from 200 to 600, and we find that could balance the geometric quality and the appearance alignment.
C.2 Analysis of various sampling time step ranges
We also investigate the effect of various sampling time steps in SDS process. The experimental results are shown in Figure 13. We conduct several experiments using different sampling ranges. We observe that adding noise at large time steps can improve the geometry quality but reduce the alignment and potentially saturate textures. And the diffusion prior does not provide adequate supervision at small time steps. In our method, we exclude small and large time steps and instead randomly sample time step from 200 to 600.
C.3 Analysis of texture initialization and point descriptors
We conduct ablation studies on texture enhancement process. We explore the importance of the initialized unseen texture from NeRF and point descriptor. The qualitative results are shown in Figure 14. We can see that texture initialization is crucial for global texture enhancement. And only optimizing point color without descriptor outputs artifacts and cannot produce reasonable results.
Appendix D Additional Results
In this section, we provide additional results of creating 3D models from different reference images using our method. The results are shown on Figure 16, Figure 17, and Figure 18. Results show that our method has a strong ability on creating high-fidelity 3D content including high-quality geometries and textures using a single reference image.
Appendix E Limitations
Our method suffers from some geometry ambiguity, such as Janus problem or over-flat geometry . A depth prior can reduce this issue. However, since we only add depth constrain at a single view, the geometry ambiguity may still exist under other views. We show some failure cases in Figure 15.