Wonder3D: Single Image to 3D using Cross-Domain Diffusion
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, Wenping Wang
Introduction
Reconstructing 3D geometry from a single image stands as a fundamental task in computer graphics and 3D computer vision , offering a wide range of versatile applications such as virtual reality, video games, 3D content creation, and the precision of robotics grasping. However, this task is notably challenging since it is ill-posed and demands the ability to discern the 3D geometry of both visible and invisible parts. This ability requires extensive knowledge of the 3D world.
Recently, the field of 3D generation has experienced rapid and flourishing development with the introduction of diffusion models. A growing body of research , such as DreamField , DreamFusion , and Magic3D , resort to distilling prior knowledge of 2D image diffusion models or vision language models to create 3D models from text or images via Score Distillation Sampling (SDS) . Despite their compelling results, these methods suffer from two main limitations: efficiency and consistency. The per-shape optimization process typically entails tens of thousands of iterations, involving full-image volume rendering and inferences of the diffusion models. Consequently, it often consumes tens of minutes or even hours on per-shape optimization. Moreover, the 2D prior model operates by considering only a single view at each iteration and strives to make every view resemble the input image. This often results in the generation of 3D shapes exhibiting inconsistencies, thus, often leading to the generation of 3D shapes with inconsistencies such as multiple faces (i.e., the Janus problem ).
There exists another group of works that endeavor to directly produce 3D geometries like point clouds , meshes , neural fields via network inference to avoid time-consuming per-shape optimization. Most of them attempt to train 3D generative diffusion models from scratch on 3D assets. However, due to the limited size of publicly available 3D datasets, these methods demonstrate poor generalizability, most of which can only generate shapes on specific categories.
More recently, several methods have emerged that directly generate multi-view 2D images, with representative works including SyncDreamer and MVDream . By enhancing the multi-view consistency of image generation, these methods can recover 3D shapes from the generated multi-view images. Following these works, our method also adopts a multi-view generation scheme to favor the flexibility and efficiency of 2D representations. However, due to only relying on color images, the fidelity of the generated shapes is not well-maintained, and they struggle to recover geometric details or come with enormous computational costs.
To better address the issues of fidelity, consistency, generalizability and efficiency in the aforementioned works, in this paper, we introduce a new approach to the task of single-view 3D reconstruction by generating multi-view consistent normal maps and their corresponding color images with a cross-domain diffusion model. The key idea is to extend the stable diffusion framework to model the joint distribution of two different domains, i.e., normals and colors. We demonstrate that this can be achieved by introducing a domain switcher and a cross-domain attention scheme. In particular, the domain switcher allows the diffusion model to generate either normal maps or color images, while the cross-domain attention mechanisms assist in the information exchange between the two domains, ultimately improving consistency and quality. Finally, in order to stably extract surfaces from the generated views, we propose a geometry-aware normal fusion algorithm that is robust to the inaccuracies and capable of reconstructing clean and high-quality geometries (see Figure 1).
We conduct extensive experiments on the Google Scanned Object dataset and various 2D images with different styles. The experiments validate that Wonder3D is capable of producing high-quality geometry with high efficiency in comparison with baseline methods. Wonder3D possesses several distinctive properties and accordingly has the following contributions:
Wonder3D holistically considers the issues of generation quality, efficiency, generalizability, and consistency for single-view 3D reconstruction. It has achieved a leading level of geometric details with reasonably good efficiency among current zero-shot single-view reconstruction methods.
We propose a new multi-view cross-domain 2D diffusion model to predict normal maps and color images. This representation not only adapts to the original data distribution of Stable Diffusion model but also effectively captures the rich surface details of the target shape.
We propose a cross-domain attention mechanism to produce multi-view normal maps and color images that are consistently aligned. This mechanism facilitates information perception across different domains, enabling our method to recover high-fidelity geometry.
We introduce a novel geometry-aware normal fusion algorithm that can robustly extract surfaces from the generated normal maps and color images.
Related Works
Recent compelling successes in 2D diffusion models and large vision language models (e.g., CLIP model ) provide new possibilities for generating 3D assets using the strong priors of 2D diffusion models. Pioneering works DreamFusion and SJC propose to distill a 2D text-to-image generation model to generate 3D shapes from texts, and many follow-up works follow such per-shape optimization scheme. For the task of text-to-3D or image-to-3D synthesis , these methods typically optimize a 3D representation (i.e., NeRF, mesh, or SDF), and then leverage neural rendering to generate 2D images from various viewpoints. The images are then fed into the 2D diffusion models or CLIP model for calculating SDS losses, which can guide the 3D shape optimization.
However, most of these methods always suffer from low efficiency and multi-face problem, where a per-shape optimization consumes tens of minutes and the optimized geometry tends to produce multiple faces due to the lack of explicit 3D supervision. A recent work one-2-3-45 proposes to leverage a generalizable neural reconstruction method SparseNeuS to directly produce 3D geometry from the generated images from zero123 . Although the method achieves high efficiency, its results are of low-quality and lack geometric details.
2 3D Generative Models
Instead of performing a time-consuming per-shape optimization guided by 2D diffusion models, some works attempt to directly train 3D diffusion models based on various 3D representations, like point clouds , meshes , neural fields However, due to the limited size of public available 3D assets dataset, most of the works have only been validated on limited categories of shapes, and how to scale up on large datasets is still an open problem. On the contrary, our method adopts 2D representations and, thus, can be built upon the 2D diffusion models whose pre-trained priors significantly facilitate zero-shot generalization ability.
3 Multi-view Diffusion Models
To generate consistent multi-view images, some efforts are made to extend 2D diffusion models from single-view images to multi-view images. However, most of these methods focus on image generation and are not designed for 3D reconstruction. The works first warp estimated depth maps to produce incomplete novel view images to then perform inpainting on them, but their result quality significantly degrades when the depth maps estimated by external depth estimation models are inaccurate. The recent works Viewset Diffusion , SyncDreamer , and MVDream share a similar idea to produce consistent multi-view color images via attention layers. However, unlike that normal maps explicitly encode geometric information, reconstruction from color images always suffers from texture ambiguity, and, thus, they either struggle to recover geometric details or require huge computational costs. SyncDreamer requires dense views for 3D reconstruction, but still suffers from low-quality geometry and blurring textures. MVDream still resorts to a time-consuming optimization using SDS loss for 3D reconstruction, and its multi-view distillation scheme requires 1.5 hours. In contrast, our method can reconstruct high-quality textured meshes in just 2 minutes.
Problem Formulation
Diffusion models are first proposed to gradually recover images from a specifically designed degradation process, where a forward Markov chain and a Reverse Markov chain are adopted. Given a sample drawn from the data distribution , the forward process of denoising diffusion models yields a sequence of noised data with , where is random noise drawn from distribution , and are fixed sequence of the noise schedule. The forward process will be iteratively applied to the target image until the image becomes complete Gaussian noise at the end. On the contrary, the reverse chain then is employed to iteratively denoise the corrupted image, i.e., recovering from by predicting the added random noise . The readers can refer to for more details about image diffusion models.
2 The Distribution of 3D Assets
Unlike that prior works adopt 3D representations like point clouds, tri-planes, or neural radiance fields, we propose that the distribution of 3D assets, denoted as , can be modeled as a joint distribution of its corresponding 2D multi-view normal maps and corresponding color images. Specifically, given a set of cameras and a conditional input image , we have
where is the distribution of the normal maps and color images observed from 3D assets conditioned on an image . For simplicity, we omit the symbol for this equation in the following discussions. Therefore, our goal is to learn a model that synthesizes multiple normal maps and color images of a set of camera poses denoted as
Adopting the 2D representation enables our method to be built upon the 2D diffusion models trained on billions of images like the Stable Diffusion model , where strong priors facilitate zero-shot generalization ability. On the other hand, the normal map characterizes the undulations and variations present on the surface of the shape, thus encoding rich detailed geometric information. This allows for the high-fidelity extraction of 3D geometry from 2D normal maps.
Finally, we can formulate this cross-domain joint distribution as a Markov chain within the diffusion scheme:
where are Gaussian noises. Our key problem is to characterize the distribution , so that we can sample from this Markov chain to generate normal maps and images.
Method
As per our problem formulation in Section 3.2, we propose a multi-view cross-domain diffusion scheme, which operates on two distinct domains to generate multi-view consistent normal maps and color images. The overview of our method is presented in Figure 2. First, our method adopts a multi-view diffusion scheme to generate multi-view normal maps and color images, and enforces the consistency across different views using multi-view attentions (see Section 4.1). Second, our proposed domain switcher allows the diffusion model to operate on more than one domain while its formulation does not require a re-training of an existing (potentially single domain) diffusion model such as Stable Diffusion . Thus, we can leverage the generalizability of large foundational models, which are trained on a large corpus of data. A cross-domain attention is proposed to propagate information between the normal domain and color image domain ensuring geometric and visual coherence between the two domains (see Section 4.2). Finally, our novel geometry-aware normal fusion reconstructs the high-quality geometry and appearance from the multi-view 2D normal and color images (see Section 4.3).
The prior 2D diffusion models generate each image separately, so that the resulting images are not geometrically and visually consistent across different views. To enhance consistency among different views, similar to prior works such as SyncDreamer and MVDream , we utilize attention mechanism to facilitate information propagation across different views, implicitly encoding multi-view dependencies (as illustrated in Figure 4).
This is achieved by extending the original self-attention layers to be global-aware, allowing connections to other views within the attention layers. Keys and values from different views are connected to each other to facilitate the exchange of information. By sharing information across different views within the attention layers, the diffusion model perceives multi-view correlation and becomes capable of generating consistent multi-view color images and normal maps.
2 Cross-Domain Diffusion
Our model is built upon pre-trained 2D stable diffusion models to leverage its strong generalization. However, current 2D diffusion models are designed for a single domain, so the main challenge is how to effectively extend stable diffusion models that are capable of operating on more than one domain.
Naive Solutions. To achieve this goal, we explore several possible designs. A straightforward solution is to add four more channels to the output of the UNet module representing the extra domain. Therefore, the diffusion model can simultaneously output normals and color image domains. However, we notice that such a design suffers from low convergence speed and poor generalization. This is because the channel expansion may perturb the pre-trained weights of stable diffusion models and therefore cause catastrophic model forgetting.
Revisiting Eq. 1, it is possible to factor the joint distribution into two conditional distributions:
This equation suggests an alternative solution where we could initially train a diffusion model to generate normal maps and then train another diffusion model to produce color images, conditioning on the generated normal maps (or vice versa). Nonetheless, the implementation of this two-stage framework introduces certain complications. It not only substantially increases the computational cost but also results in performance degradation. Please refer to Section 5.6 for an in-depth discussion.
Domain Switcher. To overcome these difficulties mentioned above, we design a cross-domain diffusion scheme via a domain switcher, denoted as . The switcher is a one-dimensional vector that labels different domains, and we further feed the switcher into the diffusion model as an extra input. Therefore, the formulation of Eq. 2 can be extended as:
The domain switcher is first encoded via positional encoding and subsequently concatenated with the time embedding. This combined representation is then injected into the UNet of the stable diffusion models. Interestingly, experiments show that this subtle modification does not significantly alter the pre-trained priors. As a result, it allows for fast convergence and robust generalization, without requiring substantial changes to the stable diffusion models.
Cross-domain Attention. Using the proposed domain switcher, the diffusion model can generate two different domains. However, it is important to note that for a single view, there is no guarantee that the generated color image and the normal map will be geometrically consistent. To address this issue and ensure the consistency between the generated normal maps and color images, we introduce a cross-domain attention mechanism to facilitate the exchange of information between the two domains. This mechanism aims to ensure that the generated outputs align well in terms of geometry and appearance.
The cross-domain attention layer maintains the same structure as the original self-attention layer and is integrated before the cross-attention layer in each transformer block of the UNet, as depicted in Figure 4. In the cross-domain attention layer, the keys and values from the normal and color image domains are combined and processed through attention operations. This design ensures that the generations of color images and normal maps are closely correlated, thus promoting geometric consistency between the two domains.
3 Textured Mesh Extraction
To extract explicit 3D geometry from 2D normal maps and color images, we optimize a neural implicit signed distance field (SDF) to amalgamate all 2D generated data. Unlike alternative representations like meshes, SDF offers compactness and differentiability, making them ideal for stable optimization.
Nonetheless, adopting existing SDF-based reconstruction methods, such as NeuS , proves unviable. These methods were tailored for real-captured images and necessitate dense input views. In contrast, our generated views are relatively sparse, and the generated normal maps and color images may exhibit subtle inaccurate predictions of some pixels. Regrettably, these errors accumulate during the geometry optimization, leading to distorted geometries, outliers, and incompleteness. To overcome the challenges above, we propose a novel geometric-aware optimization scheme.
Optimization Objectives. With the obtained normal maps and color images , we first leverage segmentation models to segment the object masks from the normal maps or color images. Specifically, we perform the optimization by randomly sampling a batch of pixels and their corresponding rays in world space , where is normal value of the sampled pixel, is color value of the pixel, is mask value of the pixel, and is the direction of the corresponding sampled ray, from all views at each iteration.
The overall objective function is defined as
where denotes the normal loss term that will be discussed later, denotes a MSE loss term that calculates the errors between rendered colors and generated colors , denotes a binary cross-entropy loss term that calculating errors between the rendered mask and the generated mask , denotes eikonal regularization term that encourages the magnitude of the SDF gradients to be unit length, denotes a sparsity regularization term that avoid floaters of SDF, and denotes a 3D smoothness regularization term that enforces the SDF gradients to be smooth in 3D space.
Geometry-aware Normal Loss. Thanks to the differentiable nature of SDF representation, we can easily extract normal values of the optimized SDF via calculating the second-order gradients of SDF. We maximize the similarity of the normal of SDF and our generated normal to provide 3D geometric supervision. To tolerate trivial inaccuracies of the generated normals from different views, we introduce a geometry-aware normal loss:
where is error between the normal of SDF and the generated normal for the sampled ray, denotes cosine function, and is a geometric-aware weight defined as
Here denotes exponential function, denotes absolute function, is a negative threshold closing to zero, and we measure the cosine value of the angle between the generated normal and the ray’s viewing direction .
The design rationale behind this approach lies in the orientation of normals, which are deliberately set to face outward, while the viewing direction is inward-facing. This configuration ensures that the angle between the normal vector and the viewing ray remains not less than . A deviation from this criterion would imply inaccuracies in the generated normals.
Furthermore, it’s worth noting that a 3D point on the optimized shape can be visible from multiple distinct viewpoints, thereby being influenced by multiple normals corresponding to these views. However, if these multiple normals do not exhibit perfect consistency, the geometric supervision may become somewhat ambiguous, leading to imprecise geometry. To address this issue, rather than treating normals from different views equally, we introduce a weighting mechanism. We assign higher weights to normals that form larger angles with the viewing rays. This prioritization enhances the accuracy of our geometric supervision process.
Outlier-dropping Losses. Besides the normal loss, mask loss and color loss are also adopted for optimizing geometry and appearance. However, it is inevitable that there exist some inaccuracies in the masks and color images, which will accumulate in the optimization and thus cause noisy surfaces and holes.
To mitigate the issues, we employ a simple yet effective strategy named outlier-dropping loss. Taking the color loss calculation as an example, instead of simply summing up the color errors of all sampled rays at each iteration, we first sort these errors in a descending order and then discard the top largest errors according to a predefined percentage. This approach is motivated by the fact that erroneous predictions lack sufficient consistency with other views, making them less amenable to effective minimization during optimization, and they often result in large errors. By implementing this strategy, the optimized geometry can eliminate incorrect isolated geometries and distorted textures.
Experiments
We train our model on the LVIS subset of the Objaverse dataset , which comprises approximately 30,000+ objects following a cleanup process. Surprisingly, even with fine-tuning on this relatively small-scale dataset, our method demonstrates robust generalization capabilities. To create the rendered multi-view dataset, we first normalized each object to be centered and of unit scale. Then we render normal maps and color images from six views, including the front, back, left, right, front-right, and front-left views, using Blenderproc . Additionally, to enhance dataset diversity, we applied random rotations to the 3D assets during the rendering process.
We fine-tune our model starting from the Stable Diffusion Image Variations Model, which has previously been fine-tuned with image conditions. We retain the optimizer settings and -prediction strategy from the previous fine-tuning. During fine-tuning, we use a reduced image size of 256 256 and a total batch size of 512 for training. The fine-tuning process involves training the model for 30,000 steps. This entire training procedure typically requires approximately 3 days on a cluster of 8 Nvidia Tesla A800 GPUs. To reconstruct 3D geometry from the 2D representations, our method is built on the instant-NGP based SDF reconstruction method .
2 Baselines
We adopt Zero123 , RealFusion , Magic123 , One-2-3-45 , Point-E , Shap-E and a recent work SyncDreamer as baseline methods. Given an input image, zero123 is capable of generating novel views of arbitrary viewpoints, and it can be incorporated with SDS loss for 3D reconstruction (we adopt the implementation of ThreeStudio ). RealFusion and Magic123 leverage Stable Diffusion and SDS loss for single-view reconstruction. One-2-3-45 directly predict SDFs via SparseNeuS by taking the generated multiple images of Zero123 . Point-E and Shap-E are 3D generative models trained on a large internal OpenAI 3D dataset, both of which are able to convert a single-view image into a point cloud or an implicit representation. SyncDreamer aims to generate multi-view consistent images from a single image for deriving 3D geometry.
3 Evaluation Protocol
Evaluation Datasets. Following prior research , we adopt the Google Scanned Object dataset for our evaluation, which includes a wide variety of common everyday objects. Our evaluation dataset matches that of SyncDreamer , comprising 30 objects that span from everyday items to animals. For each object in the evaluation set, we render an image with a size of 256×256, which serves as the input. Additionally, to assess the generalization ability of our model, we include some images with diverse styles collected from the internet in our evaluation.
Metrics. To evaluate the quality of the single-view reconstructions, we adopt two commonly used metrics Chamfer Distances (CD) and Volume IoU between ground-truth shapes and reconstructed shapes. Since different methods adopt various canonical systems, we first align the generated shapes to the ground-truth shapes before calculating the two metrics. Moreover, we adopt the metrics PSNR, SSIM and LPIPS for evaluating the generated color images.
4 Single View Reconstruction
We evaluate the quality of the reconstructed geometry of different methods. The quantitative results are summarized in Table 1, and the qualitative comparisons are presented in Fig. 6. Shap-E tends to produce incomplete and distorted meshes. SyncDreamer generates shapes that are roughly aligned with the input image but lack detailed geometries, and the texture quality is subpar. One-2-3-45 attempts to reconstruct meshes from the multiview-inconsistent outputs of Zero123 . While it can capture coarse geometries, it loses important details in the process. In comparison, our method stands out by achieving the highest reconstruction quality, both in terms of geometry and textures.
5 Novel View Synthesis
We evaluate the quality of novel view synthesis for different methods. The quantitative results are presented in Table 2, and the qualitative results can be found in Figure 3. Zero123 produces visually reasonable images, but they lack multi-view consistency since it operates on each view independently. Although SyncDreamer introduces a volume attention scheme to enhance the consistency of multi-view images, their model is sensitive to the elevation degrees of the input images and tends to produce unreasonable results. In contrast, our method is capable of generating images that not only exhibit semantic consistency with the input image but also maintain a high degree of consistency across multiple views in terms of both colors and geometry.
6 Discussions
In this section, we conduct a set of studies to verify the effectiveness of our designs as well as the properties of the method.
To validate the effectiveness of our proposed cross-domain diffusion scheme, we study the following settings: (a) cross-domain model with cross-domain attention; (b) cross-domain model without cross-domain attention; (c) sequential model rgb-to-normal: first train a multi-view color diffusion model then train a multi-view normal diffusion model conditioned on the previously generated color images; (d) sequential model normal-to-rgb: first train a multi-view normal diffusion model then train a multi-view color diffusion model conditioned on the previously generated normal images.
As shown in (a) and (b) of Figure 7, it’s evident that the cross-domain attentions significantly enhance the consistency between color images and normals, particularly in terms of the detailed geometries of objects like the ice-cream and Pharaoh sculpture. From (c) and (d) of Figure 7, while the normals and color images generated by sequential models maintain some consistency, their results suffer from performance drops.
For the sequential model rgb-to-normal, conditioning on the separately generated normal maps, the generated color images exhibit color aberrations in comparison to the input image, as shown in (c) of Figure 7. Conversely, for the sequential model normal-to-rgb, conditioning on the separately generated color images, the normal maps give unreasonable geometry, as illustrated in (d) of Figure 7. These experiments demonstrate that jointly predicting normal maps and color images through the cross-domain attention mechanism can facilitate a comprehensive perception of information from different domains. We also speculate that in the context of sequential models, the generated color images or normal maps of stage 1 may exhibit a minor domain gap to the ground truth data trained in stage 2. Therefore, compared to sequential prediction, the cross-domain approach proves to be more effective in enhancing the quality of each domain as well as the overall prediction.
Multi-view Consistency. We conducted an analysis of the effectiveness of the multi-view attention mechanism, as illustrated in Figure 9. Our findings show that the multi-view attention greatly enhances the 3D consistency of the generated multi-view images, particularly for the rear views. In the absence of the multi-view attention, the color images of the rear views exhibited unrealistic predictions.
Normal Fusion. To assess the efficacy of our normal fusion algorithm, we conducted experiments using the complex lion model, which is rich in geometric details, as illustrated in Figure 8. The baseline model’s surfaces exhibited numerous holes and noises. Utilizing either the geometry-aware normal loss or the outlier-dropping loss helps mitigate the noisy surfaces. Finally, combining both strategies yields the best performance, resulting in clean surfaces while preserving detailed geometries.
Generalization. To demonstrate the generalization capability of our method, we conducted evaluations using diverse image styles, including sketches, cartoons, and images of animals, as shown in Figure 5 and Figure 10. Despite variations in lighting effects and geometric complexities among these images, our method consistently generated multi-view normal maps and color images, ultimately yielding high-quality geometries.
Conclusions and Future Works
Conclusions. In this paper, we present Wonder3D, an innovative approach designed for efficiently generating high-fidelity textured meshes from single-view images. When provided with a single image, Wonder3D initiates the process by generating consistent multi-view normal maps and paired color images. Subsequently, it utilizes a novel normal fusion algorithm to extract highly-detailed geometries from these multi-view 2D representations. Experimental results demonstrate that our method upholds good efficiency and robust generalization, and delivers high-quality geometry.
Future Works. While Wonder3D has demonstrated promising performance in reconstructing 3D geometry from single-view images, there are still some limitations that the current framework does not fully address. First, the current implementation of Wonder3D only produces normals and color images from six views. This limited number of views makes it challenging for our method to accurately reconstruct objects with very thin structures and severe occlusions. Additionally, expanding Wonder3D to incorporate more views would demand increased computational resources during training. To address this issue, Wonder3D may benefit from leveraging more efficient multi-view attention mechanisms to handle a greater number of views effectively.
Acknowledgements
Thanks for the GPU support from VAST, the valuable suggestions from Wei Yin, the help in data rendering from Dehu Wang.