Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors

Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Skorokhodov, Peter Wonka, Sergey Tulyakov, Bernard Ghanem

Introduction

Despite observing the world in 2D, human beings have a remarkable capability to navigate, reason, and engage with their 3D surroundings. This points towards a deep-seated cognitive understanding of the characteristics and behaviors of the 3D world - a truly impressive facet of human nature. This ability is taken to another level by artists who can produce detailed 3D replicas from a single image. Contrarily, from the perspective of computer vision, the task of 3D reconstruction from an unposed image - which encompasses the creation of geometry and textures - remains an unresolved, ill-posed problem, despite decades of exploration and development .

The recent advances in deep learning have allowed an increasing number of 3D generation tasks to become learning-based. Even though deep learning has accomplished significant strides in image recognition and generation , the particular task of single-image 3D reconstruction in the wild is still lagging. We attribute this considerable discrepancy in 3D reconstruction abilities between humans and machines to two primary factors: (i) a deficiency in large-scale 3D datasets that impedes large-scale learning of 3D geometry, and (ii) the trade-off between the level of detail and computational resources when working on 3D data.

One possible approach to tackle the problem is to employ 2D priors. The pool of realistic 2D image data available online is voluminous. LAION , one of the most extensive text-image pair datasets, aids in training modern image understanding and generation models like CLIP and Stable Diffusion . With the increasing generalization capabilities of 2D generation models, there has been a notable rise in approaches that use 2D models as priors for generating 3D content. DreamFusion serves as a trailblazer for this 2D prior-based methodology for text-to-3D generation. The technique demonstrates an exceptional capacity to guide novel views and optimize a neural radiance field (NeRF) in a zero-shot setting. Drawing upon DreamFusion, recent work such as RealFusion and NeuralLift , have endeavored to adapt these 2D priors for single image 3D reconstructions.

Another approach is to employ 3D priors. Earlier attempts at 3D reconstruction leveraged 3D priors like topology constraints to assist in 3D generation . However, these manually-crafted 3D priors fall short of generating high-quality 3D content. Recently, approaches like Zero-1-to-3 and 3Dim adapted a 2D diffusion model to become view-dependent and utilized this view-dependent diffusion as a 3D prior.

We analyzed the behavior of both 2D and 3D priors and found that they both have advantages and disadvantages. 2D priors exhibit impressive generalization for 3D generation that is unattainable with 3D priors (e.g., the dragon statue example in Fig.2). However, methods relying on 2D priors alone inevitably compromise on 3D fidelity and consistency due to their restricted 3D knowledge. This leads to unrealistic geometry like multiple faces (Janus problems), mismatched sizes, inconsistent texture, and so on. An instance of a failure case can be observed in the teddy bear example in Fig.2. On the other hand, a strict reliance on 3D priors alone is unsuitable for in-the-wild reconstruction due to the limited 3D training data. Consequently, as illustrated in Fig.2, while 3D prior-based solution effectively processes common objects (for instance, the teddy bear example in the top row), it struggles with less common ones, yielding oversimplified, sometimes even flat 3D geometries (e.g., dragon statue at bottom left).

In this paper, rather than solely relying on a 2D or 3D prior, we advocate for the simultaneous use of both priors to guide novel views in image-to-3D generation. By modulating the simple yet effective tradeoff parameter between the potency of the 2D and 3D priors, we can manage the balance between exploration and exploitation in the generated 3D geometry. Prioritizing the 2D prior can enhance imaginative 3D capabilities to compensate for the incomplete 3D information inherent in a single 2D image, but this may result in less accurate 3D geometry due to a lack of 3D knowledge. In contrast, prioritizing the 3D prior can lead to more 3D-constrained solutions, generating more accurate 3D geometry, albeit with reduced imaginative capabilities and diminished ability to discover plausible solutions for challenging and uncommon cases. We introduce Magic123, a novel image-to-3D pipeline that yields high-quality 3D outputs through a two-stage coarse-to-fine optimization process utilizing both 2D and 3D priors. In the coarse stage, we optimize a neural radiance field (NeRF) . NeRF learns an implicit volume representation, which is highly effective for complex geometry learning. However, NeRF demands significant memory, resulting in low-resolution rendered images passed to the diffusion models, making the output for the image-to-3D task low-quality. Even the more resource-efficient NeRF alternative, Instant-NGP , can only reach a resolution of 128×128128\times 128 in the image-to-3D pipeline on a 16GB memory GPU. Hence, to improve the quality of the 3D content, we introduce a second stage, employing a memory-efficient and texture-decomposed SDF-Mesh hybrid representation known as Deep Marching Tetrahedra (DMTet) . This approach enables us to increase the resolution up to 1K and refine the geometry and texture of the NeRF separately. In both stages, we leverage a combination of 2D and 3D priors to guide the novel views.

We summarize our contributions as follows:

We introduce Magic123, a novel image-to-3D pipeline that uses a two-stage coarse-to-fine optimization process to produce high-quality high-resolution 3D geometry and textures.

We propose to use 2D and 3D priors simultaneously to generate faithful 3D content from any given image. The strength parameter of priors allows for the trade-off between geometry exploration and exploitation. Users therefore can play with this trade-off parameter to generate desired 3D content.

Moreover, we find a balanced trade-off between 2D and 3D priors, leading to reasonably realistic and detailed 3D reconstructions. Using the exact same set of parameters for all examples without any additional reconfiguration, Magic123 achieves state-of-the-art results in 3D reconstruction from single unposed images in both real-world and synthetic scenarios.

Methodology

We propose a two-stage framework, Magic123, that generates 3D content from a single reference image in a coarse to fine fashion, as shown in Fig. 3. In the coarse stage, Magic123 learns a coarse geometry and texture by optimizing a NeRF. In the fine stage, Magic123 improves the quality of 3D content by directly optimizing a memory-efficient differentiable mesh representation with high-resolution renderings. In both stages, Magic123 uses joint 2D and 3D diffusion priors to trade off geometry exploration and geometry exploitation, yielding reliable 3D content with high generalizability.

Image preprocessing. Magic123 is a pipeline for object-level image-to-3D generation. Given an image with a background, Magic123 requires a preprocessing step to extract the foreground object. We leverage an off-the-shelf segmentation model, Dense Prediction Transformer , to segment the object. The extracted mask, denoted as M\mathbf{M} is a binary segmentation mask and will be used in the optimization. To prevent flat geometry collapse, \iethe model generates textures that only appear on the surface without capturing the actual geometric details, we further extract the depth map from the reference view by the pretrained MiDaS . The foreground image is used as the input, while the mask and the depth map are used in the optimization as regularization priors. These reference images are assigned fixed camera poses, assumed to the front view. More details in camera settings can be found in Sec.3.2.

The coarse stage of our Magic123 is aimed at learning underlying geometry that respects the reference image. Due to its strong ability in handling complex topological changes in a smooth and continuous fashion, we adopt NeRF in this stage.

Instant-NGP and its optimization. We leverage Instant-NGP as our NeRF implementation because of its fast inference and ability to recover complex geometry. To reconstruct 3D faithfully from a single image, the optimization of NeRF requires at least two loss functions: (i) the reference view reconstruction supervision; and (ii) the novel view guidance.

Reference view reconstruction loss Lrec\mathcal{L}_{rec} is imposed in our pipeline as one of the major loss functions to ensure the rendered image from the reference viewpoint (vr\mathbf{v}^{r}, assumed to be front view) is as close to the reference image Ir\mathbf{I}^{r} as possible. We adopt the mean squared error (MSE) loss on both the reference image and its mask as follows:

where θ\theta is the NeRF parameters to be optimized, ⊙\odot is Hadamard product, Gθ(vr)G_{\theta}(\mathbf{v}^{r}) is NeRF rendered view from vr\mathbf{v}^{r} viewpoint, M()M() is the foreground mask acquired by integrating the volume density along the ray of each pixel. Since the foreground object is extracted as input, we do not model any background and simply use pure white for the background rendering for all experiments. λrgb,λmask\lambda_{rgb},\lambda_{mask} are the weights for the foreground RGB and the mask.

Novel view guidance Lg\mathcal{L}_{g} is necessary since multiple views are required to train a NeRF. We follow the pioneering work in text/image-to-3D and use diffusion priors to guide the novel view generation. As a significant difference from previous works, we do not rely solely on a 2D prior or a 3D prior, but we use both of them to guide the optimization of the NeRF. See §2.2 for details.

Depth prior Ld\mathcal{L}_{d} is exploited to avoid overly-flat or caved-in 3D content. Using only the appearance reconstruction losses might yield poor geometry due to the inherent ambiguity of reconstructing 3D content from 2D images: the content of 3D may lie at any distance and still be rendered as the same image. This ambiguity might result in flat or curved-in geometry as noted in previous works . We alleviate this issue by leveraging a depth regularization. A pretrained monocular depth estimator is leveraged to acquire the pseudo depth drd^{r} on the reference image. The depth output dd from the NeRF model from the reference viewpoint should be close to the depth prior. However, due to the value mismatch of two different sources of depth estimation, an MSE loss is not an ideal loss function. We use the normalized negative Pearson correlation as the depth regularization:

where cov(⋅)\text{cov}(\cdot) denotes covariance and σ(⋅)\sigma(\cdot) measures standard deviation.

Normal smoothness Ln\mathcal{L}_{n}. One of the NeRF limitations is the tendency to produce high-frequency artifacts on the surface of the object. To this end, we enforce the smoothness of the normal maps of geometry for the generated 3D model following . We use the finite differences of the depth to estimate the normal vector of each point, render a 2D normal map n\mathbf{n} from the normal vector, and impose a loss as follows:

where τ(⋅)\tau(\cdot) denotes the stopgradient operation and g(⋅)g(\cdot) is a Gaussian blur. The kernel size of the blurring kk is set to 9×99\times 9.

Overall, the coarse stage is optimized by a combination of losses:

where λd,λn\lambda_{d},\lambda_{n} are the weights of depth and normal regularizations.

1.2 Fine stage

The coarse stage offers a low-resolution 3D model, possibly with noise due to the tendency of NeRF to create high-frequency artifacts. Our fine stage aims to refine the 3D model and obtain a high-resolution and disentangled geometry and texture. To this end, we adopt DMTet , which is a hybrid SDF-Mesh representation and is capable of generating high-resolution 3D shapes due to its high memory efficiency. Note the fine stage is identical to the coarse stage except for the 3D representation and rendering.

2 Joint 2D and 3D priors for image-to-3D generation

2D priors. Using a single reference image is insufficient to train a complete NeRF model without any priors . To address this issue, DreamFusion proposes to use a 2D diffusion model as the prior to guide the novel views via the proposed score distillation sampling (SDS) loss. SDS exploits a 2D text-to-image diffusion model , encodes the rendered view as latent, adds noise to it, and guesses the clean novel view guided by the input text prompt. Roughly speaking, SDS translates the rendered view into an image that respects both the content from the rendered view and the prompt. The SDS loss is illustrated in the upper part of Fig. 4 and is formulated as:

where I\mathbf{I} is a rendered view, and zt\mathbf{z}_{t} is the noisy latent by adding a random Gaussian noise of a time step tt to the latent of I\mathbf{I}. ϵ,ϵϕ\epsilon,\epsilon_{\phi}, ϕ\phi, θ\theta are the added noise, predicted noise, parameters of the 2D diffusion prior, and the parameters of the 3D model. θ\theta can be MLPs of NeRF for the coarse stage, or SDF, triangular deformations, and color field for the fine stage. DreamFusion further points out that the Jacobian term of the image encoder ∂z∂I\frac{\partial\mathbf{z}}{\partial\mathbf{I}} in Eq. (5) can be further eliminated, making the SDS loss much more efficient in terms of both speed and memory. In our experiments, we utilize the SDS loss with Stable Diffusion v1.5 as our 2D prior. The rendered images are interpolated to 512×512512\times 512 as required by the image encoder in .

Textural inversion. Note the prompt e\mathbf{e} we use for each reference image is not a pure text chosen from tedious prompt engineering. Using pure text for image-to-3D generation most likely results in inconsistent texture and geometry due to the limited expressiveness of the human language. For example, using “A high-resolution DSLR image of a colorful teapot” will generate different geometry and colors that do not respect the reference image. We thus follow RealFusion to leverage the same textual inversion technique to acquire a special token to represent the object in the reference image. We use the same prompt for all examples: “A high-resolution DSLR image of ”. We find that Stable Diffusion can generate the teapot with a more similar texture and style to the reference image with the textural inversion technique compared to the results without it.

Overall, the 2D diffusion priors exhibit a remarkable capacity for exploring the space of geometry, thereby facilitating the generation of diverse geometric representations with a heightened sense of imagination. This exceptional imaginative capability compensates for the inherent limitations associated with the availability of incomplete 3D information in a single 2D image. Moreover, the utilization of 2D prior-based techniques for 3D reconstruction reduces the likelihood of overfitting in certain scenarios, owing to their training on an extensive dataset comprising over a billion images. However, it is crucial to acknowledge that the reliance on 2D priors may introduce inaccuracies in the generated 3D representations, thereby potentially deviating from true fidelity. This low-fidelity generation happens because 2D priors lack 3D knowledge. For instance, the utilization of 2D priors may yield imprecise geometries, such as Janus problems and mismatched sizes as depicted in Fig. 2 and Fig. 8.

Using only the 2D prior is not sufficient to capture detailed and consistent 3D geometry. Zero-1-to-3 thus proposes a 3D prior solution. Zero-1-to-3 finetunes Stable Diffusion into a view-dependent version on Objaverse , the largest open-source 3D dataset that consists of 818K models. Zero-1-to-3 takes a reference image and a viewpoint as input and can generate a novel view from the given viewpoint. Zero-1-to-3 thereby can be used as a strong 3D prior for 3D reconstruction. The usage of Zero-1-to-3 in an image-to-3D generation pipeline using SDS loss is formulated as:

where R,TR,T are the camera poses passed to Zero-1-to-3, the view-dependent diffusion model. The difference between using the 3D prior and the 2D prior is illustrated in Fig. 4, where we show that the 2D prior uses text embedding as guidance while the 3D prior uses the reference view Ir\mathbf{I}^{r} with the novel view camera poses as guidance. The 3D prior utilizes camera poses to encourage 3D consistency and enable the usage of more 3D information compared to the 2D prior.

Overall, the utilization of 3D priors demonstrates a commendable capacity for effectively harnessing the expansive realm of geometry, resulting in the generation of significantly more accurate geometric representations compared to their 2D counterparts. This heightened precision particularly applies when dealing with objects that are commonly encountered within the pre-trained 3D dataset. However, it is essential to acknowledge that the generalization capability of 3D priors is comparatively lower than that of 2D priors, thereby potentially leading to the production of geometric structures that may appear implausible. This low generalization results from the limited scale of available 3D datasets, especially in the case of high-quality real-scanned objects. For instance, in the case of uncommon objects, the employment of Zero-1-to-3 often tends to yield overly simplified geometries, \egflat surfaces without details in the back view (see Fig. 2 and Fig. 8).

2.2 Joint 2D and 3D priors

We find that the 2D and 3D priors are complementary to each other. Instead of relying solely on 2D or 3D prior, we propose to use both priors in 3D generation. The 2D prior is used to explore the geometry space, favoring high imagination but might lead to inaccurate geometry. We name this characteristic of the 2D prior as geometry exploration. On the other hand, the 3D prior is used to exploit the geometry space, constraining the generated 3D content to fulfill the implicit requirement of the underlying geometry, favoring precise geometry but with less generalizability. In the case of uncommon objects, the 3D prior might result in over-simplified geometry. We name this feature of using the 3D prior as geometry exploitation. In our image-to-3D pipeline, we propose a new prior loss for the novel view supervision to combine both 2D and 3D priors:

where λ2D/3D\lambda_{2D/3D} and λ3D\lambda_{3D} determine the strength of 2D and 3D prior, respectively. Weighting more on λ2D/3D\lambda_{2D/3D} leads to more geometry exploration, while weighting more on λ3D\lambda_{3D} results in more geometry exploitation. However, tuning two parameters at the same time is not user-friendly. Interestingly, through both qualitative and quantitative experiments, we find that Zero-1-to-3, the 3D prior we use, is much more tolerant to λ3D\lambda_{3D} than Stable Diffusion to λ2D\lambda_{2D}. When only the 3D prior is used, \ieλ2D=0\lambda_{2D}=0, Zero-1-to-3 generates consistent results for λ3D\lambda_{3D} ranging from 10 to 60. On the contrary, Stable Diffusion is rather sensitive to λ2D\lambda_{2D}. When setting λ3D\lambda_{3D} to and using the 2D prior only, the generated geometry varies a lot when λ2D\lambda_{2D} is changed from 11 to 22. This observation leads us to fix λ3D=40\lambda_{3D}=40 and to rely on tuning the λ2D\lambda_{2D} to trade off the geometry exploration and exploitation. We set λ2D/3D=1.0\lambda_{2D/3D}=1.0 for all results throughout the paper, but this value can be tuned according to the user’s preference. More details and discussions on the choice of 2D and 3D priors weights are available in Sec.3.4.

Experiments

NeRF4. We introduce a NeRF4 dataset that we collect from 4 scenarios, chair, drums, ficus, and microphone, out of the 8 test examples from the synthetic NeRF dataset . These four scenarios cover complex objects (drums and ficus), a hard case (the back view of the chair), and a simple case (the microphone). The other four examples are removed since they are not subject to the front view assumption, requiring further camera pose estimation or a manual tuning of the camera pose, which is out of the scope of this work.

RealFusion15. We further use the dataset collected and released by RealFusion , consisting of 15 natural images that include bananas, birds, cacti, barbie cakes, cat statues, teapots, microphones, dragon statues, fishes, cherries, and watercolor paintings \etc.

2 Implementation details

Optimizing the pipeline. We use exactly the same set of hyperparameters for all experiments and do not perform any per-object hyperparameter optimization. Both coarse and fine stages are optimized using Adam with 0.0010.001 learning rate and no weight decay for 5,0005,000 iterations. λrgb,λmask,λd\lambda_{rgb},\lambda_{mask},\lambda_{d} are set to 5,0.5,0.0015,0.5,0.001 for both stages. λ2D\lambda_{2D} and λ3D\lambda_{3D} are set to 11 and 4040 for the first stage and are lowered to 0.0010.001 and 0.010.01 in the second stage for refinement to alleviate oversaturated textures. We adopt the Stable Diffusion model of V1.5 as the 2D prior. The guidance scale of the 2D prior is set to 100100 following . For the 3D prior, Zero-1-to-3 (105,000105,000 iterations finetuned version) is leveraged. The guidance scale of Zero-1-to-3 is set to 55 following . The NeRF backbone is implemented by three layers of multi-layer perceptrons with 6464 hidden dims. Regarding lighting and shading, we keep nearly the same as . The difference is we set the first 3,0003,000 iterations in the first stage to normals’ shading to focus on learning geometry. For other iterations as well as the fine stage, we use diffuse shading with a probability 0.750.75 and textureless shading with a probability 0.250.25. The rendering resolutions are set to 128×128128\times 128 and 1024×10241024\times 1024 for the coarse and the fine stage, respectively.

Camera setting. Since the reference image is unposed, we assume its camera parameters are as follows. First, the reference image is assumed to be shot from the front view, \iepolar angle 90°90\degree, azimuth angle 0°0\degree. Second, the camera is placed 1.81.8 meters from the coordinate origin, \iethe radial distance is 1.81.8. Third, the field of view (FOV) of the camera is 40°\degree. We highlight that the 3D reconstruction performance is not sensitive to camera parameters, as long as they are reasonable, \egFOV between 2020 and 6060, and radial distance between 11 to 44 meters. Note this camera setting works for images subject to the front-view assumption. For images taken deviating from the front view, a manual change of polar angle or a camera estimation is required.

3 Results

Evaluation metrics. For a comprehensive evaluation, we adhere to the metrics employed in prior studies , namely PSNR, LPIPS , and CLIP-similarity . PSNR and LPIPS are gauged in the reference view to measure reconstruction quality and perceptual similarity. CLIP-similarity calculates an average CLIP distance between rendered image and the reference image to measure 3D consistency through appearance similarity across novel views and the reference view.

Quantitative and qualitative comparisons. We compare Magic123 against the state-of-the-art PointE , Shap-E , 3DFuse , NeuralLift , RealFusion and Zero-1-to-3 in both NeRF4 and RealFusion15 datasets. For Zero-1-to-3, we adopt the implementation here , which yields better performance than the original implementation. For other works, we use their officially released code. As shown in Table 1, Magic123 achieves Top-1 performance across all the metrics in both datasets when compared to previous approaches. It is worth noting that the PSNR and LPIPS results demonstrate significant improvements over the baselines, highlighting the exceptional reconstruction performance of Magic123. The improvement of CLIP-Similarity reflects the great 3D coherency regards to the reference view. Qualitative comparisons are available in Fig. 5. Magic123 achieves the best results in terms of both geometry and texture. Note how Magic123 greatly outperforms the 3D-based zero-1-to-3 especially in complex objects like the dragon statue and the colorful teapot in the first two rows, while at the same time greatly outperforming 2D-based RealFusion in all examples. This performance demonstrates the superiority of Magic123 over the state-of-the-art and its ability to generate high-quality 3D content.

4 Ablation and analysis

Magic123 introduces a coarse-to-fine pipeline for single image reconstruction and a joint 2D and 3D prior for novel view guidance. We provide analysis and ablation studies to show their effectiveness.

The effect of two stages. We study in Fig. 6 and Fig. 7 the effect of using the fine stage of our pipeline on the performance of Magic123. We note that a consistent improvement in terms of both qualitative and quantitative performance is observed throughout different setups when the fine stage is combined with the coarse stage. The use of a textured mesh DMTet representation enables higher quality 3D content that fits the objective and produces more compelling and higher resolution 3D consistent visuals.

3D priors only. We first turn off the guidance of the 2D prior by setting λ2D=0\lambda_{2D}=0, such that we only use the 3D prior Zero-1-to-3 as the guidance. We study the effects of λ3D\lambda_{3D} by setting it to 10,20,40,6010,20,40,60. Interestingly, we find that Zero-1-to-3 is very robust to the change of λ3D\lambda_{3D}. Tab. 2 demonstrates that different λ3D\lambda_{3D} lead to a consistent quantitative result. We thus simply set λ3D=40\lambda_{3D}=40 throughout the experiments since it achieves a slightly better CLIP-similarity score than other values.

2D priors only. We then turn off the 3D prior and study the effect of λ2D\lambda_{2D} in the image-to-3D task. As shown in Tab. 2, with the increase of λ2D\lambda_{2D}, an increase in CLIP-similarity is observed. This is due to the fact that a larger 2D prior weight leads to more imagination but unfortunately might result in the Janus problem.

Combining both 2D and 3D priors and the trade off factor λ2D/3D\lambda_{2D/3D}. In Magic123, we propose to use both 2D and 3D priors. Fig. 6 demonstrates the effectiveness of combining the 2D and 3D priors on the quantitative performance of image-to-3D generation. In Fig. 8, we further analyze the tradeoff hyperparameter λ2D/3D\lambda_{2D/3D} from Eq. (7). We start from λ2D/3D\lambda_{2D/3D}= to use only the 3D prior and gradually increase λ2D/3D\lambda_{2D/3D} to 0.1,0.5,1.0,2,50.1,0.5,1.0,2,5, and finally ∞\infty to use only the 2D prior with λ2D\lambda_{2D}=11 and λ3D\lambda_{3D}=. The key observations include: (1) Relying solely on the 3D prior results in precise geometry (as observed in the teddy bear) but falters in generating complex and uncommon objects, often rendering oversimplified geometry with minimal details (as seen in the dragon statue); (2) Relying solely on the 2D prior significantly improves performance in conjuring complex scenes like the dragon statue but simultaneously triggers the Janus problem in simple examples such as the bear; (3) As λ2D/3D\lambda_{2D/3D} escalates, the imaginative prowess of Magic123 is enhanced and more details become evident, but there is a tendency to compromise 3D consistency. We assign λ2D/3D\lambda_{2D/3D}=11 as the default value for all examples. However, this parameter could also be fine-tuned for even better results on certain inputs.

Related work

Multi-view 3D reconstruction. Multi-view 3D reconstruction aims to recover the 3D structure of a scene from its 2D RGB images captured from different camera positions . Classical approaches usually recover a scene’s geometry as a point cloud using SIFT-based point matching . More recent methods enhance them by relying on neural networks for feature extraction (\eg). The development of Neural Radiance Fields (NeRF) has prompted a shift towards reconstructing 3D as volume radiance , enabling the synthesis of photo-realistic novel views . Subsequent works have also explored the optimization of NeRF in few-shot (\eg) and one-shot (\eg) settings. NeRF does not store any 3D geometry explicitly (only the density field), and several works propose to use a signed distance function to recover a scene’s surface , including in the few-shot setting as well (\eg).

In-domain single-view 3D reconstruction. 3D reconstruction from a single view requires strong priors on the object geometry since even epipolar constraints cannot be imposed in such a setup. Direct supervision in the form of 3D shapes or keypoints is a robust way to impose such constraints for a particular domain, like human faces , heads , hands or full bodies . Such supervision requires expensive 3D annotations and manual 3D prior creation. Thus several works explore unsupervised learning of 3D geometry from object-centric datasets (\eg). These methods are typically structured as auto-encoders or generators with explicit 3D decomposition under the hood. Due to the lack of large-scale 3D data, these methods are limited to simple shapes (\egchairs, cars) and cannot generalize to more complex or uncommon objects (\egdragons, statues).

Zero-shot single-view 3D reconstruction. Foundational multi-modal networks have enabled various zero-shot 3D synthesis tasks. Earlier works employed CLIP guidance for 3D generation and manipulation from text prompts. Modern zero-shot text-to-image generators allowed to improve these results by providing stronger synthesis priors . DreamFusion is a seminal work that proposed to distill an off-the-shelf diffusion model into a NeRF for a given text query. It sparked numerous follow-up approaches for text-to-3D synthesis (\eg) and image-to-3D reconstruction (\eg). The latter is achieved via additional reconstruction losses on the frontal camera position and/or subject-driven diffusion guidance . The developed methods improved the underlying 3D representation and 3D consistency of the supervision ; explored task-specific priors and additional controls . Similar to the recent image-to-3D generators , we also follow the DreamFusion pipeline, but focus on reconstructing a high-resolution, textured 3D mesh using a joint 2D and 3D priors.

Conclusion and discussion

This work presents Magic123, a two-stage coarse-to-fine solution for generating high-quality, textured 3D meshes from a single unposed image. By leveraging both 2D and 3D priors, our approach overcomes the limitations of existing studies and achieves state-of-the-art results in image-to-3D reconstruction. The trade-off parameter between the 2D and 3D priors allows for control over the balance between exploration and exploitation of the generated geometry. Our method outperforms previous techniques in terms of both realism and level of detail, as demonstrated through extensive experiments on real-world images and synthetic benchmarks. Our findings contribute to narrowing the gap between human abilities in 3D reasoning and those of machines, and pave the way for future advancements in single image 3D reconstruction. The availability of our code, models, and generated 3D assets will further facilitate research and applications in this field.

Limitation. One of the limitations is that we assume the reference image is taken from the front view. This assumption leads to poor geometry when the reference image does not conform to the front-view assumption, \ega photo of a dish on the table taken from the up view. Our method will instead focus on generating the bottom of the dish and table instead of the dish geometry itself. This limitation can be alleviated by a manual reference camera pose tuning or camera estimation. Another limitation of our work is the dependency on the preprocessed segmentation and the monocular depth estimation model . Any error on these modules will creep into the later stages and affect the overall generation quality. Similar to previous work, Magic123 also tends to generate over-saturated textures due to the usage of the SDS loss. The over-saturation issue becomes more severe for the second stage because of the higher resolution.

Acknowledgement. The authors would like to thank Xiaoyu Xiang for the insightful discussion and Dai-Jie Wu for sharing Point-E and Shap-E results. This work was supported by the KAUST Office of Sponsored Research through the Visual Computing Center funding, as well as, the SDAIA-KAUST Center of Excellence in Data Science and Artificial Intelligence (SDAIA-KAUST AI). Part of the support is also coming from KAUST Ibn Rushd Postdoc Fellowship program.

References