SDFusion: Multimodal 3D Shape Completion, Reconstruction, and Generation
Yen-Chi Cheng, Hsin-Ying Lee, Sergey Tulyakov, Alexander Schwing, Liangyan Gui
Introduction
Generating 3D assets is a cornerstone of immersive augmented/virtual reality experiences. Without realistic and diverse objects, virtual worlds will look void and engagement will remain low. Despite this need, manually creating and editing 3D assets is a notoriously difficult task, requiring creativity, 3D design skills, and access to sophisticated software with a very steep learning curve. This makes 3D asset creation inaccessible for inexperienced users. Yet, in many cases, such as interior design, users more often than not have a reasonably good understanding of what they want to create. In those cases, an image or a rough sketch is sometimes accompanied by text indicating details of the asset, which are hard to express graphically for an amateur.
Due to this need, it is not surprising that democratizing the 3D content creation process has become an active research area. Conventional 3D generative models require direct 3D supervision in the form of point clouds , signed distance functions (SDFs) , voxels , etc. Recently, first efforts have been made to explore the learning of 3D geometry from multi-view supervision with known camera poses by incorporating inductive biases via neural rendering techniques . While compelling results have been demonstrated, training is often very time-consuming and ignores available 3D data that can be used to obtain good shape priors. We foresee an ideal collaborative paradigm for generative methods where models trained on 3D data provide detailed and accurate geometry, while models trained on 2D data provide diverse appearances. A first proof of concept is shown in Figure 1.
In our pursuit of flexible and high-quality 3D shape generation, we introduce SDFusion, a diffusion-based generative model with a signed distance function (SDF) under the hood, acting as our 3D representation. Compared to other 3D representations, SDFs are known to represent well high-resolution shapes with arbitrary topology. However, 3D representations are infamous for demanding high computational resources, limiting most existing 3D generative models to voxel grids of resolution and point clouds of points. To side-step this issue, we first utilize an auto-encoder to compress 3D shapes into a more compact low-dimensional representation. Because of this, SDFusion can easily scale up to a resolution. To learn the probability distribution over the introduced latent space, we leverage diffusion models, which have recently been used with great success in various 2D generation tasks . Furthermore, we adopt task-specific encoders and a cross-attention mechanism to support multiple conditioning inputs, and apply classifier-free guidance to enable flexible conditioning usage. Because of these strategies, SDFusion can not only use a variety of conditions from multiple modalities, but also adjust their importance weight, as shown in Figure 1. Compared to a recently proposed autoregressive model that also adopts an encoded latent space, SDFusion achieves superior sample quality, while offering more flexibility to handle multiple conditions and, at the same time, features reduced memory usage. With SDFusion, we study the interplay between models trained on 2D and 3D data. Given 3D shapes generated by SDFusion, we take advantage of an off-the-shelf 2D diffusion model , neural rendering , and score distillation sampling to texture the shapes given text descriptions as conditional variables.
We conduct extensive experiments on the ShapeNet , BuildingNet , and Pix3D datasets. We show that SDFusion quantitatively and qualitatively outperforms prior work in shape completion, 3D reconstruction from images, and text-to-shape tasks. We further demonstrate the capability of jointly controlling the generative model via multiple conditioning modalities, the flexibility of adjusting relative weight among modalities, and the ability to texture 3D shapes given textual descriptions, as shown in Figure 1.
We summarize the main contributions as follows:
We propose SDFusion, a diffusion-based 3D generative model which uses a signed distance function as its 3D representation and a latent space for diffusion.
SDFusion enables conditional generation with multiple modalities, and provides flexible usage by adjusting the weight among modalities.
We demonstrate a pipeline to synthesize textured 3D objects benefiting from an interplay between 2D and 3D generative models.
Related Work
Different from 2D images, it is less clear how to effectively represent 3D data. Indeed, various representations with different pros and cons have been explored, particularly when considering 3D generative models. For instance, 3D generative models have been explored for point clouds , voxel grids , meshes , signed distance functions (SDFs) , etc. In this work, we aim to generate an SDF. Compared to other representations, SDFs exhibit a reasonable trade-off regarding expressivity, memory efficiency, and direct applicability to downstream tasks. Moreover, conditioning 3D generation of SDFs on different modalities further enables many applications, including shape completion, 3D reconstruction from images, 3D generation from text, etc. The proposed framework can handle these tasks in a single model which makes it different from prior work.
Recently, thanks to the advancement of neural rendering , a new stream of research has emerged to learn 3D generation and manipulation from only 2D supervision . We believe the interplay between two streams of work is promising in the foreseeable future.
Diffusion Models.
Diffusion models have recently emerged as a popular family of generative models with competitive sample quality. In particular, diffusion models have shown impressive quality, diversity, and expressiveness in various tasks such as image synthesis , super-resolution , image editing , text-to-image synthesis , etc. In contrast to the flourishing research on diffusion models for 2D data, diffusion models have not yet been fully explored for 3D data. Notable exceptions include attempts to apply diffusion models to point clouds .
Differently, in this work, we apply diffusion models on SDF representations. As reasonable resolutions of SDFs are demanding to model, we study the use of a latent diffusion technique and the classifier-free conditional generation mechanism , both of which have been shown to yield promising results when being used in 2D diffusion models.
Approach
We aim at synthesizing 3D shapes using diffusion models. Towards this goal, we model the distribution over 3D shapes , a volumetric Truncated Signed Distance Field (T-SDF). However, applying diffusion models directly on reasonably high-resolution 3D shapes is computationally very demanding. Therefore, we first compress the 3D shape into a discretized and compact latent space (Section 3.1). This allows us to apply diffusion models in a lower-dimensional space (Section 3.2). The proposed framework can further incorporate various user conditions such as partial shapes, images, and text (Section 3.3). Finally, we showcase an interplay between the proposed framework and diffusion models trained on 2D data to texture 3D shapes (Section 3.4).
2 Latent Diffusion Model for SDF
Using the trained encoder , we can now encode any given SDF into a compact and low-dimensional latent variable . We can then train a diffusion model on this latent representation. Fundamentally, a diffusion model learns to sample from a target distribution by reversing a progressive noise diffusion process. Given a sample , we obtain by gradually adding Gaussian noise with a variance schedule. Then we use a time-conditional 3D UNet as our denoising model. To train the denoising 3D UNet, we adopt the simplified objective proposed by Ho et al. :
At inference time, we sample by gradually denoising a noise variable sampled from the standard normal distribution , and leverage the trained decoder to map the denoised code back to a 3D T-SDF shape representation , as shown in Figure 2.
3 Learning the Conditional Distribution
Being able to randomly sample shapes provides limited ability for interaction. Therefore, learning of a conditional distribution is essential for user applications. Importantly, multiple forms of conditional inputs are desirable such that the model can account for various kinds of scenarios. Thanks to the flexible conditional mechanism provided by a latent diffusion model , we can incorporate multiple conditional input modalities at once with task-specific encoders and a cross-attention module. To further allow for more flexibility in controlling the distribution, we adopt classifier-free guidance for conditional generation. The objective function reads as follows:
where is the task-specific encoder for the modality, is a dropout operation enabling classifier-free guidance, and is a feature aggregation function. In this work, refers to a simple concatenation.
At inference time, given conditions from multiple modalities, we perform classifier-free guidance as follows
where denotes the weight of conditions from the modality and denotes a condition filled with zeros. Intuitively, modalities with larger weights play more important roles in guiding the conditional generation.
In this work, we study SDFusion combined with three conditional modalities applied separately or jointly. For shape completion, given a partial observation of a shape, we perform blended diffusion similar to . For single-view 3D reconstruction, we adopt CLIP as the image encoder. For text-guided 3D generation, we adopt BERT as the text encoder. The encoded features are then used to modulate the diffusion process with cross-attention.
4 3D Shape Texturing with a 2D Model
While sampled 3D data often exhibits compellingly detailed geometry, textures of 3D data are generally more difficult to collect and even more often of limited quality. Here, we explore how to make the best use of 2D data and models, so as to aid 3D asset generation. Thanks to the recent success of neural rendering and 2D text-to-image models trained on extremely large-scale data , using a 2D model to perform 3D synthesis is made possible with a score distillation sampling technique .
where is a time-dependent weight function defined in the stable diffusion model. The mechanism is called Score Distillation Sampling. Please refer to for details. The score then provides an update direction to . We illustrate the process in Figure 3.
Experiments
In this section, we conduct extensive qualitative and quantitative experiments to demonstrate the efficacy and generalizability of SDFusion. We evaluate methods on three tasks: shape completion, single-view 3D reconstruction, and text-guided generation. We then demonstrate two additional use cases: multi-conditional generation and 3D shape texturing.
We evaluate the shape completion task on the ShapeNet and the BuildingNet datasets. ShapeNet is a large-scale 3D CAD model dataset with 16 common object classes. We use the train/test splits provided by Xu et al. . BuildingNet is a new large-scale 3D building model dataset. Compared to objects in ShapeNet, building models provide more geometric details and thus require higher resolution for representing the data. Hence, we use an SDF of resolution for ShapeNet, and a resolution for BuildingNet. For both datasets, we compare the completion quantitatively by providing the bottom half of the ground truth shape as input, and evaluating the completed shapes generated by different methods.
We compare SDFusion to the state-of-the-art point cloud completion method MPC and the autoregressive SDF generation method AutoSDF . We adopt metrics from MPC . For each partial shape, we generate complete shapes. To evaluate completion fidelity, we measure the Unidirectional Hausdorff Distance (UHD) between partial shapes and generated shapes. To evaluate completion diversity, we measure the Total Mutual Difference (TMD) by computing the average Chamfer distance among generated shapes. We use in the experiments.
As shown in Table 1, the proposed SDFusion performs favorably compared to all methods in the completion fidelity metric, and outperforms the baselines substantially in completion diversity. The advantages of fidelity and diversity are also apparent in Figure 4. Especially on the BuildingNet dataset, SDFusion shows its advantages in modeling high-resolution and diverse data. AutoSDF and MPC struggle to model the distribution correctly.
2 Single-view 3D Reconstruction
Next, we assess 3D shape reconstruction from a single image on the real-world benchmark Pix3D dataset. We use the provided train/test splits on the chair category. In the absence of official splits for other categories, we randomly split the dataset into disjoint train/test splits.
We compare with the ResNet2TSDF and ResNet2Voxel baselines. Both encode images and directly output 3D shapes in the form of a T-SDF and a voxel grid. We also compare to two state-of-the-art methods for 3D reconstruction, i.e., Pix2Vox and AutoSDF . We evaluate all methods after aligning resolutions to voxels. We use Chamfer Distance (CD), and F-score as evaluation metrics.
Quantitatively, as shown in Table 2, SDFusion outperforms other methods on both metrics. Qualitatively, SDFusion generates 3D shapes that are of higher visual quality and are more visually consistent with the objects shown in the images, regardless of camera poses, as shown in Figure 5.
3 Text-guided Generation
Next, we evaluate 3D shape generation conditioned on text input. For a qualitative comparison, we use the Text2shape dataset that provides descriptions for the ‘chair’ and ‘table’ categories in ShapeNet. For a quantitative evaluation, we adopt the ShapeGlot dataset which provides text utterances describing the difference between a target shape and two distractors based on the ShapeNet dataset. We compare SDFusion with AutoSDF, which recently demonstrated state-of-the-art results on the text-guided 3D shape generation task.
We follow the evaluation pipeline proposed by ShapeGlot . We train a neural evaluator to distinguish the target shape from a distractor given the description. Given two shapes from different methods, the neural evaluator provides a confidence score for each of them based on the binary classification logits. For the absolute difference between two confidence scores , we count the comparison as confused (conf.).
As shown in Table 3, SDFusion quantitatively outperforms AutoSDF by a large margin with a low confusion rate. SDFusion also performs better than AutoSDF when compared with ground truth data. Qualitatively, we show in Figure 5 that SDFusion not only generates shapes with better quality, but that the generated shapes are also more diverse. Notably, SDFusion reacts to very specific descriptions like “L-shaped table” and “table with two surfaces.” It generates objects of high diversity while remaining faithful to the provided description.
4 Multi-conditional Generation
In addition to the conditional generation tasks which take a single conditioning variable as input, we further demonstrate the efficacy of SDFusion in handling multiple modalities. First, SDFusion can jointly consider multiple conditioning modalities. On the left of Figure 7, we present the diverse generation conditional on both partial shapes and text. On the right of Figure 7, we show that given partial shapes and images, SDFusion can complete the different parts based on images. When there is ambiguity in images (e.g., rear-view of chairs), SDFusion can produce diverse predictions. Second, as shown in Figure 8, SDFusion can not only be jointly conditioned on multiple inputs, but a weight can be used to control the importance of the conditioning modalities, enabling more flexible user control. For example, for the left sample in Figure 8, the larger the weight for the input image, the more similar the results are to the shapes in the image. Similarly, the larger the weight for input text, the more “egg-shaped” the results. We envision such a fine-grained form of control to be particularly useful for interactive user applications.
5 3D Shape Texturing
Finally, we showcase an application that uses SDFusion to generate 3D shapes of detailed geometry, and uses a pretrained text-to-image 2D diffusion model to provide textures. As shown in Figure 9, the diffusion model pre-trained on large-scale 2D data can provide semantically meaningful and diverse guidance to texture the 3D shapes. The model shows superior expressiveness to interpret abstract concepts (e.g., Chinese- and Halloween-style) and materials (e.g., gingerbread, cream cheese). Given a single description, the texturing pipeline can also generate diverse results, as shown in the rightmost part of Figure 9.
Conclusion
In this work, we present SDFusion, an attempt to adopt diffusion models to signed distance functions for 3D shape generation. To alleviate the computationally demanding nature of 3D representations, we first encode 3D shapes into an expressive low-dimensional latent space, which we use to train the diffusion model. To enable flexible conditional usage, we adopt class-specific encoders along with a cross-attention mechanism for handling conditions from multiple modalities, and leverage classifier-free guidance to facilitate weight control among modalities. Foreseeing the potential of a collaborative symbiosis between models trained on 2D and 3D data, we further demonstrate an application that takes advantage of a pretrained 2D text-to-image model to texture a generated 3D shape.
Although the results look promising and exciting, there are quite a few future directions for improvement. First, SDFusion is trained on high-quality signed distance function representations. To make the model more general and to enable the use of more diverse data, a model that operates on various 3D representations simultaneously is desirable. Another future direction is related to the diversity of the data: we currently apply SDFusion on object-centric data. It is interesting to apply the model to more challenging scenarios (e.g., entire 3D scenes). Finally, we believe there is room to further explore how to combine models trained on 2D and 3D data.
Acknowledgements: Work supported in part by NSF under Grants 2008387, 2045586, 2106825, MRI 1725729, and NIFA award 2020-67021-32799. Thanks to NVIDIA for providing a GPU for debugging.