Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion Models

Jiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao, Ying Shan, Xiaohu Qie, Shenghua Gao

Introduction

Text-to-3D synthesis endeavors to create 3D content that is coherent with an input text, which has the potential to benefit a wide range of applications such as animations, games, and virtual reality. Recently developed zero-shot text-to-image models have made remarkable progress and can generate diverse, high-fidelity, and imaginative images from various text prompts. However, extending this success to the text-to-3D synthesis task is challenging because it is not practically feasible to collect a comprehensive paired text-3D dataset.

Zero-shot text-to-3D synthesis , which eliminates the need for paired data, is an attractive approach that typically relies on powerful vision-language models such as CLIP . There are two main categories of this approach. 1) CLIP-based generative models, such as CLIP-Forge . They utilize images as an intermediate bridge and train a mapper from the CLIP image embeddings of ShapeNet renderings to the shape embeddings of a 3D shape generator, then switch to the CLIP text embedding as the input at test time. 2) CLIP-guided 3D optimization methods, such as DreamFields and PureCLIPNeRF . They continuously optimize the CLIP similarity loss between a text prompt and rendered images of a 3D scene representation, such as neural radiance fields . While the first category heavily relies on 3D shape generators trained on limited 3D shapes and seldom has the capacity to adjust its shape structures, the second category has more creative freedom with the “dreaming ability” to generate diverse shape structures and textures.

We develop our method building upon CLIP-guided 3D optimization methods. Although these methods can produce remarkable outcomes, they typically fail to create precise and accurate 3D structures that conform to the input text (Fig. 1, \nth2 row)). Due to the scratch training and random initialization without any prior knowledge, these methods tend to generate highly-unconstrained “adversarial contents” that have high CLIP scores but low visual quality. To address this issue and synthesize more faithful 3D contents, we suggest generating a high-quality 3D shape from the input text first and then using it as an explicit “3D shape prior” in the CLIP-guided 3D optimization process. In the text-to-shapeThroughout this paper, we use the term “shape” to refer to 3D geometric models without textures, while some works also use this term for textured 3D models. stage, we begin by synthesizing a 3D shape without textures of the main common object in the text prompt. We then use it as the initialization of a voxel-based neural radiance field and optimize it with the full prompt.

The text-to-shape generation itself is a challenging task. Previous methods are often trained on images and tested with texts, and use CLIP to bridge the two modalities. However, this approach leads to a mismatching problem due to the gap between the CLIP text and image embedding spaces. Additionally, existing methods cannot produce high-quality 3D shapes. In this work, we propose to directly bridge the text and image modalities with a powerful text-to-image diffusion model, i.e., Stable Diffusion . We use the text-to-image diffusion model to synthesize an image from the input text and then feed the image into an image-to-shape generator to produce high-quality 3D shapes. Since we use the same procedure in both training and testing, the mismatching problem is largely reduced. However, there is still a style domain gap between the images synthesized by Stable Diffusion and the shape renderings used to train the image-to-shape generator. Inspired by recent work on controllable text-to-image synthesis , we propose to jointly optimize a learnable text prompt and fine-tune the Stable Diffusion to address this domain gap. The fine-tuned Stable Diffusion can reliably synthesize images in the style of shape renderings used to train the image-to-shape module without suffering from the domain gap.

To summarize, 1) We make the first attempt to introduce the explicit 3D shape prior into CLIP-guided 3D optimization methods. The proposed method can generate more accurate and high-quality 3D shapes conforming to the corresponding text, while still enjoying the “dreaming” ability of generating diverse shape structures and textures (Fig. 1, \nth1 row). Therefore, we name our method “Dream3D” as it has both strengths. 2) Regarding text-to-shape generation, we present a straightforward yet effective approach that directly connects the text and image modalities using a powerful text-to-image diffusion model. To narrow the style domain gap between the synthesized images and shape renderings, we further propose to jointly optimize a learnable text prompt and fine-tune the text-to-image diffusion model for rendering-style image generation. 3) Our Dream3D can generate imaginative 3D content with better visual quality and shape accuracy than state-of-the-art methods. Additionally, our text-to-shape pipeline can produce 3D shapes of higher quality than previous work.

Related Work

3D Shape Generation. Generative models for 3D shapes have been extensively studied in recent years. It is more challenging than 2D image generation due to the expensive 3D data collection and the complexity of 3D shapes. Various 3D generators employ different shape representations, e.g., voxel grids , point clouds , meshes , and implicit fields . These generators are trained to model the distribution of shape geometry (and optionally, texture) from a collection of 3D shapes. Some methods attempt to learn a 3D generator using only 2D image supervision. These methods incorporate explicit 3D representations, such as meshes and neural radiance fields , along with surface or volume-based differentiable rendering techniques , to enable the learning of 3D awareness from images.

Text-to-Image. Previous studies in text-to-image synthesis have focused mainly on domain-specific datasets and utilized GANs . However, recent advances in scalable generative architectures and large-scale text-image datasets have enabled unprecedented performance in zero-shot text-to-image synthesis. DALL⋅\cdotE and GLIDE , as pioneering works, employ auto-regressive model and diffusion model as their architectures, respectively. DALL⋅\cdotE 2 utilizes a diffusion prior network to translate CLIP text embeddings to CLIP image embeddings, and an unCLIP module to synthesize images from CLIP image embeddings. In Imagen and Stable Diffusion , a large pre-trained text encoder is employed to guide the sampling process of a diffusion model in pixel space and latent space, respectively.

Zero-Shot Text-to-3D. Zero-shot text-to-3D generation techniques exploit the joint text-image modeling capability of pre-trained vision-language models such as CLIP to obviate the need for paired text-3D data. CLIP-Forge trains a normalizing flow model to convert CLIP image embeddings to VAE shape embeddings, and switches the input to CLIP text embeddings at the inference time. ISS trains a mapper to map the CLIP image embedding into the latent shape code of a pre-trained single-view reconstruction (SVR) network , which is then fine-tuned by taking the CLIP text embedding as input. DreamFields and CLIP-Mesh are pioneering works that explore zero-shot 3D content creation using only CLIP guidance. The former optimizes a randomly-initialized NeRF, while the latter optimizes a spherical template mesh as well as random texture and normal maps. PureCLIPNeRF enhances DreamFields with grid-based representation and more diverse image augmentations. Recently, DreamFusion has gained popularity in the research community due to its impressive results. Powered by a strong text-to-image model, Imagen , it can generate high-fidelity 3D objects using the score distillation loss.

Method

Our objective is to generate 3D content that aligns with the given input text prompt yy. As illustrated in Fig. 2, our framework for text-guided 3D synthesis comprises two stages. In the first stage (Sec. 3.2), we obtain an explicit 3D shape prior SS using a text-guided 3D shape generation process. The text-guided shape generation process involves a text-to-image phase that employs a fine-tuned Stable Diffusion model GIG_{I} (Sec. 3.3), and an image-to-shape phase that employs a shape embedding generation network GMG_{M} and a high-quality 3D shape generator GSG_{S}. In the second stage (Sec. 3.1), we utilize the 3D prior SS to initialize a neural radiance field , and optimize it with CLIP guidance to generate the 3D content. Our framework only requires a collection of textureless 3D shapes without any text labels to train the 3D generator GSG_{S}, and the fine-tuning process of GIG_{I} converges rapidly.

Background: 3D Optimization with CLIP Guidance. CLIP is a powerful vision-language model that comprises a text encoder ETE_{T}, and an image encoder EIE_{I}. By maximizing the cosine similarity between the text embedding and the image embedding encoded by ETE_{T} and EIE_{I} respectively on a large-scale paired text-image dataset, CLIP aligns the text and image modalities in a shared latent embedding space.

Prior research leverages the capability of CLIP to generate 3D contents from text. Starting from a randomly-initialized 3D representation parameterized by θ\theta, they render images from multiple viewpoints and optimize θ\theta by minimizing the CLIP similarity loss between the rendered image R(vi;θ)\mathcal{R}(\boldsymbol{v}_{i};\theta) and the text prompt yy:

where R\mathcal{R} denotes the rendering process, and vi\boldsymbol{v}_{i} denotes the rendering viewpoint at the ii-th optimization step. Specifically, DreamFields and PureCLIPNeRF employ neural radiance fields (NeRF) as the 3D representation θ\theta, while CLIP-Mesh uses a spherical template mesh with associated texture and normal maps.

Observations and Motivations. Though these CLIP-guided optimization methods can generate impressive results, we observe that they often fall short in producing precise and detailed 3D structures that accurately match the text description. As depicted in Fig. 1, we employ these methods to create 3D content featuring common objects, but the outcomes exhibited distortion artifacts and appeared unusual, adversely affecting their visual quality and hindering their use in real-world applications.

We attribute the failure of previous works to generate accurate and realistic objects to two main factors: (i) The optimization process begins with a randomly-initialized 3D representation lacking any explicit 3D shape prior, making it very challenging for the models to conjure up the scene from scratch. (ii) The CLIP loss in Eq. 1 prioritizes global consistency between the rendered image and the text prompt, rather than offering robust and precise guidance on the synthesized 3D structure. As a result, the optimization output is significantly unconstrained.

Optimization with 3D Shape Prior as Initialization. To address the aforementioned issue and generate more faithful 3D content, we propose to use a text-to-shape generation process to create a high-quality 3D shape SS from the input text prompt yy. Subsequently, we use it as an explicit “3D shape prior” to initialize the CLIP-guided 3D optimization process. As illustrated in Fig. 2, for the text prompt “a park bench overgrown with vines”, we first synthesize “a park bench” without textures in the text-to-shape stage. We then use it as the initialization of a neural radiance field and optimize it with the full prompt, following previous works.

Here, sigmoid⁡(x)=1/(1+e−x)\operatorname{sigmoid}(x)=1/(1+e^{-x}) and softplus⁡−1(x)=log⁡(ex−1)\operatorname{softplus}^{-1}(x)=\log(e^{x}-1). Eq. 2a converts SDF values to density for volume rendering, where β>0\beta>0 is a hyper-parameter controlling the sharpness of the shape boundary (smaller β\beta leads to sharper shape boundary, β=0.05\beta=0.05 in our experiments). Eq. 2b transforms the density into pre-activated density. To ensure that the distribution of the accumulated transmittance is the same as DVGO, we clamp the minimum value of the density outside the shape prior as .

With the density grid Vdensity\boldsymbol{V}_{\text{density}} initialized by the 3D shape SS and the color MLP frgbf_{\text{rgb}} initialized randomly, we render image R(Vdensity,frgb;vi)\mathcal{R}(\boldsymbol{V}_{\text{density}},f_{\text{rgb}};\boldsymbol{v}_{i}) from viewpoint vi\boldsymbol{v}_{i} and optimize θ=(Vdensity,frgb)\theta=(\boldsymbol{V}_{\text{density}},f_{\text{rgb}}) with the CLIP loss in Eq. 1. Following DreamFields and PureCLIPNeRF , we perform background augmentations for the rendered images and leverage the transmittance loss introduced by to reduce noise and spurious density. Besides, since CLIP loss cannot provide accurate geometrical supervision, the 3D shape prior may be gradually disturbed and “forgotten”, thus we also adopt a shape-prior-preserving loss to preserve the global structure of the 3D shape prior:

where \mathds1(⋅)\mathds{1}(\cdot) is the indicator function, and alpha⁡(⋅)\operatorname{alpha}(\cdot) transforms the density into the opacity representing the probability of termination at each position in volume rendering.

By initializing NeRF with an explicit 3D shape prior, we give extra knowledge on how the 3D content should look like and prevent the model from imagining from scratch and generating “adversarial contents” that have high CLIP scores but low visual quality. Based on the initialization, the CLIP-guided optimization further provides flexibility and is able to synthesize more diverse structures and textures.

2 Stable-Diffusion-Assisted Text-to-Shape Generation as 3D Shape Prior

To obtain the 3D shape prior, a text-guided shape generation scheme is required, which is a challenging task due to the lack of paired text-shape datasets. Previous approaches typically first train an image-to-shape model using rendered images, and then bridge the text and image modalities using the CLIP embedding space.

CLIP-Forge trains a normalizing flow network to map CLIP image embeddings of shape renderings to latent embeddings of a volumetric shape auto-encoder, and at test time, it switches to CLIP text embeddings as input. However, the shape auto-encoder has difficulty in generating high-quality and diverse 3D shapes, and directly feeding CLIP text embeddings to the flow network trained on CLIP image embeddings suffers from the gap between the CLIP text and image embedding spaces. ISS trains a mapper network to map CLIP image embeddings of shape renderings to the latent space of a pre-trained single-view reconstruction (SVR) model, and fine-tunes the mapper at test time by maximizing the CLIP similarity between the text prompt and the images rendered from synthesized shapes. While the test-time fine-tuning alleviates the gap between the CLIP text and image embeddings, it is cumbersome to fine-tune the mapper for each text prompt.

In contrast to the aforementioned methods that connect the text and image modalities in the CLIP embedding space, we use a powerful text-to-image diffusion model to directly bridge the two modalities. Specifically, we first synthesize an image from the input text and then feed it into an image-to-shape module to generate a high-quality 3D shape. This pipeline is more concise and naturally eliminates the gap between CLIP text and image embeddings. However, it introduces a new domain gap between the images generated by the text-to-image diffusion model and the shape renderings used to train the image-to-shape module. We will introduce a novel technique to alleviate this gap in Sec. 3.3.

Text-to-Image Diffusion Model. Diffusion models are generative models trained to reverse a diffusion process. The diffusion process begins with a sample from the data distribution, x0∼q(x0)\boldsymbol{x}_{0}\sim q\left(\boldsymbol{x}_{0}\right), which is gradually corrupted by Gaussian noise over TT timesteps: xt=αtxt−1+1−αtϵt−1,t=1,2,…,T\boldsymbol{x}_{t}=\sqrt{\alpha_{t}}\boldsymbol{x}_{t-1}+\sqrt{1-\alpha_{t}}\boldsymbol{\epsilon}_{t-1},t=1,2,\ldots,T, where αt\alpha_{t} defines the noise level and ϵt−1\boldsymbol{\epsilon}_{t-1} denotes the noise added at timestep t−1t-1. To reverse this process, a denoising network ϵθ\epsilon_{\theta} is trained to estimate the added noise at each timestep. During inference, samples can be generated by iteratively denoising pure Gaussian noise. Text-to-image diffusion models further condition the denoising process on texts. Given a text prompt yy and a text encoder cθc_{\theta}, the training objective is:

In this work, we use Stable Diffusionhttps://github.com/CompVis/stable-diffusion , an open-source text-to-image diffusion model, which employs a CLIP ViT-L/14 text encoder as cθc_{\theta} and is trained on the large-scale LAION-5B dataset . Stable Diffusion is known for its ability to generate diverse and imaginative images in various styles from heterogeneous text prompts.

High-quality 3D generator. To provide a more precise initialization for optimization, high-quality 3D shapes are highly desirable. In this work, we utilize the SDF-StyleGAN , a state-of-the-art 3D generative model, to generate high-quality 3D priors. SDF-StyleGAN is a StyleGAN2-like architecture that maps a random noise z∼N(0,I)\boldsymbol{z}\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{I}) to a latent shape embedding eS∈W\boldsymbol{e}_{S}\in\mathcal{W} and synthesizes a 3D feature volume FV\boldsymbol{F}_{V}, which is an implicit shape representation. We can query the SDF value at arbitrary position x\boldsymbol{x} by feeding the interpolated feature from FV\boldsymbol{F}_{V} at x\boldsymbol{x} into a jointly trained MLP. We improve upon the original SDF-StyleGAN, which trains one network for each shape category, by training a single 3D shape generator GSG_{S} on the 13 categories of ShapeNet . This modification provides greater flexibility in the text-to-shape process.

Shape Embedding Mapping Network. To bridge the image and shape modalities, we further train a shape embedding mapping network GMG_{M}. Firstly, we utilize GSG_{S} to generate a large set of 3D shapes {Si}i=1N\{S^{i}\}_{i=1}^{N} and shape embeddings {eSi}i=1N\{\boldsymbol{e}_{S}^{i}\}_{i=1}^{N}. Then, we render SiS^{i} from KK viewpoints to obtain shape renderings {Irj}j=1NK\{\boldsymbol{I}_{r}^{j}\}_{j=1}^{NK} and the corresponding image embeddings {eIj}j=1NK\{\boldsymbol{e}_{I}^{j}\}_{j=1}^{NK} with the CLIP image encoder EIE_{I}, forming a paired image-shape embedding dataset {(eIj,eSj)}j=1NK\{(\boldsymbol{e}_{I}^{j},\boldsymbol{e}_{S}^{j})\}_{j=1}^{NK}. Finally, we use this dataset to train a conditional diffusion model GMG_{M} which can synthesize shape embeddings from image embeddings of shape renderings.

To prepare the dataset for training GMG_{M}, we generate N=64000N=64000 shapes using GSG_{S} and render K=24K=24 views for each shape. The ranges of azimuth and elevation angles of the rendered views are [−90∘,90∘]\left[-90^{\circ},90^{\circ}\right] and [20∘,30∘]\left[20^{\circ},30^{\circ}\right], respectively. We employ an SDF renderer since we represent the synthesized shapes with SDF grids.

3 Fine-tuning Stable Diffusion for Rendering-Style Image Generation

As stated in Sec. 3.2, we use Stable Diffusion to directly bridge the text and image modalities for text-to-shape generation. Nonetheless, the image-to-shape module is trained on shape renderings, which exhibit a significant style domain gap from the images produced by Stable Diffusion. Previous research has attempted to combine a text-to-image model with an image-to-shape model for text-to-shape generation. However, this approach is plagued by the aforementioned style domain gap, leading to flawed geometric structures and diminished performance.

Inspired by recent work on controllable text-to-image generation such as textual inversion and DreamBooth , we propose a method for addressing the domain gap problem by fine-tuning Stable Diffusion into a stylized generator. Our core idea is to enable Stable Diffusion to replicate the style of shape renderings used to train the image-to-shape module outlined in Sec. 3.2. This allows us to seamlessly input the generated stylized images into the image-to-shape module without being affected by the domain gap.

Fine-tuning Process. The fine-tuning process is illustrated in Fig. 3. To fine-tune Stable Diffusion, we need a dataset that consists of shape renderings and related stylized text prompts. For each shape SS in the ShapeNet dataset, we generate a set of shape renderings {ISj}j=1NS\{\boldsymbol{I}_{S}^{j}\}_{j=1}^{N_{S}}. Subsequently, each rendering ISj\boldsymbol{I}_{S}^{j} is linked with a stylized text prompt ySjy_{S}^{j} in the format of “a CLS in the style of ∗*”, where CLS denotes the shape category name and ∗* represents a placeholder token that requires optimization for its text embedding. For instance, if the image is rendered from a chair shape, then the associated text prompt will be “a chair in the style of ∗*”. The paired dataset DS=(ISj,ySj)j=1NSD_{S}={(\boldsymbol{I}_{S}^{j},y_{S}^{j})}_{j=1}^{N_{S}} is then utilized to fine-tune Stable Diffusion by minimizing Ldiffusion\mathcal{L}_{\text{diffusion}} presented in Eq. 4.

During fine-tuning, we freeze the CLIP text encoder of Stable Diffusion, and optimize two objectives: (i) the text embedding of the placeholder token ∗*, denoted as v∗\boldsymbol{v}_{*}, and (ii) the parameters θ\theta of the diffusion model ϵθ\epsilon_{\theta}. Optimizing the text embedding v∗\boldsymbol{v}_{*} aims to learn a virtual word that captures the style of the rendered images best, even though it is not present in the vocabulary of the text encoder. Fine-tuning the parameters θ\theta of the diffusion model further enhances the ability to capture the style precisely since it is hard to control the synthesis of Stable Diffusion solely on the language level. Our experiments demonstrate stable convergence of the fine-tuning process in approximately 2000 optimization steps, requiring only 40 minutes on a single Tesla A100 GPU. We show some synthesized results using the fine-tuned model in Fig. 3.

Dataset Scale and Background Augmentation. We have identified two essential techniques empirically that enable the fine-tuned model to synthesize stylized images in a stable manner. Firstly, unlike textual inversion or DreamBooth which utilize only 3−53{-}5 images, fine-tuning with a larger set of shape renderings containing thousands of images helps the model capture the style more precisely. Secondly, fine-tuning Stable Diffusion using shape renderings with a pure-white background results in a chaotic and uncontrollable background during inference. However, augmenting the shape renderings with random solid-color backgrounds allows the fine-tuned model to synthesize images with solid-color backgrounds stably, making it easy to remove the background if necessary. Further details can be found in the supplementary material.

Experiments

In this section, we evaluate the efficacy of our proposed text-to-3D synthesis framework. Initially, we compare our results with state-of-the-art techniques (Sec. 4.1). Subsequently, we demonstrate the effectiveness of our Stable-Diffusion-assisted approach for text-to-shape generation (Sec. 4.2). Furthermore, we conduct ablation studies to assess the effectiveness of critical components of our framework (Sec. 4.3).

Dataset. Our framework requires only a set of untextured 3D shapes for training the 3D shape generator GSG_{S}. Specifically, we employ 13 categories from ShapeNet and utilize the data preprocessing procedure of Zheng et al. to generate 1283128^{3} SDF grids from the original meshes. During the fine-tuning of Stable Diffusion, we employ the SDF renderer to produce a shape rendering dataset.

Implementation details. The 3D generator GSG_{S} utilizes the SDF-StyleGAN architecture. The diffusion-model-based shape embedding mapping network GMG_{M} is based on an open-source DALL⋅\cdotE 2 implementationhttps://github.com/lucidrains/DALLE2-pytorch. We train GMG_{M} by extracting image embeddings from shape renderings using the CLIP ViT-B/32 image encoder. The stylized text-to-image generator GIG_{I} is fine-tuned from Stable Diffusion v1.4. In the optimization stage, we set the learning rates for the density grid Vdensity\boldsymbol{V}_{\text{density}} and color MLP frgbf_{\text{rgb}} to 5×10−15\times 10^{-1} and 5×10−35\times 10^{-3} respectively, and we adopt the CLIP ViT-B/16 encoder as the guidance model. For each text prompt, we optimize for 5000 steps, while previous NeRF-based text-to-3D methods typically require 10000 steps or more.

Evaluation Metrics. Regarding the primary results of our framework, i.e., text-guided 3D content synthesis, we report the CLIP retrieval precision on a manually created dataset of diverse text prompts and objects. For specifics regarding the dataset, please refer to the supplementary material. This metric quantifies the percentage of generated images that the CLIP encoder associates with the correct text prompt used for generation. We utilize Fréchet Inception Distance (FID) to evaluate the shape generation quality for the initial text-to-shape generation stage.

We compare our method with three state-of-the-art baseline methods on the task of text-guided 3D synthesis, i.e., DreamFields , CLIP-Mesh , and PureCLIPNeRF . We conduct tests using the default settings and official implementations for all baseline methods. In particular, we utilize the medium-quality configuration of DreamFields and the implicit architecture variant of PureCLIPNeRF owing to its superior performance.

Time cost. Thanks to the shape prior initialization, our CLIP-guided optimization process exhibits significantly improved efficiency compared to previous NeRF-based methods. Dream3D optimizes for only 5000 steps within 25 minutes, while DreamFields and PureCLIPNeRF require more than 10000 steps, taking over an hour (measured on 1 A100 GPU). Training the 3D shape generator GSG_{S} takes 7 days on 4 A100 GPUs and training the shape embedding generation network GMG_{M} takes 1 day on 1 A100 GPU. It is worth noting that these models are only trained once and their inference time can be neglected compared to the optimization cost.

Quantitative Results. We report the CLIP retrieval precision metrics in Tab. 1. It is noteworthy that both the baseline methods and our approach utilize the CLIP ViT-B/16 encoder for optimization, and both the CLIP ViT-B/16 and CLIP ViT-B/32 encoders are employed as retrieval models. As Table Tab. 1 shows, our method achieves the highest CLIP R-Precision with both retrieval models. Moreover, our framework exhibits a significantly smaller performance gap between the two retrieval models compared to the baseline methods. By leveraging the 3D shape prior, our method initiates the optimization from a superior starting point, thereby mitigating the adversarial generation problem that prioritizes obtaining high CLIP scores while neglecting the visual quality. Consequently, our method demonstrates more robust performance across different CLIP models.

Qualitative Results. The qualitative comparison is presented in Fig. 4, indicating that the baseline methods encounter challenges in generating precise and realistic 3D objects, resulting in distorted and unrealistic visuals. DreamFields’ results are frequently blurry and diffuse, while PureCLIPNeRF tends to synthesize symmetric objects. CLIP-Mesh experiences difficulties in generating intricate visual effects due to its explicit mesh representation. In contrast, our method effectively generates higher-quality 3D structures by incorporating explicit 3D shape priors.

2 Text-to-Shape Generation

The research on zero-shot text-to-shape generation is limited, and we compare our approach with CLIP-Forge and measure the quality of shape generation using the Fréchet Inception Distance (FID). Specifically, we synthesize 3 shapes for each prompt from a dataset of 233 text prompts provided by CLIP-Forge. Then, we render 5 images for each synthesized shape and compare them to a set of ground truth ShapeNet renderings to compute the FID. The ground truth images are obtained by randomly choosing 200 shapes from the test set of each ShapeNet category and rendering 5 views for each shape.

As shown in Tab. 2, our approach achieves a lower FID than CLIP-Forge. CLIP-Forge employs a volumetric shape auto-encoder to generate 3D shapes. However, the qualitative results in Fig. 5 indicate poor shape generation capability, making it difficult to generate plausible 3D shapes. A high-quality 3D shape prior is also advantageous for the optimization process as an excellent initialization.

3 Ablation Studies

Effectiveness of Fine-tuning Stable Diffusion. Our approach employs a fine-tuned Stable Diffusion to establish a connection between the text and image modalities. To evaluate its efficacy, we employ the original Stable Diffusion model to produce images from the text prompts of CLIP-Forge . Subsequently, we employ these images to generate 3D shapes using the image-to-shape module. The results in the \nth3 row of Tab. 2 indicate a decline in FID performance, suggesting that this approach would negatively impact the shape generation process. Furthermore, we directly test using text embedding to generate shape embeddings with GMG_{M}, which also leads to a decline in performance as seen in the \nth2 row of Tab. 2.

Effectiveness of 3D Shape Prior. To validate the efficacy of the 3D shape prior, we eliminate the first stage of our framework and optimize from scratch using the same text prompts as presented in Sec. 4.1. Subsequently, we evaluate the CLIP retrieval precision. The outcomes presented in Tab. 1 indicate that optimizing without 3D shape prior results in a considerable decline in performance, thereby demonstrating its effectiveness.

Effectiveness of Lprior\mathcal{L}_{\text{prior}}. The 3D prior preserving loss Lprior\mathcal{L}_{\text{prior}} shown in Eq. 3 aims to reinforce the 3D prior during the optimization process in case that the prior is gradually disturbed and discarded. To demonstrate its effectiveness, we synthesized “a park bench” as a prior for the prompt “A park bench overgrown with vines”, and then optimize with and without Lprior\mathcal{L}_{\text{prior}} and compared the results, which are visualized in Fig. 6. The results indicate that using Lprior\mathcal{L}_{\text{prior}} during optimization helps maintain the structure of the 3D prior shape, while discarding it causes distortion and discontinuity artifacts, thereby disturbing the initial shape.

Limitations and Future Work

Our framework relies on a fine-tuned Stable Diffusion to generate rendering-style images. Despite its strong generation capability, Stable Diffusion may produce shape images that fall outside the distribution of the training data of the image-to-shape module. This is due to the fact that Stable Diffusion is trained on an internet-scale text-image dataset, whereas the 3D shape generator is trained on ShapeNet. Furthermore, the quality of text-to-shape synthesis in our framework is heavily reliant on the generation capability of the 3D generator. Our future work will explore incorporating stronger 3D priors into our framework to enable it to work with a wider range of object categories.

Additionally, our framework is indeed orthogonal to score distillation-based text-to-3D methods , as we can also utilize the score distillation sampling objective for optimization. We believe that incorporating 3D shape priors can enhance the quality and diversity of the generation results, as DreamFusion acknowledged.

Conclusion

This paper introduces Dream3D, a text-to-3D synthesis framework that can generate diverse and imaginative 3D content from text prompts. Our approach incorporates explicit 3D shape priors into the CLIP-guided optimization process to generate more plausible 3D structures. To address the text-to-shape generation, we propose a straightforward yet effective method that utilizes a fine-tuned text-to-image diffusion model to bridge the text and image modalities. Our method is shown to generate 3D content with superior visual quality and shape accuracy compared to previous work, as demonstrated by extensive experiments.

Acknowledgements. The work was supported by National Key R&D Program of China (2018AAA0100704), NSFC #61932020, #62172279, Science and Technology Commission of Shanghai Municipality (Grant No. 20ZR1436000), Program of Shanghai Academic Research Leader, and “Shuguang Program” supported by Shanghai Education Development Foundation and Shanghai Municipal Education Commission.

References

We adopt the architecture of SDF-StyleGAN as our 3D generator. As Fig. 7 shows, it maps a random noise z∼N(0,I)\boldsymbol{z}\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{I}) to a latent shape embedding eS∈W\boldsymbol{e}_{S}\in\mathcal{W} and synthesizes a 3D feature volume FV\boldsymbol{F}_{V}, which is an implicit representation of the generated shape. We can query the SDF value at arbitrary position x\boldsymbol{x} by feeding the interpolated feature from FV\boldsymbol{F}_{V} at x\boldsymbol{x} into a jointly trained MLP network. During training, a global discriminator and a local discriminator are used simultaneously to supervise the generated SDF grids at the coarse and fine level respectively. Different from the original SDF-StyleGAN that trains one network for one shape category, we train one 3D shape generator GSG_{S} on 13 categories of the ShapeNet dataset to enlarge the shape generation capability.

The shape embedding mapping network GMG_{M} is a diffusion-model-based generative network that can generate shape embeddings eS\boldsymbol{e}_{S} from the CLIP image embeddings eI\boldsymbol{e}_{I} of shape renderings. The network architecture and training strategy of GMG_{M} are based on an open-source DALL-E-2 implementationhttps://github.com/lucidrains/DALLE2-pytorch. Specifically, GMG_{M} is equivalent to the diffusion prior network in DALL-E-2 which generates CLIP image embeddings from CLIP text embeddings. Here we replace the input with CLIP image embeddings of shape renderings and the output with shape embeddings. We use the DiffusionPrior class in the codebase to implement GMG_{M} and the train_diffusion_prior.py script to train GMG_{M}. The model and training hyperparameters are listed in Tab. 3.

C Details of Fine-tuning Stable Diffusion

In our framework, we connect the text and image modalities by fine-tuning the Stable Diffusion into a stylized generator with a set of shape renderings {ISj}j=1NS\{\boldsymbol{I}_{S}^{j}\}_{j=1}^{N_{S}} and give it the ability to synthesize images in the “rendering” style. In experiments, we find it crucial to utilize a large set of shape renderings for fine-tuning and to augment the backgrounds of the shape renderings with random colors.

We tried fine-tuning Stable Diffusion using shape renderings with three different types of backgrounds: 1) solid white background, 2) solid green background and 3) random-color background. We visualize the shape renderings used for fine-tuning and the images synthesized by the fine-tuned Stable Diffusion in Fig. 8. As Fig. 8(a) and Fig. 8(b) show, although shape renderings with solid-white or solid-green backgrounds can make the fine-tuned Stable Diffusion capture the “rendering” style of the object successfully, the backgrounds in the synthesized images are out of control, i.e., the fine-tuned Stable Diffusion fails to synthesize images with solid-color backgrounds. This will increase the difficulty of separating the foreground objects from the backgrounds and affect the stability of the subsequent image-to-shape generation since the shape embedding mapping network GMG_{M} is trained on shape renderings with solid-color backgrounds. In comparison, augmenting the backgrounds of the shape renderings with random colors leads to a stable stylized generator that can synthesize solid-color-background images consistently, as Fig. 8(c) shows.

During fine-tuning, we indeed expect the Stable Diffusion model to capture two types of styles: 1) the “rendering” style of the foreground object and 2) the “solid-color” style of the background. Similar to the observation that the foreground “rendering” style requires a large set of rendered images to learn, we consider that a single-color background is too few to be recognized as a “solid-color background style” by the Stable Diffusion model, while showing a lot of different solid-color examples to the model can make it notice the solid-color background style and capture it during fine-tuning.

To better demonstrate the importance of the random-color background augmentation, we also evaluate the Fréchet Inception Distance (FID) between the shape renderings used for fine-tuning and the images synthesized by the fine-tuned Stable Diffusion in Tab. 4. For each type of background, we render 10001000 images with that background for each ShapeNet category, forming a shape rendering dataset containing 13000 images in total (denoting as DSD_{S}). Then we leverage DSD_{S} to fine-tune the Stable Diffusion model for 5000 steps, and utilize the fine-tuned Stable Diffusion to synthesize 100 images for each shape category using the text prompt ”a CLS in the style of *”, leading to a set of 1300 generated images (denoting as DgenD_{gen}). Finally, we compute the FID between DSD_{S} and DgenD_{gen}. As Tab. 4 shows, augmenting the backgrounds of shape renderings with random colors significantly boosts the FID, which demonstrates its effectiveness.

D Details of 3D Optimization with 3D Shape Prior

To render the color of a pixel C^(r)\hat{C}(r), we cast the ray rr from the camera center through the pixel, and sample KK points between the pre-defined near and far planes. We then query the densities and colors of the KK ordered sampled points {(σi,ci)}i=1K\left\{\left(\sigma_{i},c_{i}\right)\right\}_{i=1}^{K} using Eq. 5. Finally, we accumulate the KK queried results into a single color with the volume rendering process:

where αi\alpha_{i} denotes the opacity representing the probability of termination at point ii, TiT_{i} denotes the accumulated transmittance from the near plane to point ii, δi\delta_{i} denotes the distance to the adjacent sampled point, and cbgc_{bg} demotes a pre-defined background color.

Following DVGO, all values in Vdensity\boldsymbol{V}_{density} are initialized as and the bias term in Eq. 5b is set to

where αinit\alpha_{init} is a hyperparameter and is set to 10−610^{-6} in practice. With such an initialization, the accumulated transmittance TiT_{i} is decayed by 1−αinit≈11-\alpha_{init}\approx 1 for a ray that traces forward a distance of a voxel size ss, making the scene “transparent” at the beginning of optimization.

D.2 Shape Prior Initialization and Optimization

With such an initialization, the density values on the shape surface will be close to 12β\frac{1}{2\beta} (log⁡(exp⁡(1β⋅sigmoid⁡(0))−1)≈12β\log(\exp(\frac{1}{\beta}\cdot\operatorname{sigmoid}(0))-1)\approx\frac{1}{2\beta}). The area inside the shape surface will have larger density values (>12β>\frac{1}{2\beta}), and the density values outside the shape will decrease with the distance from the shape surface. We set the minimum density value outside the shape to so the area far from the shape surface has the same initialization as the original DVGO.

As Fig. 9 shows, at the beginning of the 3D optimization process (step=0), the 3D shape prior is visible due to the larger density values around the shape surface. As a result, the area around the shape surface will dominate the volume rendering, and the density/color values in this area will be updated faster than the area far from the surface. Based on the initialization, Then subsequent CLIP-guided optimization process further provides more flexibility and is able to synthesize more diverse structures and textures.

E Additional Results on Text-to-Shape Generation

We show additional qualitative text-guided 3D shape generation results in Fig. 10. Compared to CLIP-Forge , our method produces more plausible 3D shapes thanks to the high-quality 3D generator, while the shapes generated by suffer from rough surfaces and discontinuities.

Besides, we also provide more quantitative comparisons with CLIP-Forge on text-to-shape generation. We generate 3 shapes for each text prompt in the text prompt set provided by CLIP-Forge and measure three metrics: 1) Fréchet Inception Distance (FID) between 5 rendered images for each shape with different camera poses and a set of images rendered from the ground truth shapes in the ShapeNet dataset with the same camera poses. 2) Fréchet Point Distance (FPD) , for each generated shape and each ground truth shape in the ShapeNet test set, we extract the mesh at 64364^{3} resolution and sample 2048 points from the mesh surface, then pass the points to a DGCNN backbone network pre-trained on the point cloud classification task and use the feature of the last layer to compute this metric. 3) Maximum Measure Distance (MMD), for each generated shape represented by a 32332^{3} occupancy grid, we match a shape in the ShapeNet test set based on the highest IOU, and then average the IOU across all the text queries. As Tab. 5 shows, our text-to-shape generation method outperforms CLIP-Forge on all three metrics.

F Additional Results on Text-to-3D Synthesis

In this section, we show additional qualitative comparison results on text-to-3D synthesis with baseline methods in Fig. 11 and more diversified generation results of our method in Fig. 12. It can be seen that our method can synthesize plausible 3D structures with the help of 3D shape priors. To better visualize the 3D structures generated by different methods, we also show video examples in the attached MP4 file.

G Integration with SVR Models

Single-view reconstruction (SVR) models can reconstruct a 3D shape from a single input image. We then ask, can we use an SVR model directly as the image-to-shape module in our framework? To answer this question, we conduct the same fine-tuning process to fine-tune a Stable Diffusion model with the ShapeNet renderings provided by Choy et al. which are commonly used by many SVR methods. We find that although the shape renderings in Choy et al. have more complex textures, the fine-tuned model can still capture the style successfully and synthesize novel images imitating the style. With such a fine-tuned Stable Diffusion, we can solve the text-to-shape generation in a precise way: synthesize an image using the fine-tuned Stable Diffusion with text prompt in the format of ”a CLS in the style of *”, and then directly feed the synthesized image into the SVR model. We show some text-guided shape generation results using two SVR methods, i.e., occupancy networks and DVR , in Fig. 13 and Fig. 14, respectively. Both methods are trained with the shape renderings provided by Choy et al. . The occupancy networks only predict shape, while DVR can predict both shape and color. As Fig. 13 and Fig. 14 show, we achieve text-to-shape generation successfully with the synthesized images, which proves the strong generation ability of Stable Diffusion and the effectiveness of our fine-tuning pipeline.

A recent work named ISS also utilizes an SVR model to perform text-to-shape generation. However, the pipeline of ISS is much more complicated. It trains a mapper network to map CLIP features to the latent space of the SVR model, which requires a two-stage fine-tuning to align the text and shape feature spaces. At inference time, ISS needs to fine-tune the mapper network for each text prompt, which is redundant in our pipeline. With the help of the fine-tuned Stable Diffusion, we can directly generate an image from the text prompt and feed the image into the SVR model to synthesize a 3D shape. Besides, thanks to the strong generation ability of Stable Diffusion, we can enjoy a much larger generation diversity and synthesize as many 3D shapes as we want for each text prompt.

G.2 Text-to-3D Synthesis using SVR models

Despite the success in text-guided shape generation with SVR models, we find that current SVR models are very sensitive to the input images. Although we can successfully capture the style of the shape renderings using the fine-tuned Stable Diffusion, some minor flaws in the synthesized images such as offsets of the objects from the image center and unrealistic artifacts (e.g., a chair lacks a leg) are inevitable. These minor flaws may lead to failed shape reconstructions, whose quality affects 3D shape priors. This sensitiveness makes the 3D prior generation in the first stage of our framework unstable. Therefore, we choose to use a 3D generator associated with a shape embedding mapping network to generate 3D shapes in the latent shape embedding space, instead of directly using an SVR model in our framework.

We visualize six text-to-3D synthesis results using 3D shape priors produced by the occupancy networks in Fig. 15. The successful results in the first two rows show the probability of integrating as SVR model into our framework. In the last row, we show two failure cases in which the SVR model fails to reconstruct plausible 3D shape priors to illustrate the drawbacks of using SVR models. We can observe that the discontinuity in the “bedside lamp” shape leads to discontinuity in the final optimization result, while the failed truck shape results in total chaos.