Control3D: Towards Controllable Text-to-3D Generation
Yang Chen, Yingwei Pan, Yehao Li, Ting Yao, Tao Mei
Introduction
3D content creation has significantly impacted the multimedia field by enabling the construction of immersive and engaging digital worlds. The applications from video games and animated films to virtual reality present vast massive opportunities for 3D professionals to deliver compelling experiences that captivate audiences. The conventional process of 3D content creation is commonly time-consuming, necessitating the expertise of professional designers with significant experience in 3D software tools. As such, there is a strong urge to automate this creative process, not only for enhancing the efficiency of 3D designers with domain expertise but also for democratizing 3D content creation for novices.
Recent advancements in 3D-specific generative adversarial networks (GANs) (Nguyen-Phuoc et al., 2019; Chan et al., 2022; Liao et al., 2020; Niemeyer and Geiger, 2021; Chan et al., 2021; Gu et al., 2022) have reduced the technical barriers to creating 3D content. However, constructing such GAN models requires cost-expensive large-scale 3D data collection and meticulous pre-processing. Therefore, these models are typically confined to a pre-defined single object category, which severely restricts the diversity of synthetic 3D content and their practical applicability. In contrast, more recent text-to-3D generation works (Poole et al., 2023; Lin et al., 2023; Wang et al., 2023; Metzer et al., 2023) demonstrated their remarkable ability to create 3D scenes solely based on human-written text prompts, yielding various impressive 3D assets. Such automatic 3D content creation from input text prompts alone can greatly emancipate ordinary users from the need to acquire all the skills required for creating 3D assets. Nevertheless, simplifying the 3D generation interface as a text-only format may impede users’ ability to fully articulate their desired specifications (e.g., visual prompts like sketch). Accordingly, it is crucial to explore more robust interfaces that offer comprehensive control signals for 3D content creation.
In this work, we propose Control3D, the first attempt that enhances users’ controllability by upgrading text-to-3D generation with additional hand-drawn sketch conditions. The incorporation of sketching aligns with human’s innate ability in drawing and painting, offering a more intuitive and natural way for users to interactively control 3D content creation. To achieve this, we draw inspiration from recent works (Poole et al., 2023; Wang et al., 2023) which leverage large-scale pre-trained text-to-image diffusion models to optimize a Neural Radiance Field (NeRF) by applying score distillation sampling (SDS). SDS estimates the optimization direction of NeRF such that the distributions of rendered images derived from the 3D model are pushed to a higher density probability region determined by the input text prompt. Consequently, the generated 3D scenes are semantically aligned with the input text prompts. Herein, we take one step further by extending the typical text-driven SDS with more conditions of sketch. In particular, we propose to integrate the optimization of NeRF with an image-conditioned text-to-image diffusion model (ControlNet (Zhang and Agrawala, 2023)), which triggers the control of diffusion models with additional sketch conditions. Although this way simply guides text-to-3D generation with the control signals of sketches, we observe that such implicit control process is insufficient to produce high-quality 3D scene that precisely maintains the same geometric structure of given sketch.
To alleviate this issue, we further design a novel sketch consistency loss that explicitly encourages the geometric consistency between synthetic 3D scene and given sketch. Technically, in each training step, we utilize a pre-trained differentiable photo-to-sketch model (Li et al., 2019) to estimate the sketch of the rendered image. We then employ the sketch consistency loss to match the embeddings between the estimated sketch and the given input sketch. This design encourages NeRF model to generate outputs that adhere to the sketch-specified geometry from arbitrary poses.
We conduct thorough experiments to verify the effectiveness of our proposed method. Experimental results show that our proposed Control3D is capable of generating realistic 3D scenes with remarkable likeness to the given sketch while also respecting the contexts present in the input text prompt. In sum, we have made the following contributions:
We propose Control3D, a new framework to create realistic 3D scene conditioned on a text prompt and a visual prompt (sketch image). To the best of our knowledge, this is the first attempt to control text-to-3D generation with a human-drawn sketch.
We additionally introduce a novel sketch consistency loss that explicitly enforces the synthetic 3D scene to precisely preserve the same geometric structure as in the given sketch.
We perform extensive experiments to demonstrate that our controllable text-to-3D generation results not only have plausible appearances and shapes, but also faithfully conform to the given prompt and sketch.
Related Work
Diffusion models. Diffusion models (Sohl-Dickstein et al., 2015; Ho et al., 2020; Nichol and Dhariwal, 2021; Ho and Salimans, 2022) have emerged as the new trend of generative models for generating diverse, high-quality content. Especially, they have recently been used to form state-of-the-art text-to-image (T2I) models (such as DALL-E2 (Ramesh et al., 2022) and Imagen (Saharia et al., 2022)) with the help of large-scale datasets. These models can generate high-quality images of objects and scenes that are aligned with a natural language text prompt given by the user. In order to reduce the computation resources and improve the inference speed, Latent Diffusion Model (LDM) was further proposed (Rombach et al., 2022). Motivated by the success of these models, many works attempt to control pre-trained T2I diffusion models to support additional input conditions. Textual Inversion (Gal et al., 2023) and DreamBooth (Ruiz et al., 2023) are proposed to personalize the contents in the generated images using a small set of images with the same subjects. Recent work ControlNet (Zhang and Agrawala, 2023) proposes to control large image diffusion model (i.e., Stable Diffusion) by additional condition inputs like edge maps, segmentation maps, sketches, etc. Despite the advances in controllable 2D T2I generation, using text prompts and additional condition images to describe and control 3D generation remains an open and challenging problem in the multimedia field.
Text-to-3D generation. Recently, significant advancements have been made in multimedia content generation (Ho et al., 2020; Ramesh et al., 2022; Saharia et al., 2022; Chen et al., 2019b, a; Poole et al., 2023; Jain et al., 2022; Pan et al., 2017; Zhang et al., 2023). In between, with recent notable advancements in T2I generation and NeRF based 3D reconstruction, there has been a growing interest in text-to-3D generation. While large-scale paired text-image data is available for T2I generation, paired text-3D data is currently unavailable on a similar scale. To liberate the need for training data, DreamField (Jain et al., 2022) and CLIP-Mesh (Mohammad Khalid et al., 2022) leverage cross-modal knowledge from a pre-trained image-text model (i.e., CLIP model) to optimize underlying 3D representations (NeRFs and Meshes). However, these models tend to produce less photorealistic 3D results. More recently, sparked by the success of diffusion models in 2D image generation, DreamFusion (Poole et al., 2023) and SJC (Wang et al., 2023) utilize pre-trained T2I diffusion models for text-to-3D generation and demonstrates impressive results. Following work Magic3D (Lin et al., 2023) further improves the generation quality with a coarse-to-fine strategy that leverages both low- and high-resolution diffusion priors.
Existing methods (Jain et al., 2022; Poole et al., 2023) can generate 3D assets matching the input text prompt. However, they are unable to control the 3D generation process with additional freehand interfaces. In this paper, we instead tackle a novel and challenging problem, which is to control the text-to-3D generation process with a hand-drawn sketch. Latent-NeRF (Metzer et al., 2023) is perhaps the most related work that uses a 3D mesh as an additional constraint to guide the generation process. However, the 3D mesh is too complex and difficult for ordinary users to design and produce. In contrast, our proposed control signal, presented in the form of a sketch, is more intuitive and user-friendly, making it a more accessible method for users to interact with and control the 3D generation.
Sketch-based visual synthesis. Using a sketching interface to guide computers to generate content can be traced back to Ivan Sutherland’s SketchPad (Sutherland, 1964). This tradition has continued in the area of sketch-based visual synthesis. One common approach is using GANs to learn the mapping between sketches and images. SketchyGAN (Chen and Hays, 2018) presents an edge-preserving data augmentation technique to train a GAN that can synthesize plausible images from sketches. ContextualGAN (Lu et al., 2018) proposes to learn the joint distribution of sketch and image for faithful sketch-to-image generation. Recent works (Voynov et al., 2022; Zhang and Agrawala, 2023) involve pre-trained T2I models in sketch-based visual synthesis. Given a sketch and a text prompt, these models use the sketch to control the diffusion model, producing results that align with the text prompt and follow the spatial layout of the sketch. In this work, we go one step further and make the first attempt that demonstrates sketch controlling in the realm of text-to-3D generation. We notice a related work Sketch2Mesh (Guillard et al., 2021) that focuses on 3D generation from sketches. However, our work targets creating realistic 3D scenes conditioned on a text prompt plus a visual prompt (sketch image), which is more challenging.
NeRF with Regularizations. Recently, Neural Radiance Fields (NeRF) (Mildenhall et al., 2020) has received significant attention due to its powerful representation ability for 3D scenes. Although NeRF achieves state-of-the-art performance in view synthesis, its ability to reconstruct scenes from a sparse set of input views is significantly limited. The performance drops severely when only a few input views are available. Various external regularizations have been proposed to address this problem (Deng et al., 2022; Jain et al., 2021; Kim et al., 2022; Niemeyer et al., 2022; Xu et al., 2022). Specifically, DietNeRF (Jain et al., 2021) introduces semantic consistency constraints that align input and novel views. InfoNeRF (Kim et al., 2022) proposes a ray entropy minimization regularization to implicitly regularize the density field, while DS-NeRF (Deng et al., 2022) explicitly incorporates additional depth supervision. RegNeRF (Niemeyer et al., 2022) introduces a normalizing flow depth smoothness regularization, and SinNeRF (Xu et al., 2022) proposes multiple semantic and geometry regularizations in a semi-supervised perspective. These advancements are highly significant in the development of NeRF and provide valuable insights for the field of view synthesis. However, the aforementioned regularizations often rely on ground truth views of the 3D scene, while our Control3D focuses on a more challenging setting that only has a text prompt and sketch as input. In order to pursue better controllable text-to- 3D generation, we design a novel sketch consistency loss to regularize the NeRF optimization.
Method
In this section, we elaborate our proposed Control3D, which leverages a hand-drawn sketch to guide text-to-3D generation. Our approach provides an intuitive and user-friendly control for text-to-3D generation. We start by briefly reviewing the background of Neural Radiance Fields and diffusion models. We then continue to introduce our method and how we apply the image-conditioned 2D diffusion model and novel sketch consistency loss functions to enable controllable text-to-3D generation. Figure 2 depicts an overview of our Control3D model.
where the ray originating at the camera center through the pixel along direction , and the accumulated transmittance weights the radiance by the ray travels from the image plane at to unobstructed. To approximate the integral, NeRF employs a hierarchical sampling algorithm to select points within near and far bound and along each ray. Since all processes are fully differentiable, NeRF training loss is formulated as a pixel-wise photometric reconstruction error between rendered pixel color and the ground truth color . To render an image, a collection of rays are sampled corresponding to all the pixels in that image, and the resulting color values are arranged into a 2D image.
Diffusion Models. Diffusion models (DMs) are generative models that can generate samples from a Gaussian distribution to match target data distribution by a gradual denoising process (Ho et al., 2020). In the forward diffusion process , Diffusion models gradually add Gaussian noises to a ground truth image according to a predetermined schedule :
where is a noised ample with noise level . The reverse process consists of denoising steps that progressively remove noise by modeling a neural network with parameters that predicts the noise contained in a noisy image at step . The loss function for training the diffusion model is formulated as follows:
where t uniformly sampled from and is a weighting function that depends on the timestep . Then can be reconstructed from by removing the predicted noise:
where , and , .
A T2I diffusion model builds upon the above theory and receives a text prompt as an additional condition. Given a text prompt , a text encoder first maps it into text embedding. Then the text embedding is injected into the diffusion model via attention mechanism widely adopted in Vision Transformers (Yao et al., 2023; Li et al., 2022; Yao et al., 2022). Formally, the T2I diffusion model can be denoted as .
2. Control3D
In pursuit of facilitating controllable text-to-3D generation, our method takes hand-drawn sketches as an additional condition to guide text-to-3D generation. In this section, we first describe how we integrate text-to-3D generation with a 2D conditioned diffusion model (ControlNet) in the process of text-to-3D generation, then describe our proposed sketch consistency loss for pursuing better controllable text-to-3D generation.
Text-to-3D Generation with Score distillation Sampling. A recent pioneering practice (Dreamfusion (Poole et al., 2023)) designs Score distillation sampling (SDS), which enables utilizing a pre-trained T2I diffusion model to optimize a NeRF model solely based on a text prompt . Formally, let the NeRF model parameterized by and be a differentiable volumetric renderer that can produce an image at a given camera pose , i.e., . The SDS loss provides the gradient direction to update NeRF parameters :
The SDS loss perturbs the rendered image (i.e., one view of the NeRF’s output) into a noisy sample at arbitrary timestep as described in the forward diffusion process. Then, and the input text are taken as inputs of diffusion model to predict the noise , which should be the same as the added noise . Intuitively, by doing so, SDS loss pushes the rendered images towards the higher-density regions under the text-conditioned diffusion prior, i.e., to be realistic and resemble the given input text prompt.
Sketch-controlled Text-to-3D Generation. Given a sketch image and a text prompt , our goal is to generate a realistic 3D scene that not only follows the sketch outline but also respects the contexts present in the input text prompt. To achieve this goal, we remould the standard SDS based text-to-3D pipeline by exploiting an image conditioned diffusion model (ControlNet (Zhang and Agrawala, 2023)) to trigger sketch-controlled text-to-3D. ControlNet is an end-to-end neural network architecture that controls a large-scale pre-trained image diffusion model (Stable Diffusion) to learn task-specific input conditions. Specifically, herein we use ControlNet-scribblehttps://huggingface.co/lllyasviel/sd-controlnet-scribble as our diffusion prior model, which is trained on large-scale sketch-image-text pairs and can enable sketch-guided text-to-image generation. As ControlNet-scribble has two conditions, sketch and text prompt , the noise is estimated as follows:
where is the scale of classifier-free guidance (Ho and Salimans, 2022) and is a hyper-parameter that determines the control degree of the conditioned sketch image . Note that when , the ControlNet-scribble is degraded as a Stable Diffusion model, which will generate images only from the text prompt while ignoring the sketch condition. In this way, our proposed method elegantly incorporates the input sketch image and text prompt in a unified fashion. Similar to Eq. 5, we update the NeRF model by the following gradient:
where is the parameters of the pre-trained ControlNet-scribble.
Intuitively, previous SDS-based text-to-3D generation (Eq. 5) ensures rendered views of the 3D scene lie in a higher probability density region conditioned on a single text prompt under the diffusion prior. However, our conditioned SDS (Eq. 7) encourages the rendered images also align with the input sketch, where the probability density region is further narrowed down by the sketch condition. As a result, we can obtain a 3D scene that aligns closely with the input text prompt and sketch. Following (Poole et al., 2023; Wang et al., 2023), our diffusion loss also employed with view-dependent prompting (e.g., adding “front view”, “side view”, or “back view” with respect to the camera position to the main prompt). We set for the sketch image corresponded view and for other views to avoid the learned 3D scene being overfitted to the viewpoint of the input sketch image. Nevertheless, through our experiments, we found that implicitly controlling 3D generation solely using 2D diffusion prior commonly fails to ensure the generated 3D scene precisely aligns with the geometry cues of the input sketch.
Sketch Consistency Loss. To mitigate the aforementioned issues, we propose a novel sketch consistency loss to encourage the geometry described by the input sketch to be highly preserved in the 3D generation. One intuitive way to achieve this goal is to leverage the input sketch image to directly constrain NeRF rendered images. Nevertheless, the target NeRF rendered images are photo-realistic and thus have a huge domain/style gap with the input sketch image. Thus, it is not trivial to directly encourage the similarity between rendered images and input sketch image. In contrast, we propose to utilize an off-the-shelf photo-to-sketch model (Li et al., 2019) to estimate the sketch of the rendered images. By doing so, we can compare the estimated sketch with the input sketch, thereby easily encouraging the synthetic results geometrically consistent with input sketches.
Next, a natural solution to compare the ground-truth input sketch with the estimated sketch is to use traditional mean squared error loss. However, such pixel-wise comparison might be misleading. This is because the comparison is only accurate when the estimated sketch is perfectly aligned with the original sketch image’s pose, which is often not accessible. Instead, an alternative way is to generally compare the semantic-level representation of the input sketch and estimated sketches captured from different viewpoints. To fulfill this goal, inspired by (Jain et al., 2021), we utilize a CLIP image encoder (Radford et al., 2021) to extract normalized image embeddings of the estimated sketch and input sketch. On the one hand, the CLIP image encoder is trained on hundreds of millions of web images that allow the network to understand sketch modality images (Radford et al., 2021; Sain et al., 2023). On the other hand, it can capture consistent semantic-level representation across varied viewpoints (Jain et al., 2021). Then our sketch consistency loss can be formulated by minimizing their cosine similarity:
where is a rendered image by the NeRF model from an arbitrary viewpoint. Although there exists pixel-wise misalignment between the estimated sketch and input sketch as they have different scales and viewpoints, we observe that the sketch consistency loss is robust to supervise the NeRF model to generate output that adheres to the sketch-specified geometry from arbitrary poses.
Overall Training. Finally, the overall objective to train a NeRF for our controllable text-to-3D generation is given by:
Experiments
We implement the proposed Control3D mainly based on the Score Jacobian Chaining (SJC) (Wang et al., 2023) codebase. Following SJC, we use the voxel radiance field (Chen et al., 2022) to implement the underlying NeRF. SJC uses emptiness loss and center depth loss to regularize the NeRF learning. Our method also leverages these regularizations. For more details please refer to the original SJC (Wang et al., 2023). Following (Jain et al., 2021), the image encoder used in Eq. 8 is the pre-trained CLIP ViT B/32 (Radford et al., 2021). We resize the estimated sketch and input sketch to resolution to match the input resolution of CLIP image encoder architecture. All experiments of Control3D are conducted on a single NVIDIA V100 GPU. We train the model for 10,000 iterations and the whole training process takes approximately an hour for each scene.
2. Performance Comparison and Analysis
Visualization of Controllable Text-to-3D Generation via our Control3D. Here we show the qualitative examples of our controllable text-to-3D generation with sketch guidance in Figure 3. For each scene, we show several different views. In general, our Control3D manages to produce 3D scenes using simple hand-drawn sketches plus corresponding text prompts. We clearly observe that the synthetic 3D scenes faithfully respect both the semantic context present in the input text prompt and the geometric structure specified in the input sketch. Note that the input hand-drawn sketch is not required to be too strictly accurate or tight. Even when the contour curve of the input sketches only roughly describes the shapes of target 3D assets, our method can generate the corresponding results that basically align with the coarse geometry defined by the hand-drawn sketch (see the first row in Figure 3.
In addition, as shown in the first two rows of Figure 3, given a fixed sketch, our method has the ability to generate different photo-realistic 3D scenes that conform to the corresponding text prompts. For instance, Control3D can generate “peacock themed” and “ocean themed” dresses that match the same input sketch by using different text prompts. Meanwhile, we also present another interesting case, by re-using the same text prompt and feeding different input sketches. As shown in the first two rows in Figure 3, our method has the flexibility to demonstrate shape controls while preserving the same text-driven appearances. For example, Control3D is able to generate a “full skirt” and a “midi skirt” with the same appearance theme. The above observations demonstrate that our Control3D may potentially enable many interesting 3D applications (such as recontextualization and reshaping), which would otherwise require tedious manual effort to tackle using traditional 3D modeling techniques.
Qualitative Comparisons. To the best of our knowledge, our work is the first attempt to perform controllable text-to-3D generation with hand-drawn sketches. Hence, in the absence of an existing benchmark for comparison, we have to compare our method with existing text-to-3D generation methods which are solely conditioned on text prompts. Herein we compare our Control3D with five typical baselines. 1) CLIP-Mesh (Mohammad Khalid et al., 2022), a zero-shot text-to-3D generation method using a pre-trained image-text model (i.e., CLIP (Radford et al., 2021)). 2) DreamField (Jain et al., 2022), which combines neural radiance fields with CLIP to synthesis diverse 3D objects form text prompt. 3) DreamFusion*: As primary DreamFusion (Poole et al., 2023) leverages image diffusion priors from their private model Imagen (Saharia et al., 2022), we capitalize on the publicly available 2D diffusion model (Stable Diffusion) and reimplement DreamFusion based on (Tang, 2022), namely DreamFusion*. 4) Latent-NeRF (Metzer et al., 2023), which learns a NeRF model on a latent feature space instead of in RGB pixel space, using a score distillation sampling loss in the latent space of Stable Diffusion. 5) Score Jacobian Chaining (SJC) (Wang et al., 2023), is another score distillation sampling baesd text-to-3D framework. It is worthy to note that the recent Magic3D (Lin et al., 2023) has shown high-quality text-to-3D generation results, it is excluded from comparisons since it relies on a private diffusion model eDiff-I (Balaji et al., 2022) that is unavailable to the research community.
We depict the qualitative comparisons in Figure 4. As illustrated in this figure, CLIP-Mesh and DreamField show somewhat inferior capability of shape generation, making it difficult to generate plausible 3D shapes. Taking the fifth row (Figure 4(d)) as an example, when using the text prompt “an expensive office chair”, CLIP-Mesh and DreamField generate chairs that are distorted and do not accurately match real-world chairs’ structures. Although Dreamfusion* can generate reasonable 3D shapes, it encounters challenges in the generation of precise and realistic 3D textures, which consequently lead to unrealistic visual appearances. For instance, given the text prompt “an imperial state crown of england” and “blue bird, highly detailed”, DreamFusion* can generate the accurate shapes of “crown” and “bird”, but falls short in rendering fine texture details. While NeRFs operate in image space, DreamFusion* encodes rendered RGB images to a latent space in each and every training step for applying score distillation sampling with the publicly available Latent Diffusion Model (i.e., Stable Diffusion). Compare with the original DreamFusion which performs score distillation sampling in the standard RGB space by using their private RGB space diffusion model, this degraded guidance in latent space of DreamFusion* is somewhat insufficient and thus result in degenerated text-to-3D solutions. Instead, Latent-NeRF formulates the NeRF in the latent space, where the NeRF is optimized to render 2D feature maps in Stable Diffusion’s latent space. These feature maps can easily be transformed back to RGB space through Stable Diffusion’s image decoder. In this way, Latent-NeRF produces more textual details than DreamFusion*. However, Latent-NeRF’s results still frequently suffer from blurry and diffuse issues.
In contrast, SJC and our Control3D generate much better 3D structures than the aforementioned baselines. Furthermore, when compared to SJC, our Control3D achieves higher 3D quality in terms of both geometry and texture. On the one hand, the hand-drawn sketch already depicts a well-drafted geometry and thus guides the NeRF model to generate plausible 3D shapes through our well-designed conditioned score sampling distillation and sketch consistency losses. On the other hand, with the sketch guidance, the text-conditioned probability density from the large-scale diffusion model has been narrowed down to a more compact region, which makes the underlying NeRF model easier to learn a high-fidelity texture. Accordingly, Our Control3D manages to control text-to-3D generation with a human-drawn sketch, while all the baseline methods lack this ability.
User study. We additionally conducted a user study to quantitatively evaluate Control3D against two diffusion based baseline models (i.e., Latent-NeRF and SJC) by comparing each pair. We invite 6 participants and show them two videos side by side in each test case. The videos are rendered from a canonical view by two different methods using the same text prompt. We then ask participants to choose the better one by jointly considering the following three aspects: (1) the alignment to the text prompt, (2) the fidelity of the visual appearance and (3) the accuracy of the geometry. According to all participants’ feedback, we measure the user preference score of one method as the percentage of its generated results that are preferred. Table 1 shows the results of the user study. In general, our Control3D significantly outperforms the baseline methods with higher user preference rates.
Ablation study. To enable controllable text-to-3D generation with a hand-drawn sketch, we design two loss terms: the conditioned score sampling distillation loss () in Eq. 7 and sketch consistency loss () in Eq. 8. In this section, we investigate the effectiveness of each design. We depict the results of each ablated run in Figure 5. Text-only is the base model SJC that creates 3D scenes only adhering to the semantics of the input text prompt. Instead, when is employed, the generated 3D scenes conform to both the input sketch and text prompt. This highlights the critical effectiveness of for text-to-3D generation with sketch condition. However, when only is applied, the generated 3D shapes don’t precisely match the input sketch and may be distorted in local region. By utilizing an additional sketch consistency constraint , the shape mismatch and artifact issue is clearly alleviated. This demonstrates the advantage of our designed sketch consistency loss in Eq. 8.
Conclusion
In this paper, we have proposed Control3D, the first attempt to enhance user controllability in text-to-3d generation by incorporating hand-drawn sketch conditions. Specifically, a 2D conditioned diffusion model (ControlNet) is remoduled to optimize a Neural Radiance Field (NeRF), encouraging each view of the 3D scene to align with the given text prompt and hand-drawn sketch. Moreover, we propose a novel sketch consistency loss that explicitly encourages the geometric consistency between synthetic 3D scene and the given sketch. The extensive experiments demonstrate that the proposed method can generate accurate and faithful 3D scenes that closely align with the input text prompts and sketches. Our Control3D provides a promising foundation for future research in controllable text-to-3D generation, which will lead to more creative and intuitive ways to generate 3D content.