Latent-NeRF for Shape-Guided Generation of 3D Shapes and Textures
Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, Daniel Cohen-Or
Introduction
Text-guided image generation has seen tremendous success in recent years, primarily due to the breathtaking development in Language-Image models and diffusion models . These breakthroughs have also resulted in fast progression for text-guided shape generation . Most recently, it has been shown that one can directly use score distillation from a 2D diffusion model to guide the generation of a 3D object represented as a Neural Radiance Field (NeRF) .
While Text-to-3D can generate impressive results, it is inherently unconstrained and may lack the ability to guide or enforce a 3D structure. In this paper, we show how to introduce shape-guidance to the generation process to guide it toward a specific shape, thus allowing increased control over the generation process. Our method builds upon two models, a NeRF model , and a Latent Diffusion Model (LDM) . Latent Models, which apply the entire diffusion process in a compact latent space, have recently gained popularity due to their efficiency and publicly available pretrained checkpoints. As score distillation was previously applied only on RGB diffusion models, we first present two key modifications to the NeRF model that are better paired with guidance from a latent model. First, instead of representing our NeRF in the standard RGB space, we propose a Latent-NeRF which operates directly in the latent space of the LDM, thus avoiding the burden of encoding a rendered RGB image to a latent space for each and every guiding step. Secondly, we show that after training, one can easily transform a Latent-NeRF back into a regular NeRF. This allows further refinement in RGB space, where we can also introduce shading constraints or apply further guidance from RGB diffusion models . This is achieved by introducing a learnable linear layer that can be optionally added to a trained Latent-NeRF, where the linear layer is initialized using an approximate mapping between the latent and RGB values .
Our first form of shape-guidance is applied using a coarse 3D model, which we call a Sketch-Shape. Given a Sketch-Shape, we apply soft constraint during the NeRF optimization process to guide its occupancy based on the given shape. Easily combined with Latent-NeRF optimization, the additional constraint can be tuned to meet a desired level of strictness. Using a Sketch-Shape allows users to define their base geometry, where Latent-NeRF then refines the shape and introduces texture based on a guiding prompt.
We further present Latent-Paint, another form of shape-guidance where the generation process is applied directly on a given 3D mesh, and we have not only the structure but also the exact parameterization of the input mesh. This is achieved by representing a texture map in the latent space and propagating the guidance gradients directly to the texture map through the rendered mesh. By doing so, we allow for the first time to colorize a mesh using guidance from a pretrained diffusion model and enjoy its expressiveness.
We evaluate our different forms of guidance under a variety of scenarios and show that together with our latent-based guidance, they offer a compelling solution for constrained shape and texture generation.
Related Work
3D shape synthesis is a longstanding problem in computer graphics and computer vision. In recent years, with the emergence of neural networks, the research in 3D modeling has immensely advanced. The most conventional supervision type is applied directly with 3D shapes, through different representations such as implicit functions , meshes or point clouds . As 3D supervision is often difficult to obtain, other works use images to guide the generative task . In fact, even when 3D data is available, 2D renderings are sometimes chosen as the supervising primitive . For example, in GET3D , two generators are trained, one generates a 3D SDF, and the other a texture field. The output textured mesh is then obtained in a differentiable manner by utilizing DMTet . These generators are adversarially trained with a dataset of 2D images. In a diffusion model has been used to generate multiple views of a given input image. Yet, it has been trained in a supervised manner on a multi-view dataset, unlike our work which does not require a dataset.
Text-to-3D with 2D Supervision
Recently, the success of text-guided synthesis in numerous domains , has motivated a surge of works that use Language-Image models to guide 3D scenes representations. CLIP-Forge consists of two separate components, an implicit autoencoder conditioned on shape codes, and a normalizing flow model that is trained to generate shape codes according to CLIP embeddings. CLIP-Forge exploits the fact that CLIP has a joint text-image embedding space to train on image embeddings and infer on text embeddings, achieving text-to-shape capabilities. Text2Mesh introduced mesh colorization and geometric fine-tuning by optimizing an initial mesh through differential rendering and CLIP guidance. TANGO follows a similar optimization scheme, while improving results by considering an explicit shading model. CLIP-Mesh optimizes an initial spherical mesh according to a target text prompt, using a modified CLIP loss that accounts for the gap and ambiguity between image/text CLIP embeddings. Similarly to our method, they also use UV texture mapping to bake colors into the mesh. DreamFields employs CLIP guidance as well, but uses NeRFs to represent the 3D object instead of an explicit triangular mesh, together with a dedicated sparsity loss. CLIPNeRF pretrains a disentangled NeRF representation network on rendered object datasets, which is then used to constraint a NeRF scene optimization under CLIP loss, between random renderings of the NeRF and target image or text CLIP embedding. DreamFusion introduced, for the first time, the use of largely successful pretrained 2D diffusion models for text-guided 3D object generation. DreamFusion uses a proprietary 2D diffusion model to supervise the generation of 3D objects represented by NeRFs. To guide a NeRF scene using the pretrained diffusion model, the authors derive a Score-Distillation loss, see Section 3.1 for more details.
Neural Rendering
The recent rapid progression of neural networks has immensely advanced the performance of differential renderers. Particularly NeRF have shown astounding performance on novel view generation and relighting, also extending to other applications like 3D reconstruction . Thanks to their differentiable nature, it has been recently shown that one can introduce different neural objectives during training to guide the 3D modeling.
Method
Here we present our shape-guidance solution. We describe the Latent-NeRF framework, presented in Figure 2, and then introduce different guidance controls that can be combined with Latent-NeRF for controlling its generation. Yet, before showing our solution, we provide a quick overview of two recently proposed techniques that are highly relevant to our method.
Score Distillation is a method that enables using a diffusion model as a critic, i.e., using it as a loss without explicitly back-propagating through the diffusion process. It has been introduced in DreamFusion for guiding 3D generation using the Imagen model . To perform score distillation, noise is first added to a given image (e.g., one view of the NeRF’s output). Then, the diffusion model is used to predict the added noise from the noised image. Finally, the difference between the predicted and added noises is used for calculating per-pixel gradients. For NeRF, the gradients are back-propagated for updating the 3D NeRF model.
Going into more detail, at each iteration of the score distillation optimization, a rendered image is noised to a randomly drawn time step ,
where , and is a time-dependent constant specified by the diffusion model. Then, the per-pixel score distillation gradients are taken to be
where is the diffusion model’s denoiser (which approximates the noise to be removed), are the denoiser’s parameters, is an optional guiding text prompt, and is a constant multiplier that depends on . During training, gradients are propagated from the pixel gradients to the NeRF parameters and gradually change the 3D object. Please refer to for the complete details and derivation of Score Distillation. Note that DreamFusion uses the proprietary Imagen model that is very computationally demanding. We rely on the publicly available Stable Diffusion model and the re-implementation of DreamFusion (it operates in the RGB space and not the latent as we propose).
2 Latent-NeRF
We now turn to describe our Latent-NeRF approach. In this method, a NeRF model is optimized to render 2D feature maps in Stable Diffusion’s latent space . Latent-NeRF outputs four pseudo-color channels, , corresponding to the four latent features that stable diffusion operates over, and a volume density . Figure 2 illustrates this process. Representing the scene using NeRF implicitly imposes spatial consistency between different views, due to the spatial radiance field and rendering equation. Still, the fact that can be represented by a NeRF with spatial consistencies is non-trivial. Previous works showed that super-pixels in depend mainly on individual patches in the output image. This can be attributed to the high resolution () and low channel-wise depth () of this latent space, which encourages local dependency over the autoencoder’s image and latent spaces. Assuming is a near patch level representation of its corresponding RGB image makes the latents nearly equivariant to spatial transformations of the scene, which justifies the use of NeRFs for representing the 3D scenes.
The vanilla form of Latent-NeRF is text-guided, with no other constrains for the scene generation. In this setting, we employ the following loss:
where is the Score-Distillation loss depicted in Figure 2. Note, that the exact value of this loss is not accessible. Instead, the gradients implied by it are approximated by a single forward pass through the denoiser. These gradients are directly passed to the autograd solver. The loss suggested in prevents floating “radiance clouds” by penalizing the binary entropy of ill-defined background masks . Namely, it encourages a strict blending of the object NeRF and background NeRF.
RGB Refinement
Using Latent-NeRF, one may successfully learn to represent 3D scenes even when optimizing solely in latent space. Still, in some cases, it could be beneficial to further refine the model by fine-tuning it in pixel space, and have the NeRF model operate directly in RGB. To do so, we must convert the NeRF that was trained in latent space to a NeRF that operates in RGB. This requires converting the MLP’s output from the four latent channels to three RGB channels such that the initial rendered RGB image is close to the decoder output, when applied to the rendered latent of the original model. Interestingly, it has been shown that a linear approximation is sufficient to predict plausible RGB colors given a single four-channel latent super pixel, via the following transformation
which was calculated using pairs of RGB images and their corresponding latent codes over a collection of natural images. As our NeRF model is already composed of a set of fully connected layers, we simply add another linear layer that is initialized using the weights in Equation 3. This converts our Latent-NeRF to operate in pixel space and ensures that our refinement process starts from a valid result. The additional layer is then fine-tuned together with the rest of the model, to create the refined and final output. The overall fine-tuning procedure is illustrated in Figure 3.
3 Sketch-Shape Guidance
Next, we introduce a novel technique for guiding the Latent-NeRF generation based on a coarse geometry, which we call a Sketch-Shape. A Sketch-Shape is an abstract coarse alignment of simple 3D primitives like spheres, boxes, cylinders, etc., that together depict an outline of a more complex object. Figures 9, 10, 11 illustrate such simple shapes. Ideally, we would like the output density of our MLP to match that of the Sketch-Shape, such that the output Latent-NeRF result resembles the input shape. Nevertheless, we would also like the new NeRF to have the capacity to create new details and geometries that match the input text prompt and improve the fidelity of the shape. To achieve this lenient constraint, we encourage the NeRF’s occupancy to match the winding-number indicator of the Sketch-Shape, but with decaying importance near the surface to allow new geometries. This loss reads as
This loss implies that the occupancy should be well constrained away from the surface, and free to be set by score distillation near the surface. This loss is applied in addition to the Latent-NeRF loss, over the entire point set that is used by the NeRF’s volumetric rendering. represents the distance of from the surface, and is a hyperparameter that controls how lenient the loss is, i.e., lower values imply a tighter constraint to the input Sketch-Shape. Applying the loss only on the sampled point-set , makes it more efficient as these points are already evaluated as part of the Latent-NeRF rendering process.
4 Latent-Paint of Explicit Shapes
We now move to a more strict constraint, where the guidance is based on an exact structure of a given shape, e.g., provided in the form of a mesh. We call this approach Latent-Paint, which leads to the generation of novel textures for a given shape. Our method generates texture over a UV texture map, which can either be supplied by the input mesh, or calculated on-the-fly using XAtlas . To color a mesh, we first initialize a random latent texture image of size , where and can be chosen according to the desired texture granularity. We set them both to be in our experiments.
Figure 4 presents the training process. At each score distillation iteration, we render the mesh with a differentiable renderer to obtain a feature map that is pseudo-colored by the latent texture image. Then, we apply the score distillation loss from Equation 2 in the same way it is applied for Latent-NeRF. Yet, instead of back-propagating the loss to the NeRF’s MLP parameters, we optimize the deep texture image by back-propagating through the differentiable renderer. To get the final RGB texture image, we simply pass the latent texture image through Stable Diffusion’s decoder once, to get a larger high-quality RGB texture.
Evaluation
We now validate the effectiveness of our different forms of guidance through a variety of experiments.
We use the Stable Diffusion implementation by HuggingFace Diffusers, with the v1-4 checkpoint. For score distillation, we use the code-base provided by , with Instant-NGP as our NeRF model. Latent-NeRF usually takes less than 15 minutes to converge on a single V100, while using an RGB-NeRF with Stable Diffusion takes about 30 minutes, due to the increased overhead from encoding into the latent space. Note that DreamFusion takes about 1.5 hours on 4 TPUs. This clearly shows the computational advantage of Latent-NeRF.
1 Text-Guided Generation
We begin by demonstrating the effectiveness of the latent rendering approach with Latent-NeRF. In Figs. 1, 5 and 7, we show several results obtained by our method. In the supplementary material we provide additional results of different objects, including video visualizations. In Figure 5, we show the consistency of our learned shapes from several viewpoints. Next, we use the baseline set by DreamFusion to qualitatively compare our approach (with the proposed RGB refinement) against other methods. As can be seen in Figure 6, Latent-NeRF achieves significantly better results than DreamFields and CLIPMesh . We believe that the better quality of DreamFusion can be attributed to the high quality of Imagen , but unfortunately, we cannot validate this as the model is not publicly available to the community.
RGB Refinement
Figure 7 shows the quality improvement achieved by our RGB refinement method. It reveals that RGB refinement is mostly useful for complex objects or for regions with detailed textures. Refinement iterations in the RGB space are about slower than iterations in latent space, thus, increasing the runtime to more than 30 minutes. Thanks to our linear mapping from a Latent-NeRF to an RGB-NeRF, practitioners may apply the refinement method only after the 3D shape has already converged with the more efficient latent training. This allows a fast exploration of 3D shapes, and an optional polishing step with RGB refinement.
Textual-Inversion
As our Latent-NeRF is supervised by Stable-Diffusion, we can also use Textual Inversion tokens as part of the input text prompt. This allows conditioning the object generation on specific objects and styles, defined only by input images. Results using Textual Inversion are presented in Figure 8.
2 Sketch-Shape Guidance
Figure 9 shows different Sketch-Shape results with the same conditioning mesh. The different text prompts are able to guide the shape toward refined geometries that better match the text prompt. The rough Sketch-Shape in this figure, was quickly designed in Blender and allows us to easily define a coarse shape that guides the Latent-NeRF. Moreover, Figure 17 depicts an ablation over the lenient parameter from Eq. 4. When is set to , the generated mesh takes the form of the conditioning shape (shown in Figure 9). As grows, more details are added on top of the base shape, until little to no resemblance is observed at . Figure 10 contains additional results of different shapes generated with the same conditioning mesh, here a coarse house shape. The normals visualization (bottom row) shows that our method can add fine geometric details.
Figure 11 demonstrates that our proposed approach can successfully work with a variety of different Sketch-Shapes. Notice that our method handles a variety of different shapes and also works well with shapes extruded from 2D sketches. We also exhibit in Figure 14 the effectiveness of shape-guidance, by showing results of the same prompts with and without the shape loss.
3 Latent-Paint Generation
We tested Latent-Paint on a variety of input shapes shown in Figs. 12 and 13. As all of the shapes in these figures do not contain precomputed UV parameterization, we use XAtlas to compute such parameterization automatically. In contrast, the fish mesh in Figure 15 (obtained from ) already contains high quality UV parameterization, which we are also able work with. Figure 13 compares our Latent-Paint approach to two closely related methods, Tango and CLIPMesh . As can be seen, our approach achieves more precise textures thanks to the guidance from the diffusion model.
Note that Latent-Paint can work without assuming or computing any UV-map by simply optimizing per-face latent attributes. Yet, we found it better to use a UV-map for two main reasons: (i) The UV map makes the texture granularity independent of the geometric resolution, i.e., coarse geometries do not imply course colorization; and (ii) texture maps are easier to use with downstream applications like MeshLab and Blender .
Limitations
Our presented technique is yet a preliminary step towards the challenging goal of a comprehensive text-to-shape model that uses no 3D supervision. Still, the proposed latent framework has its limitations. To attain a plausible 3D shape, we use the same “prompt tweaking” used by DreamFusion , i.e., adding a directional text-prompt (e.g., ”front”, ”side” with respect to the camera) to the input text prompt. We find that this assistance tends to fail with our approach when applied to certain objects. Moreover, we find that even Stable Diffusion tends to generate unsatisfactory images when specifying the desired direction as shown in Figure 16. Additionally, similar to most works that employ diffusion models, there is a stochastic behavior to the results, such that the quality of the results may significantly vary between different seeds.
Conclusions
In this work, we introduced a latent framework for generating 3D shapes and textures using different forms of text and shape guidance. We first adapted the score distillation loss for LDMs, enabling the use of recent, powerful and publicly available text-to-image generation models for 3D shape generation. Successfully applying score distillation on LDMs results in a fast and flexible object generation framework. We then introduced shape-guided control on the generated model. We showed two versions of shape-guided generation, Sketch-Shape and Latent-Paint, and demonstrated their effectiveness for providing additional control over the generation process.
Typically, the notion of rendering refers to generating an image in pixel space. Here, we have presented a method that renders directly into the latent space of a neural model. We believe that our Latent-NeRF approach opens the avenue for more latent space rendering methods, which can gain from a compact and effective latent representation, and advance the use of neural models that operate in latent space rather than pixel space. Furthermore, our novel approach and its ease-of-use nature would encourage further research toward effective text-guided shape generation.
We thank Yuval Alaluf, Rinon Gal and Kfir Goldberg for their insightful comments. We would also like to thank Ido Richardson for his excellent mesh model designs used throughout the paper.