Zero-Shot Text-Guided Object Generation with Dream Fields
Ajay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel, Ben Poole
Introduction
Detailed 3D object models bring multimedia experiences to life. Games, virtual reality applications and films are each populated with thousands of object models, each designed and textured by hand with digital software. While expert artists can author high-fidelity assets, the process is painstakingly slow and expensive. Prior work leverages 3D datasets to synthesize shapes in the form of point clouds, voxel grids, triangle meshes, and implicit functions using generative models like GANs . These approaches only support a few object categories due to small labeled 3D shape datasets. But multimedia applications require a wide variety of content, and need both 3D geometry and texture.
In this work, we propose Dream Fields, a method to automatically generate open-set 3D models from natural language prompts. Unlike prior work, our method does not require any 3D training data, and uses natural language prompts that are easy to author with an expressive interface for specifying desired object properties. We demonstrate that the compositionality of language allows for flexible creative control over shapes, colors and styles.
A Dream Field is a Neural Radiance Field (NeRF) trained to maximize a deep perceptual metric with respect to both the geometry and color of a scene. NeRF and other neural 3D representations have recently been successfully applied to novel view synthesis tasks where ground-truth RGB photos are available. NeRF is trained to reconstruct images from multiple viewpoints. As the learned radiance field is shared across viewpoints, NeRF can interpolate between viewpoints smoothly and consistently. Due to its neural representation, NeRF can be sampled at high spatial resolutions unlike voxel representations and point clouds, and are easy to optimize unlike explicit geometric representations like meshes as it is topology-free.
However, existing photographs are not available when creating novel objects from descriptions alone. Instead of learning to reconstruct known input photos, we learn a radiance field such that its renderings have high semantic similarity with a given text prompt. We extract these semantics with pre-trained neural image-text retrieval models like CLIP , learned from hundreds of millions of captioned images. As NeRF’s volumetric rendering and CLIP’s image-text representations are differentiable, we can optimize Dream Fields end-to-end for each prompt. Figure Zero-Shot Text-Guided Object Generation with Dream Fields illustrates our method.
In experiments, Dream Fields learn significant artifacts if we naively optimize the NeRF scene representation with textual supervision without adding additional geometric constraints (Figure 3). We propose general-purpose priors and demonstrate that they greatly improve the realism of results. Finally, we quantitatively evaluate open-set generation performance using a dataset of diverse object-centric prompts.
Using aligned image and text models to optimize NeRF without 3D shape or multi-view data,
Dream Fields, a simple, constrained 3D representation with neural guidance that supports diverse 3D object generation from captions in zero-shot, and
Simple geometric priors including transmittance regularization, scene bounds, and an MLP architecture that together improve fidelity.
Related Work
Our work is primarily inspired by DeepDream and other methods for visualizing the preferred inputs and features of neural networks by optimizing in image space . These methods enable the generation of interesting images from a pre-trained neural network without the additional training of a generative model. Closest to our work is , which studies differentiable image parameterizations in the context of style transfer. Our work replaces the style and content-based losses from that era with an image-text loss enabled by progress in contrastive representation learning on image-text datasets . The use of image-text models enables easy and flexible control over the style and content of generated imagery through textual prompt design. We optimize both geometry and color using the differentiable volumetric rendering and scene representation provided by NeRF, whereas was restricted to fixed geometry and only optimized texture. Together these advances enable a fundamentally new capability: open-ended text-guided generation of object geometry and texture.
Concurrently to Dream Fields, a few early works have used CLIP to synthesize or manipulate 3D object representations. CLIP-Forge generates multiple object geometries from text prompts using a CLIP embedding-conditioned normalizing flow model and geometry-only decoder trained on ShapeNet categories. Still, CLIP-Forge generalizes poorly outside of ShapeNet categories and requires ground-truth multi-view images and voxel data. Text2Shape learns a text-conditional Wasserstein GAN to synthesize novel voxelized objects, but only supports finite resolution generation of individual ShapeNet categories. In , object geometry is optimized evolutionarily for high CLIP score from a single view then manually colored. ClipMatrix edits the vertices and textures of human SMPL models to create stylized, deformable humanoid meshes. creates an interactive interface to edit signed-distance fields in localized regions, though they do not optimize texture or synthesize new shapes. Text-based manipulation of existing objects is complementary to us.
For images, there has been an explosion of work that leverages CLIP to guide image generation. Digital artist Ryan Murdock (@advadnoun) used CLIP to guide learning of the weights of a SIREN network , similar to NeRF but without volume rendering and focused on image generation. Katherine Crowson (@rivershavewings) combined CLIP with optimization of VQ-GAN codes and used diffusion models as an image prior . Recent work from Mario Klingemann (@quasimondo) and have shown how CLIP can be used to guide GAN models like StyleGAN . Some works have optimized parameters of vector graphics, suggesting CLIP guidance is highly general . These methods highlighted the surprising capacity of what image-text models have learned and their utility for guiding 2D generative processes. Direct text to image synthesis with generative models has also improved tremendously in recent years , but requires training large generative models on large-scale datasets, making such methods challenging to directly apply to text to 3D where no such datasets exist.
There is also growing progress on generative models with NeRF-based generators trained solely from 2D imagery. However, these models are category-specific and trained on large datasets of mostly forward-facing scenes , lacking the flexibility of open-set text-conditional models. Shape-agnostic priors have been used for 3D reconstruction .
Background
Our method combines Neural Radiance Fields (NeRF) with an image-text loss from . We begin by discussing these existing methods, and then detail our improved approach and methodology that enables high quality text to object generation.
NeRF parameterizes a scene’s density and color using a multi-layer perceptron (MLP) with parameters trained with a photometric loss relying on multi-view photographs of a scene. In our simplified model, the NeRF network takes in a 3D position and outputs parameters for an emission-absorption volume rendering model: density and color . Images can be rendered from desired viewpoints by integrating color along an appropriate ray, , for each pixel according to the volume rendering equation:
The integral is known as “transmittance” and describes the probability that light along the ray will not be absorbed when traveling from (the near scene bound) to . In practice , these two integrals are approximated by breaking up the ray into smaller segments within which and are assumed to be roughly constant:
For a given setting of MLP parameters and pose , we determine the appropriate ray for each pixel, compute rendered colors and transmittances, and gather the results to form the rendered image, and transmittance .
In order for the MLP to learn high frequency details more quickly , the input is preprocessed by a sinusoidal positional encoding before being passed into the network:
where is referred to as the number of “levels” of positional encoding. In our implementation, we specifically apply the integrated positional encoding (IPE) proposed in mip-NeRF to combat aliasing artifacts combined with a random Fourier positional encoding basis with frequency components sampled according to
2 Image-text models
Method
In this section, we develop Dream Fields: a zero-shot object synthesis method given only a natural language caption.
Building on the NeRF scene representation (Section 3.1), a Dream Field optimizes an MLP with parameters that produces outputs and representing the differential volume density and color of a scene at every 3D point . This field expresses object geometry via the density network. Our object representation is only dependent on 3D coordinates and not the camera’s viewing direction, as we did not find it beneficial. Given a camera pose , we can render an image and compute the transmittance using segments via (4). Segments are spaced at roughly equal intervals with random jittering along the ray. The number of segments, , determines the fidelity of the rendering. In practice, we fix it to 192 during optimization.
2 Objective
How can we train a Dream Field to represent a given caption? If we assume that an object can be described similarly when observed from any perspective, we can randomly sample poses and try to enforce that the rendered image matches the caption at all poses. We can implement this idea by using a CLIP network to measure the match between a caption and image given parameters and pose :
We primarily use image and text encoders from CLIP , which has a Vision Transformer image encoder and masked transformer text encoder trained contrastively on a large dataset of 400M captioned 2242 images. We also use a baseline Locked Image-Text Tuning (LiT) ViT B/32 model from trained via the same procedure as CLIP on a larger dataset of billions of higher-resolution (2882) captioned images. The LiT training set was collected following a simplified version of the ALIGN web alt-text dataset collection process and includes noisy captions.
Figure Zero-Shot Text-Guided Object Generation with Dream Fields shows a high-level overview of our method. DietNeRF proposed a related semantic consistency regularizer for NeRF based on the idea that “a bulldozer is a bulldozer from any perspective”. The method computed the similarity of a rendered and a real image. In contrast, (7) compares rendered images and a caption, allowing it to be used in zero-shot settings when there are no object photos.
3 Challenges with CLIP guidance
4 Pose sampling
Image data augmentations such as random crops are commonly used to improve and regularize image generation in DeepDream and related work. Image augmentations can only use in-plane 2D transformations. Dream Fields support 3D data augmentations by sampling different camera pose extrinsics at each training iteration. We uniformly sample camera azimuth in 360∘ around the scene, so each training iteration sees a different orientation of the object. As the underlying scene representation is shared, this improves the realism of object geometry. For example, sampling azimuth in a narrow interval tended to create flat, billboard geometry.
5 Encouraging coherent objects through sparsity
To remove near-field artifacts and spurious density, we regularize the opacity of Dream Field renderings. Our best results maximize the average transmittance of rays passing through the volume up to a target constant. Transmittance is the probability that light along ray is not absorbed by participating media when passing between point along the ray and the near plane at (2). We approximate the total transmittance along the ray as the joint probability of light passing through discrete segments of the ray according to Eq. (4). Then, we define the following transmittance loss:
This encourages a Dream Field to increase average transmittance up to a target transparency . We use in experiments. is annealed in from over 500 iterations to smoothly introduce transparency, which improves scene geometry and is essential to prevent completely transparent scenes. Scaling preserves object cross sectional area for different focal and object distances.
When the rendering is alpha-composited with a simple white or black background during training, we find that the average transmittance approaches , but the scene is diffuse as the optimization populates the background. Augmenting the scene with random background images leads to coherent objects. Dream Fields use Gaussian noise, checkerboard patterns and the random Fourier textures from as backgrounds. These are smoothed with a Gaussian blur with randomly sampled standard deviation. Background augmentations and a rendering during training are shown in Figure 4.
We qualitatively compare (9) to baseline sparsity regularizers in Figure 5. Our loss is inspired by the multiplicative opacity gating used by . However, the gated loss has optimization challenges in practice due in part to its non-convexity. The simplified additive loss is more stable, and both are significantly sharper than prior approaches for sparsifying Neural Radiance Fields.
6 Localizing objects and bounding scene
When Neural Radiance Fields are trained to reconstruct images, scene contents will align with observations in a consistent fashion, such as the center of the scene in NeRF’s Realistic Synthetic dataset . Dream Fields can place density away from the center of the scene while still satisfying the CLIP loss as natural images in CLIP’s training data will not always be centered. During training, we maintain an estimate of the 3D object’s origin and shift rays accordingly. The origin is tracked via an exponential moving average of the center of mass of rendered density. To prevent objects from drifting too far, we bound the scene inside a cube by masking the density .
7 Neural scene representation architecture
The NeRF network architecture proposed in parameterizes scene density with a simple 8-layer MLP of constant width, and radiance with an additional two layers. We use a residual MLP architecture instead that introduces residual connections around every two dense layers. Within a residual block, we find it beneficial to introduce Layer Normalization at the beginning and increase the feature dimension in a bottleneck fashion. Layer Normalization improves optimization on challenging prompts. To mitigate vanishing gradient issues in highly transparent scenes, we replace ReLU activations with Swish and rectify the predicted density with a softplus function. Our MLP architecture uses 280K parameters per scene, while NeRF uses 494K parameters.
Evaluation
We evaluate the consistency of generated objects with their captions and the importance of scene representation, then show qualitative results and test whether Dream Fields can generalize compositionally. Ablations analyze regularizers, CLIP and camera poses. Finally, supplementary materials have further examples and videos.
3D reconstruction methods are evaluated by comparing the learned geometry with a ground-truth reference model, e.g. with Chamfer Distance. Novel view synthesis techniques like LLFF and NeRF do not have ground truth models, but compare renderings to pixel-aligned ground truth images from held-out poses with PSNR or LPIPS, a deep perceptual metric .
As we do not have access to diverse captioned 3D models or captioned multi-view data, Dream Fields are challenging to evaluate with geometric and image reference-based metrics. Instead, we use the CLIP R-Precision metric from the text-to-image generation literature to measure how well rendered images align with the true caption. In the context of text-to-image synthesis, R-Precision measures the fraction of generated images that a retrieval model associates with the caption used to generate it. We use a different CLIP model for learning the Dream Field and computing the evaluation metric. As with NeRF evaluation, the image is rendered from a held-out pose. Dream Fields are optimized with cameras at a angle of elevation and evaluated at elevation. For quantitative metrics, we render at resolution 1682 during training as in . For figures, we train with a 50 higher resolution of 2522.
We collect an object-centric caption dataset with 153 captions as a subset of the Common Objects in Context (COCO) dataset (see supplement for details). Object centric examples are those that have a single bounding box annotation and are filtered to exclude those captioned with certain phrases like “extreme close up”. COCO includes 5 captions per image, but only one is used for generation. Hyperparameters were manually tuned for perceptual quality on a set of 20-74 distinct captions from the evaluation set, and are shared across all other scenes. Additional dataset details and hyperparameters are included in the supplement.
2 Analyzing retrieval metrics
In the absence of 3D training data, Dream Fields use geometric priors to constrain generation. To evaluate each proposed technique, we start from a simplified baseline Neural Radiance Field largely following and introduce the priors one-by-one. We generate two objects per COCO caption using different seeds, for a total of 306 objects. Objects are synthesized with 10K iterations of CLIP ViT B/16 guided optimization of 168168 rendered images, bilinearly upsampled to the contrastive model’s input resolution for computational efficiency. R-Precision is computed with CLIP ViT B/32 and LiT B/32 to measure the alignment of generations with the source caption.
Table 5.2 reports results. The most significant improvements come from sparsity, scene bounds and architecture. As an oracle, the ground truth images associated with object-centric COCO captions have high R-Precision. The NeRF representation converges poorly and introduces aliasing and banding artifacts, in part from its use of axis-aligned positional encodings.
We instead combine mip-NeRF’s integrated positional encodings with random Fourier features, which improves qualitative results and removes a bias toward axis-aligned structures. However, the effect on precision is neutral or negative. The transmittance loss in combination with background augmentations significantly improves retrieval precision 18% and 15.6%, while the transmittance loss is not sufficient on its own. This is qualitatively shown in Figure 5. Our MLP architecture with residual connections, normalization, bottleneck-style feature dimensions and smooth nonlinearities further improves the R-Precision 8% and 2%. Bounding the scene to a cube improves retrieval 13% and 11%. The additional bounds explicitly mask density and concentrate samples along each ray.
We also scale up Dream Fields by optimizing with an image-text model trained on a larger captioned dataset of 3.6B images from . We use a ViT B/32 model with image and text encoders trained from scratch. This corresponds to the uu configuration from , following the CLIP training procedure to learn both encoders contrastively. The LiT ViT encoder used in our experiments takes higher resolution 2882 images while CLIP is trained with 2242 inputs. Still, LiT B/32 is more compute-efficient than CLIP B/16 due to the larger patch size in the first layer.
LiT does not significantly help R-Precision when optimizing Dream Fields with low resolution renderings, perhaps because the CLIP B/32 model used for evaluation is trained on the same dataset as the CLIP B/16 model in earlier rows. Optimizing for longer with higher resolution 2522 renderings closes the gap. LiT improves visual quality and sharpness (Appendix A), suggesting that improvements in multimodal image-text models transfer to 3D generation.
In Figure 6, we show non-cherrypicked generations that test the compositional generalization of Dream Fields to fine-grained variations in captions taken from the website of . We independently vary the object generated and stylistic descriptors like shape and materials. DALL-E also had a remarkable ability to combine concepts in prompts out of distribution, but was limited to 2D image synthesis. Dream Fields produces compositions of concepts in 3D, and supports fine-grained variations in prompts across several categories of objects. Some geometric details are not realistic, however. For example, generated snails have eye stalks attached to their shell rather than body, and the generated green vase is blurry.
Ablating sparsity regularizers While we regularize the mean transmittance, other sparsity losses are possible. We compare unregularized Dream Fields, perturbations to the density , regularization with a beta prior on transmittance , multiplicative gating versions of and our additive regularizer in Figure 5. On real-world scenes, NeRF added Gaussian noise to network predictions of the density prior to rectification as a regularizer. This can encourage sharper boundary definitions as small densities will often be zeroed by the perturbation. The beta prior from Neural Volumes encourages rays to either pass through the volume or be completely occluded:
The multiplicative loss is inspired by the opacity scaling of for feature visualization. We scale the CLIP loss by a clipped mean transmittance:
Table 2 compares the regularizers, showing that density perturbations and the beta prior improve R-Precision 12.4% and 15%, respectively. Scenes with clipped mean transmittance regularization best align with their captions, 26.8% over the baseline. The beta prior can fill scenes with opaque material even without background augmentations as it encourages both high and low transmittance. Multiplicative gating works well when clipped to a target and with background augmentations, but is also non-convex and sensitive to hyperparameters. Figure 7 shows the effect of varying the target transmittance with an additive loss.
Varying optimized camera poses Each training iteration, Dream Fields samples a camera pose to render the scene. In experiments, we used a full 360∘ sampling range for the camera’s azimuth, and fixed the elevation. Figure 8 shows multiple views of a bird when optimizing with smaller azimuth ranges. In the left-most column, a view from the central azimuth (frontal) is shown, and is realistic for all training configurations. Views from more extreme angles (right, left, rear view columns) have artifacts when the Dream Field is optimized with narrow azimuth ranges. Training with diverse cameras is important for viewpoint generalization.
There are a number of limitations in Dream Fields. Generation requires iterative optimization, which can be expensive. 2K-20K iterations are sufficient for most objects, but more detail emerges when optimizing longer. Meta-learning or amortization could speed up synthesis.
We use the same prompt at all perspectives. This can lead to repeated patterns on multiple sides of an object. The target caption could be varied across different camera poses. Many of the prompts we tested involve multiple subjects, but we do not target complex scene generation partly because CLIP poorly encodes spatial relations . Scene layout could be handled in a post-processing step.
The image-text models we use to score renderings are not perfect even on ground truth training images, so improvements in image-text models may transfer to 3D generation. Our reliance on pre-trained models inherits their harmful biases. Identifying methods that can detect and remove these biases is an important direction if these methods are to be useful for larger-scale asset generation.
Our work has begun to tackle the difficult problem of object generation from text. By combining scalable multi-modal image-text models and multi-view consistent differentiable neural rendering with simple object priors, we are able to synthesize both geometry and color of 3D objects across a large variety of real-world text prompts. The language interface allows users to control the style and shape of the results, including materials and categories of objects, with easy-to-author prompts. We hope these methods will enable rapid asset creation for artists and multimedia applications.
We thank Xiaohua Zhai, Lucas Beyer and Andreas Steiner for providing pre-trained models on the LiT 3.6B dataset, Paras Jain, Kevin Murphy, Matthew Tancik and Alireza Fathi for useful discussions and feedback on our work, and many colleagues at Google for building and supporting key infrastructure. Ajay Jain is supported in part by the NSF GRFP under Grant Number DGE 1752814.
Appendix A Qualitative results and ablations
An explanatory video with more qualitative results, code, an interactive Colab notebook, and object-centric prompts are available at https://ajayj.com/dreamfields. The video includes 360∘ renderings where the camera orbits the object, as well as associated depth maps.
Figure 9 qualitatively compares generations using guidance from three different contrastive image-text models. Dream Fields generated with LiT B/32 are generally sharper than those with CLIP models, but all three can produce objects reflecting some aspects of the prompts.
Diversity of synthesized objects
In creative applications, users often want to select between multiple synthesized results. Dream Fields can synthesize multiple objects from the same prompt by changing the random seed before optimization. The seed changes the initialization of the NeRF weights, the camera pose sampled each iteration and the random background and crop augmentations. Figure 10 shows the effect of changing the seed for four object-centric COCO prompts. Changing the seed changes scene shape and layout. For example, the bus on the left is compressed into a cubical shape, while the bus on the right is elongated. Colors and textures are often similar across the seeds, though can also vary. Changing the seed is another dimension of control in addition to prompt engineering.
Appendix B Object Centric COCO captions dataset
Our Object Centric COCO dataset includes 153 test set prompts and 74 development set prompts. Several additional prompts are used for qualitative results, and are included in the main paper alongside figures. Captions are included along with code on the project website.
Appendix C Hyperparameters and training setup
Our Fourier feature positional encodings use frequency levels, while novel view synthesis applications with image supervision commonly use to fit high-frequency details in photographs. Low-frequency ablations in Table 1 use , which can improve convergence in the absence of our other geometric priors.
Rendering
Scenes are bounded to a cube with side length 2. The camera is sampled at a fixed radius of 4 units from the center of the cubical scene bounds and an elevation of 30∘ above the equator. Near and far planes are set at units from the camera based on the minimum and maximum possible distance to the corners of the cube. During training, we sample 192 points along each ray, spaced uniformly and jittered with uniform noise. Rendered 1682 views are cropped to 1542 and upsampled to CLIP’s input resolution for scenes where we compute qualitative metrics or 2522 views are cropped to 2242 for certain higher-quality visualizations. Crop sizes are selected to cover about 80% of the image area. At test time, we sample 512 points along the rays and render at a higher resolution equal to CLIP’s input size of 2242 or LiT’s input size of 2882 for computing R-Precision and 4002 for visualizations.
Optimization
MLP parameters are initialized with the Flax defaults: LeCun normal weights and zero bias for linear layers, and unit scaling and zero bias for layer normalization. The MLP is optimized with Adam with . Learning rate warms up exponentially from to over 1500 iterations, then is held constant. The camera origin is separately tracked with an exponential moving average with decay rate of the center of mass of rendered density. Figure 11 visualizes different forms of the transmittance loss.
Hyperparameter selection
Hyperparameters are manually tuned for visual quality on a development set of 74 object-centric COCO captions distinct from the test set reported in the paper. Most tuning is done on a smaller subset of 20 of the 74 captions, and hyperparameters are shared across all scenes.
Hardware
Optimization is done on 8 preemptible TPU cores, and 10K iterations takes approximately 1 hour 12 minutes. This means each Dream Field costs approximately $3-4 to generate on Google Cloud, which is economical for applications. Training is bottlenecked by MLP inference and backpropagation during volumetric rendering, not CLIP.
Appendix D Pixel and voxel baselines
We implemented 2D image optimization with a total variation loss and generative prior (CLIP Guided Diffusion), as well as a 3D voxel baselines to replace NeRF in Fig. 12. All results for this ablation optimize CLIP ViT B/16.
The 2D image is an RGB pixel grid, composited with random backgrounds during optimization similar to Dream Fields. Optimizing a single 2D RGB image does not produce a multi-view consistent 3D object, so other viewpoints cannot be rendered. Even with transmittance and TV regularization, the resulting image is noisy.
The voxel grid stores 1283 RGB and alpha values, interpolated trilinearly at ray sample points and composited without a neural network using the PyTorch3D library. Despite the transmittance loss, data augmentations and scene bounds, the voxel grid also has significant low-density artifacts. The voxel baseline has CLIP B/32 R-Precision 37.0%3.9, while NeRF has 59.8%2.8 (Tables 5.2, 3) with 16 fewer parameters, showing that the neural representation improves consistency with the input caption in a generalizable way. Using a hybrid representation with an explicit voxel grid followed by a smaller MLP head might improve computational efficiency of Dream Fields without degrading quality.
Appendix E Signed distance field parameterization
In early experiments, we learned scene density with the VolSDF parameterization where is a signed distance function implicitly defining the object surface and is the CDF of the Laplace distribution. This allows normal vector prediction with autodifferentiation and could improve the quality of the surface extracted from the radiance field. Dream Fields successfully train with this alternate parameterization and produce visually compelling objects. The SDF Eikonal loss introduces an additional loss weight hyperparameter, which benefits from some tuning. Alternate 3D representations are an interesting avenue for future work.
Appendix F Impact of optimization time
Qualitatively, additional details and hyper-realistic effects are added over the course of long runs, shown in our supplementary video. Some details are not realistic, like floating text related to the typographic attacks identified in .
More augmentations may help further regularize the optimization. These include more aggressive 2D image augmentations such as smaller random crops, and more 3D data augmentations including varying focal length, varying distance from the subject and varying elevation. 3D data augmentations are supported by our approach.