CLIP-Forge: Towards Zero-Shot Text-to-Shape Generation
Aditya Sanghi, Hang Chu, Joseph G. Lambourne, Ye Wang, Chin-Yi Cheng, Marco Fumero, Kamal Rahimi Malekshan
Introduction
Generating 3D shapes from text input has been a challenging and interesting research problem with both significant scientific and applied value . In the artificial intelligence and cognitive science research communities, researchers have long sought to bridge the two modalities of natural language and geometric shape . In practice, text-to-shape generation models are a key enabling component to new smart tools in creative design and manufacture as well as animation and games .
Significant progress has been made to connect text and image modalities . Recently, DALL-E and its associated pre-trained visual-textual embedding model CLIP has shown promising results on the problem of text-to-image generation . Notably, they have demonstrated strong zero-shot generalization while evaluated on tasks the model has not been specifically trained on. Shape generation is a more fundamental problem than image generation, because images are projections and renderings of the inherently 3D physical world. Therefore, one may wonder if the success in 2D can be transferred to the 3D domain. This turns out to be a non-trivial problem. Unlike the text-to-image case, where paired data is abundant, it is impractical to acquire huge paired datasets of texts and shapes.
Leveraging the progress of text-to-image generation, we present CLIP-Forge. As shown in Figure 2, we overcome the limitation of shape-text pair data scarcity via a simple and effective approach. We exploit the fact that 3D shapes can be easily and automatically rendered into images using standard graphics pipelines. We then utilize pre-trained image-text joint embedding models such as , which bring text and image embeddings in a similar latent space so that they can be used interchangeably. Hence, we can train a model using image embeddings, but at inference time replace it with text embeddings.
In CLIP-Forge, we first obtain a latent space for shapes via training an autoencoder, and we then train a normalizing flow network to model the distribution of shape embeddings conditioned on the image features obtained from the pre-trained image encoder . We use the renderings of 3D shapes and hence, no labels are required for training our model. During inference, we obtain text features of the given text query via the pre-trained text encoder. We then condition the normalizing flow network with text features to generate a shape embedding, which is converted into 3D shape through the shape decoder. In this process, CLIP-Forge requires no text labels for shapes, which means it can be extended easily to larger datasets. Since our method is fully feed-forward, it also has the advantage of avoiding the expensive inference time optimizations as employed in existing 2D approaches .
The main contributions of this paper are as follows:
We present a new method, CLIP-Forge, that generates 3D shapes directly from text as shown in Figure 1, without requiring paired text-shape labels.
Our method has an efficient generation process requiring no inference time optimization, can generate multiple shapes for a given text and can be easily extended to multiple 3D representation.
We provide extensive evaluation of our method in various zero-shot generation settings qualitatively and quantitatively.
Related Work
Zero-Shot Learning. Zero-shot learning is an important paradigm of machine learning, which typically aims to make predictions on classes that have never been observed during training, by exploiting certain external knowledge source. The literature originates from the image classification problem , and has been recently extended to generative models, in particular, the task of synthesizing images from text . As far as our best knowledge, our method is the first to bring this paradigm to the 3D shape domain, which enables efficient shape generation from natural language text input.
Applications of CLIP. A major building block of our method is CLIP , which shows groundbreaking zero-shot capability using a mechanism to connect text and image by bringing them closer in the latent space. Previous work such as ALIGN , has used a similar framework on noisy datasets. Recently, pre-trained CLIP has been used for several zero-shot downstream applications . The most similar previous work to ours is zero-shot image and drawing synthesis . Typically, these methods involve iteratively optimizing a random image to increase certain CLIP activations. There is still no clear way to apply them to 3D due to the significantly higher complexity. Our approach conditions a shape prior network with CLIP features, which has the advantages of significant speed-up and the capability to generate multiple shapes from a single text.
3D Shape Generation and Language. Recently, there has been tremendous progress in 3D shape generating in different data formats such as point cloud , voxel , implicit representation and mesh . While our method is not limited to produce one 3D data format, we mainly adopt the implicit representation in this work due to their simplicity and superior quality. More recently, methods that use text to localize objects in 3D scene have been explored . A metric learning method for text-to-shape generation is presented in . The main difference and advantage of our approach is its zero-shot capability which requires no text-shape labels.
Multi-Stage Training. In this work we follow a multi-stage training approach, where we first learn the embeddings of the target data and then learn a probabilistic encoding model for the learned embeddings. Such an approach as been explored in image generation and 3D shape generation . Concretely for CLIP-Forge, we first train a 3D shape autoencoder and then model the embeddings using normalizing flow.
Normalizing Flow. Generative models have extensive use cases such as content creation and editing. Flow-based generative networks is able to perform exact likelihood evaluation, while being efficient to sample from. They have been widely applied to a variety of tasks ranging from image generation , audio synthesis and video generation . Recently, normalizing flow has been brought to the 3D domain enabling fast generation of point clouds . In this paper, we employ a normalizing flow model to model the conditional distribution of latent shape representations given text and image embeddings.
Method
Our method requires a collection of 3D shapes without any associated text labels, which takes the format of . Each shape in the collection is comprised of a rendered image , a voxel grid , a set of query points in the 3D space , and space occupancies . As an overview, the CLIP-Forge training has two stages. In the first stage, we train an autoencoder with a voxel encoder and an implicit decoder. Once the autoencoder training is completed we obtain a shape embedding for each 3D shape in . In the second stage, we train a conditioned normalizing flow network to model and generate , which is conditioned with image features obtained from the CLIP image encoder using . During inference, we first convert the text to the interchangeable text-image latent space using the CLIP text encoder. We then condition the normalizing flow network with the given text features and a random vector sampled from the uniform Gaussian distribution to obtain a shape embedding. Finally, this shape embedding is converted to a 3D shape using the implicit decoder. The overall architecture is shown in Figure 3.
The autoencoder consists of an encoder and a decoder. We use an encoder to extract the shape embedding for the training shape collection, using of resolution as the input. We use a simple voxel network that comprises of a series of batch-normalized 3D convolution layers followed by linear layers. This can be written as:
where is augmented with a Gaussian noise. We find empirically injecting this noise improves the generation quality as later shown in the ablation study. This is also theoretically verified to improve results for conditional density estimation . We then pass through an implicit decoder. Our decoder architecture is inspired by the Occupancy Networks , which takes concatenated and as input. Our implicit decoder consists of linear layers with residual connections and predicts . We use a mean squared error loss between the predicted occupancy and the ground truth occupancy. Our framework is flexible and can be adapted to different forms of architectures. To showcase this, we use a PointNet as the encoder and a FoldingNet as the decoder that generates point clouds instead of occupancies, which are trained with a Chamfer loss .
2 Stage 2: Conditional Normalizing Flow
We train a normalizing flow network using and its corresponding rendered images . Note that each can include multiple images of the same shape from different rendering settings, such as changing camera viewpoints. We model the conditional distribution of using a RealNVP network with five layers, which transforms the distribution of into a normal distribution. We obtain the condition vector by passing through the ViT based CLIP image encoder , whose weights are frozen after pre-training. is concatenated with the transformed feature vector at each scale and translation coupling layers of RealNVP:
where and stand for the scale and translation function parameterized by a neural network. The intuition here is we split the object embedding into two parts where one part is modified using a neural network that is simple to invert, but still dependent on the remainder part in a non-linear manner. The splitting can be done in several ways by using a binary mask . In particular, we investigate two strategies: checkerboard masking and dimension-wise masking. The checkerboard masking has value 1 where the sum of spatial coordinates is odd, and 0 otherwise. The dimension-wise mask has value 1 for the first half of latent vector, and 0 for the second half. The masks are reversed after every layer. Finally, we impose a density estimation loss on the shape embeddings as:
where is the normalizing flow model, and is the Jacobian of at . We model the latent distribution as an unit Gaussian distribution.
3 Inference
During the inference phase, we convert a text query t into the text embedding using the CLIP text encoder, . As the CLIP image and text encoders are trained to bring the image and text embeddings in a joint latent space, we can simply use the text embedding as the condition vector for the normalizing flow model, i.e. =. Once we obtain the condition vector we can sample a vector from the normal distribution and use the reverse path of the flow model to obtain a shape embedding in . The normal distribution allows us to sample multiple times to obtain multiple shape embeddings for a given text query. We obtain the mean shape embedding by using the mean of the normal distribution. The mean shape embedding represents the prototype for a given text query. These shape embeddings are then converted to 3D shapes using the implicit decoder trained in stage 1.
Experiments
In this section, we first describe the experimental setup and then show qualitative and quantitative results. More results can be found in the supplementary material.
Dataset. For all of our experiments, we use the ShapeNet(v2) dataset which consists of 13 rigid object classes. We use the processed version of the data which consists of rendered images, voxel grids, query points and their occupancies from shapes as provided in .
Implementation Details. For both training stages, we use the Adam optimizer with a learning rate of 1e-4 and a batch size of 32. We train the stage 1 autoencoder for 300 epochs whereas we train the stage 2 conditional normalizing flow model for 100 epochs. For all the experiments below we use a latent size of 128 with a BatchNorm based voxel encoder and a ResNet based decoder inspired by the Occupancy Network . We use a RealNVP based network with dimension-wise masking for the flow model. The design decisions are discussed in the ablation study section and further details are provided in the supplemental material.
Evaluation Metrics. To evaluate our method thoroughly, we consider four criteria and several metrics for those criteria respectively. Furthermore, for some criteria we manually define a set of 234 text queries (or prompts). These queries include direct hyponyms for the ShapeNet categories from the WordNet taxonomy, sub-categories and relevant shape attributes for a given category (e.g. a round chair, a square table, etc.) across the ShapeNet(v2) dataset. The text queries are listed in the appendix. The criteria are as follows:
Reconstruction Quality. This criteria is mainly used to check the reconstruction capabilities of the stage 1 autoencoder on the test set. We use two metrics: Mean Square Error (MSE) on 30,000 sample query points and Intersection over Union (IOU) with voxel shapes.
Generation Quality. We use this criteria to evaluate the quality of generated shapes on text queries. We consider two metrics: Fréchet inception distance (FID) and Maximum Measure Distance (MMD) using IOU . To calculate FID and MMD, we first take 224 text queries as mentioned above and generate a mean shape embedding for each text query. We then generate resolution 3D objects for all the text queries. For FID, we compare the generated 3D shapes with the test dataset of ShapeNet. FID depends on a pre-trained network, for which we train a voxel classifier on the 13 ShapeNet classes and use the feature vector from the fourth layer. We provide more details in the appendix. In the case of MMD, for each generated shape we match a shape in the test dataset based on the highest IOU. We then average the IOU across all the text queries. Note, MMD is a variation of the Minimum Measure Distance as described in , which we believe is more suitable for implicit representations as we do not need to sample the surface.
Diversity Across Categories. To make sure we generate shapes across categories we design a new criteria. First, we generate the shapes based on the text queries as mentioned above. For each text query we have an assigned label. We then pass the generated voxels through the same classifier used to calculate the FID metric. We then report the accuracy based on the assigned label. We refer to this metric as Acc. throughout the text. Also note that the FID metric gives a good measure for diversity as we compare it with the test distribution.
Human Perceptual Evaluation. To evaluate CLIP-Forge’s ability to provide control over the generated shape using attribute, common name, and sub-category information from the text prompt, we conducted a perceptual evaluation using Amazon SageMaker Ground Truth and crowd workers from Mechanical Turk . More detail is provided in section 4.3.
We compare CLIP-Forge with text-to-shape generation models that are trained with direct supervision signals. The only existing paired text-shape dataset is provided by Text2Shape , which contains 56,399 natural language descriptions for ShapeNet objects within the chair and table categories. We train two supervised models using the Text2Shape dataset: text2shape-CMA uses the cross-modality alignment loss described in and text2shape-supervised uses a direct MSE loss in the embedding space. For both supervised baseline methods, we use the same CLIP text encoder and occupancy network shape encoder and decoder to ensure fair comparison. Table 1 shows the result on our text query set. It can be seen that CLIP-Forge significantly outperforms both supervised baselines in all evaluation metrics. In particular, we observe that text2shape-CMA generates generic shapes such as boxes and spheres that do not resemble specific objects. The text2shape-supervised baseline fails to generalize and tends to generate chair- and table- like shapes that it is trained on, despite the text query is irrelevant to these two categories.
2 Qualitative Results
We qualitatively evaluate generative capabilities of our method. First, in Figure 19 we show that our network can generate multiple and diverse shapes using a single text query. This can be useful in a design process for imagining new variations. Next, we show that our network can generate shapes based on category, sub-category, common semantic words, and common shape attributes as shown in Figure 5. It can be seen that our network captures semantic notion of the text query. Finally, we show the generated shapes from interpolation between two text inputs in Figure 20. The interpolation results imply that the conditioning space is smooth.
3 Human Perceptual Evaluation
In this study we measure whether providing additional detail in the text prompt gives rise to semantically appropriate changes in the generated shape. To evaluate if the shape changes are semantically correct, we used human evaluators from Amazon Mechanical Turk . The human evaluators were presented with pairs of images as shown in Figure 6(a). One image was generated using the ShapeNet(v2) category name (for example “a car”) while the other was generated using text which described a sub-category or shape attribute (for example “a truck” or “a round car”). The human evaluators were asked to identify which image best matched the sub-category or attribute text prompt. Each image pair was shown to 9 independent human evaluators. We record the fraction of image pairs for which more than half of the evaluators selected the image generated using the sub-category or attribute augmented prompt.
The results of the perceptual study are shown in Figure 6(b). The human evaluators correctly identified the model generated by the detailed prompt for 70.83% of the image pairs, showing that our method is able to utilize attribute and sub-category information in a way which is recognizable for humans. We see the attribute prompts produced shapes which were more easily identified than those from the sub-category prompts. One reason for this result is that the attribute augmented prompts give a clear description of how the object should look, while many of the sub-categories are less easily recognized given the quality of generation. For example “A circular bench” was correctly identified by 8/9 evaluators while “A laboratory bench” was not recognized by any of the 9 humans.
4 Choice of Prefix in Text Prompt
Designing a prompt can be challenging as small changes in words can potentially have a impact on our downstream task. In this experiment, we investigate how much does prompt selection effect the performance of our method. We specifically investigate what prefix to choose before a text query. The investigations are shown in Table 2. We find that prefix selection indeed has a effect on generation quality and diversity. A interesting avenue of future research would be to investigate prompt engineering .
5 CLIP-Forge for Point Cloud
In this section, we investigate if our method can be simply applied to a different representation, namely, a point cloud. As stated earlier we use the PointNet encoder and FoldingNet decoder . We use the same flow architecture as mentioned above. We train the network on the ShapeNet(v2) dataset. The results are shown in Fig 8. It can be seen that our method does a satisfactory job in generating 3D point clouds using text queries while using off the shelf point cloud encoders and decoders.
Ablation Studies
In this section, we discuss how different components of our algorithm affect our model. For all our ablation studies, we use the above mentioned hyperparameters for the autoencoder unless otherwise stated. For the flow model, we use RealNVP model with Checkerboard masking for most experiments unless otherwise stated.
In Table 3, we experiment with different parts of the autoencoder architecture. The first subsection of Table 3 investigates how adding noise in the latent space helps our model. Empirically it can be seen from the table that adding noise not only helps the reconstruction but also improves the generation and diversity of the shapes generated. Next we investigate the size of the latent space and find that our model works reasonably well while using a smaller latent size of 128. Finally, we explore different encoders and decoders for our model. The results indicate that our model can take different representations, point cloud, as the input to the encoder. We provide more details regarding the encoder and decoder in appendix.
2 Stage 2 Prior Design Choice
In this section, we investigate the design choice for the prior network. First, we investigate different conditioning mechanisms, namely, conditioning the affine coupling layers and conditioning the prior network. From Table 4, it can be seen the choice of conditioning matters and conditioning the affine layers is the most effective. This intuitively makes sense as we are conditioning multiple times as we are concatenating each coupling layer whereas we just condition the prior once. A similar phenomena is observed in the case of an architectures like , where they concatenated the condition vector multiple times.
In Table 4, we also investigate different masking techniques (Dimension masking and Checkered masking) and a distinct flow architecture: Masked Autoregressive Flow (MAF) . It can be seen from the table that both masking techniques are effective but Dimension masking (RealNVP-D) seems to be more effective than Checkered masking (RealNVP-C). Furthermore, we find that MAF flow prior network is not as effective as RealNVP. For the remainder of the ablation studies we use the Dimension Masking for RealNVP.
3 Number of Renderings
Next, we evaluate if using more views helps the generation quality and diversity. We report the results in Table 5. The views are randomly selected from the renderings as prepared in . It can be seen that using more views in general help improve the generation quality and diversity. As we are using a pre-trained CLIP model which is trained on natural images from different viewpoints, training using multiple views of shape renderings allows us to better capture the output distribution of CLIP model.
4 CLIP Architecture
In this section, we evaluate using different CLIP models to see how increasing the size of CLIP model and using ResNet or ViT based clip model effects our downstream task. We empirically observe from Table 5 that increasing the size of the model, i.e from ViT-B/32 to ViT-B/16, does not effect the text based generation too much. A more surprising result is that ResNet based CLIP model performs inferior to Visual Transformers. We hypothesize that patch based methods such as ViT focus more on the foreground object rather than the background. This is especially helpful in the case of image renderings.
Limitations and Future Work
We believe our method can be improved in several ways. Firstly, the quality of generation is still lacking and we believe a novel future avenue would be to combine ideas from local implicit methods . Furthermore, our work currently focuses on geometry and it would be interesting to integrate texture to our model. Finally, we are limited by CLIP’s trained data distribution and a potential future direction would be to fine tune it for a specfic dataset.
In terms of potential negative impact, language-driven 3D modeling tools enabled by CLIP-Forge might lower the technical barriers to 3D modeling and potentially reduce some tedious 3D modeling tasks for 3D modelers and animators. However, it brings a greater benefit of democratizing 3D content creation to the general public, similar to that everyone can take photos and make videos today.
Conclusion
We presented a method, CLIP-Forge, that can efficiently generate multiple 3D shapes while preserving the semantic meaning from a given text prompt. Our method requires no text-shape labels as training data, offering an opportunity to leverage shape-only datasets such as ShapeNet. Finally, we showed that our model can generate results on other representations such as point clouds and we thoroughly studied different components of the method.
Appendix A Architecture and Experiment Details
For all our experiments in the ablation section of the main paper, we run the second stage network with 3 different seeds and report the mean in the main paper. We take the best seed for the experiment section in the main paper to report the qualitative and quantitative results. The text queries (or prompts) used for classification FID, MMD, and Acc. are shown in Table 6. Note that these text queries are mostly taken from WordNet with added common synonyms and shape attributes. In Table 8, we show category-wise accuracy results of CLIP-Forge in the main paper’s Table 1. For our visualizations, we output a shape with resolution and use the rendering script inspired by . We use a set of different thresholds values and pick the threshold for different category that yields the best visual result.
In the main paper, we refer to the batch normalization based voxel encoder as VoxEnc, whereas when we add residual connection to VoxEnc we refer it to as ResVoxEnc. For both of these encoders, we have 4 3D convolution layers followed by a linear layer. The input to these encoder is a voxel representation based shape. We also experiment with a point cloud based encoder which is inspired by PointNet. The PointNet encoder has 5 linear layers followed by a max pooling operation. We then use an MLP followed by a final linear layer to project it to the latent size. The input to this encoder is a point cloud with 2048 sampled points. For the decoder, we refer to the residual connection based network as RN-OccNet. In this model, we concatenate the query locations with the latent code and pass it through a 5 block ResNet based decoder. We also experiment with conditioning the batchnorm of the decoder instead of concatenating it, which we refer to as CBN-OccNet. Both these decoders are inspired by OccNet . For our point cloud based generation experiments, we use a FoldingNet based decoder, where we use two folding based operations with a single square grid.
Finally, we use RealNVP for the prior model. We use 5 blocks of coupling layer containing translation and scale along with batch norm, where the masking is inverted after each block. Each network comprises of a 2-layer MLP followed by a linear layer. A 1024 hidden vector size is used. For the MAF model, we also use the same number of blocks and hidden vectors.
Appendix B Comparison with Supervised Models
In this section, we provide a more detailed comparison between CLIP-Forge and supervised methods. Note, it is not clear how to compare our zero-shot model with supervised models. As our end goal is to generate shapes across categories and text queries, we decide to use our original text query subset (mentioned above) and Shapenet (v2) test dataset as the test set. This test set ensures we test on commonly described words for different shape category as mentioned in WordNet . We consider two datasets: T2S is the annotated text-shape description dataset from text2shape which mainly contains information regarding texture, and SN13 is the ShapeNet (v2) subset containing 13 categories from .
T2S dataset has text labels only for chair and tableclass, so we train a supervised baseline model that has a linear layer connecting a pre-trained CLIP text encoder and a pre-trained occupancy network decoder using an L2 loss in the latent space. We use the same text encoder and shape decoder in this baseline to ensure a fair comparison. We compare the baseline model with our model which does not have use any supervision from text labels and is also trained on T2S shape dataset (chair and table only). The results are shown in the first part of Table 7. It can be seen from the table that our model can generate shapes in chair and table categories based on common words with higher quality despite not using any text label information.
To test baseline models on all of ShapeNet subset (SN13), as there is no text label data, we create a simple supervision signal by directly using the category name as the text for training the supervised model. The results are shown in second part of Table 7. It can be seen that our model outperforms the supervised method, demonstrating its stronger zero-shot generalization ability. This results also indicate that our model scales better with more data without requiring text-shape labels. In Figure 22, we show qualitative results of the supervised baselines, where the model fails to generate cars when trained on T2S, and fails to capture the details of sports car when trained with SN13.
Appendix C Category-wise Accuracy Results
We also report category-wise accuracy results obtained from our classifier for our method in Table 8. It can be generally noted that our method can generate shapes across all categories of Shapenet. However, accuracy across some categories such as airplane and car are higher than other categories such as boat and loudspeaker. We hypothesize that this may be due to some categories having larger data points during training compared to others.
Appendix D Comparison with Text2Img+Img2Shape
In this section, we compare our method to off-the-shelf networks that simply generate an image from text first and then generate a 3D shape from the image. We use pre-trained DALLE-mini for converting a text to image and use a pre-trained occupancy network with image encoder to convert an image to 3D shape. The results are shown in Fig. 9. It can be seen that the resulting shapes suffer from poor quality. This is mainly due to the domain gap between generated images and natural images such as distortion artifacts and unclean background.
Appendix E Effect of Threshold Parameter
Our results are strongly affected by the threshold used to create the occupancy value. We use a constant threshold value of for our metrics (Acc., FID and MMD) and human perceptual evaluations. However, for our visual results we do a grid search and choose the best threshold value. Figure 10, shows the visual results of different thresholds. It can be seen that different shapes require different threshold which depends on the category and local details of the shape. We believe that our metrics and human evaluation results can be further improved if a better technique is discovered for threshold tuning.
Appendix F Out of Distribution Generation
We also conduct experiments to see if the network can generate shapes based on text queries which are out of distribution from its training data. The results are shown in Figure 11. It can be seen from the results that the method tries to generate the desired shape based on its training dataset. We believe extending our method to generalize on out of distribution samples might be interesting avenue to explore for future work.
Appendix G Visual Results for Different Prefixes
In Figure 12, we show results for different prefixes. They indicate that for different prefixes there are small variations in generated shape. Moreover, in some prefixes such as “a rendering of”, the visual results are worse. It would be interesting to investigate other prompts or do prompt tuning as future work.
Appendix H Visual Results for more Descriptive Texts
We show additional results using text queries that are longer and more descriptive in Figure 21. It can be seen that CLIP-Forge is able to capture certain shape-related attributes. Non-shape related descriptions such as color is not captured but could potentially bias the generation. We believe that combining our method with semi-supervised learning can enable more fine control of shape generation using text.
Appendix I Additional Qualitative Results
In this part, we show more visual results for shape generation conditioned with text based on sub-category (Figure 13 and Figure 14), synonyms (Figure 15), shape attributes (Figure 16 and Figure 17) and common names (Figure 18). Moreover, we also show more visuals for text based multiple shape generation (Figure 19) and interpolation (Figure 20). It can be seen from all these results that our method is good at generating 3D shapes based on text queries. However, in same cases for example “a swivel chair”, it cannot construct all the details. Furthermore, on some sub-categories such as “an operating table” it cannot generate accurate shapes.
Appendix J Human Perceptual Evaluation
In the human perceptual evaluation described in section 4.3 of the main paper, crowd workers recruited through Amazon Mechanical Turk were shown pairs of images, one generated from the ShapeNet(v2) category name (see the first column of Figure 23) and the other from a detailed text prompt containing either subcategory or attribute information. The crowd workers were shown the detailed text prompt and asked to identify which of the two images it best describes. Nine crowd workers viewed each image pair and we record the number of times the model from the detailed text prompt is selected. For each detailed text prompt, this gives us a score from 0 to 9 indicating how effectively Clip-Forge can produce distinctive shapes which differ from the ShapeNet categories in a way which humans find semantically meaningful. In Table 6 the human evaluation scores are shown as colors for each query text for which the evaluation was conducted. Figure 23 shows a few examples in more detail. The second column of Figure 23 shows text prompts which produced distinctive shapes and the third column shows cases where the shapes were not as easily identified based on the text. We see that when the prompt elicited a very distinctive shape (‘A monster truck”, “A fighter plane”) a high fraction of the human raters were able to identify the correct model. In some cases the low score reflects a lack of resolution (for example “A swivel chair”, “A billiard table” and “A seaplane”). In the case of “A wheelchair”, Clip-Forge was unable to generate round wheels, but as the bottom of the legs were joined up this gave enough of an impression of wheels for humans to select the model. In the case of “A muscle car” Clip-Forge attempted to create the shape of a low form of a sports car, however the shape was not far enough from the generic car for the crowd workers to select it.