Text2Mesh: Text-Driven Neural Stylization for Meshes
Oscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim, Rana Hanocka
Introduction
Editing visual data to conform to a desired style, while preserving the underlying content, is a longstanding objective in computer graphics and vision . Key challenges include proper formulation of content, style, and the constituents for representing and modifying them.
To edit the style of a 3D object, we adapt a formulation of geometric content and stylistic appearance commonly used in computer graphics pipelines . We consider content as the global structure prescribed by a 3D mesh, which defines the overall shape surface and topology. We consider style as the object’s particular appearance or affect, as determined by its color and fine-grained (local) geometric details. We propose expressing the desired style through natural language (a text prompt), similar to how a commissioned artist is provided a verbal or textual description of the desired work. This is facilitated by recent developments in joint embeddings of text and images, such as CLIP . A natural cue for modifying the appearance of 3D shapes is through 2D projections, as they correspond with how humans and machines perceive 3D geometry. We use a neural network to synthesize color and local geometric details over the 3D input shape, which we refer to as a neural style field (NSF). The weights the NSF network are optimized such that the resulting 3D stylized mesh adheres to the style described by text. In particular, our neural optimization is guided by multiple 2D (CLIP-embedded) views of the stylized mesh matching our target text. Results of our technique, called Text2Mesh, are shown in Fig. 1. Our method produces different colors and local deformations for the same 3D mesh content to match the specified text. Moreover, Text2Mesh produces structured textures that are aligned with salient features (e.g. bricks in Fig. 2), without needing to estimate sharp 3D curves or a mesh parameterization . Our method also demonstrates global understanding; e.g. in Fig. 3 human body parts are stylized in accordance with their semantic role.
We use the weights of the NSF network to encode a stylization (e.g. color and displacements) over the explicit mesh surface. Meshes faithfully portray 3D shapes and can accurately represent sharp, extrinsic features using a high level of detail. Our neural style field is complementary to the mesh content, and appends colors and small displacements to the input mesh. Specifically, our neural style field network maps points on the mesh surface to style attributes (i.e., RGB colors and displacements).
We guide the NSF network by rendering the stylized 3D mesh from multiple 2D views and measuring the similarity of those views against the target text, using CLIP’s embedding space. However, a straightforward optimization of the 3D stylized mesh which maximizes the CLIP similarity score converges to a degenerate (i.e. noisy) solution (see Fig. 5). Specifically, we observe that the joint text-image embedding space contains an abundance of false positives, where a valid target text and a degenerate image (i.e. noise, artifacts) result in a high similarity score. Therefore, employing CLIP for stylization requires careful regularization.
We leverage multiple priors to effectively guide our NSF network. The 3D mesh input acts as a geometric prior that imposes global shape structure, as well as local details that indicate the appropriate position for stylization. The weights of the NSF network act as a neural prior (i.e. regularization technique), which tends to favor smooth solutions . In order to produce accurate styles which contain high-frequency content with high fidelity, we use a frequency-based positional encoding . We garner a strong signal about the quality of the neural style field by rendering the stylized mesh from multiple 2D views and then applying 2D augmentations. This results in a system which can effectively avoid degenerate solutions, while still maintaining high-fidelity results.
The focus of our work is text-driven stylization, since text is easily modifiable and can effectively express complex concepts related to style. Text prescribes an abstract notion of style, allowing the network to produce different valid stylizations which still adhere to the text. Beyond text, our framework extends to additional target modalities, such as images, 3D meshes, or even cross-modal combinations.
In summary, we present a technique for the semantic manipulation of style for 3D meshes, harnessing the representational power of CLIP. Our system combines the advantages of explicit mesh surfaces and the generality of neural fields to facilitate intuitive control for stylizing 3D shapes. A notable advantage of our framework is its ability to handle low-quality meshes (e.g., non-manifold) with arbitrary genus. We show that Text2Mesh can stylize a variety of 3D shapes with many different target styles.
Related Work
Text-Driven Manipulation. Our work is similar in spirit to image manipulation techniques controlled through textual descriptions embedded by CLIP . CLIP learns a joint embedding space for images and text. StyleCLIP perform CLIP-guided image editing using a pre-trained StyleGAN . VQGAN-CLIP leverage CLIP for text-guided image generation. Concurrent work uses CLIP to fine-tune a pre-trained StyleGAN , and for image stylization . Another concurrent work uses the ShapeNet dataset and CLIP to perform unconditional 3D voxel generation . The above techniques leverage a pre-trained generative network or a dataset to avoid the degenerate solutions common when using CLIP for synthesis. The first to leverage CLIP for synthesis without the need for a pre-trained network or dataset is CLIPDraw . CLIPDraw generates text-guided 2D vector graphics, which conveys a type of drawing style through vector strokes. Concurrent work uses CLIP to optimize over parameters of the SMPL human body model to create digital creatures. Prior to CLIP, text-driven control for deforming 3D shapes was explored using specialized 3D datasets.
Geometric Style Transfer in 3D. Some approaches analyze 3D shapes and identify similarly shaped geometric elements and parts which differ in style . Others transfer geometric style based on content/style separation . Other approaches are specific to categories of furniture , 3D collages , LEGO , and portraits . 3DStyleNet edits shape content with a part-aware low-frequency deformation and synthesizes colors in a texture map, guided by a target mesh. Mesh Renderer changes color and geometry driven by a target image. Liu et al. stylize a 3D shape by adding geometric detail (without color), and ALIGNet deforms a template shape to a target one. The above methods rely on 3D datasets, while other techniques use a single mesh exemplar for synthesizing geometric textures or producing mesh refinements . Shapes can be edited to contain cubic stylization , or stripe patterns . Unlike these methods, we consider a wide range of styles, guided by an intuitive and compact (text) specification.
Texture Transfer in 3D. Aspects of a 3D mesh style can be controlled by texturing a surface through mesh parameterization . However, most parameterization approaches place strict requirements on the quality of the input mesh (e.g., a manifold, non-intersecting, and low/zero genus), which do not hold for most meshes in the wild . We avoid parameterization altogether and opt to modify appearance using a neural field which provides a style value (i.e., an RGB value and a displacement) for every vertex on the mesh. Recent work explored a neural representation of texture , here we consider both color and local geometry changes for the manipulation of style.
Neural Priors and Neural Fields. A recent line of work leverages the inductive bias of neural networks for tasks such as image denoising , surface reconstruction , point cloud consolidation , image synthesis, and editing . Our framework leverages the inductive bias of neural networks to act as a prior which guides Text2Mesh away from degenerate solutions present in the CLIP embedding space. Specifically, our stylization network acts as a neural prior, which leverages positional encoding to synthesize fine-grained stylization details.
NeRF and follow ups have demonstrated success on 3D scene modeling. They leverage a neural field to represent 3D objects using network weights. However, neural fields commonly entangle geometry and appearance, which limits separable control of content and style. Moreover, they struggle to accurately portray sharp features, are slow to render, and are difficult to edit. Thus, several techniques were proposed enabling ease of control , and introducing acceleration strategies . Instead, our method uses a disentangled representation of a 3D object using an explicit mesh representation of shape and a neural style field which controls appearance. This formulation avoids parametrization, and can be used to easily manipulate appearance and generate high resolution outputs.
Method
In practice, we treat the given vertices of as query points into this field, and use a differentiable renderer to visualize the style over the given triangulation. Increasing the number of triangles in for the purposes of learning a neural field with finer granularity is trivial, e.g., by inserting a degree vertex (see Appendix B). Even using a standard GPU (11GB of VRAM) our method handles meshes with up to K triangles. We are able to render stylized objects using very high resolutions, as shown in the Appendix B.
Since our NSF uses low-dimensional coordinates as input to an MLP, this exhibits a spectral bias toward smooth solutions (e.g. see Fig. 5). To synthesize high-frequency details, we apply a positional encoding using Fourier feature mappings, which enables MLPs to overcome the spectral bias and learn to interpolate high-frequency functions . For every point its positional encoding is given by:
First, we normalize the coordinates to lie inside a unit bounding box. Then, the per-vertex positional encoding features are passed as input to an MLP , which then branches out to MLPs and . Specifically, the output of is a color , and the output of is a displacement along the vertex normal . To prevent content-altering displacements, we constrain to be in the range . To obtain our stylized mesh prediction , every point is displaced by and colored by . Vertex colors propagate over the entire mesh surface using an interpolation-based differentiable renderer . During training we also consider the displacement-only mesh , which is the same as without the predicted vertex colors (replaced by gray). Without the use of in our final loss formulation (Eq. 5), the learned geometric style is noisier ( ablation in Fig. 5).
2 Text-based correspondence
Our neural optimization is guided by the multi-modal embedding space provided by a pre-trained CLIP model. Given the stylized mesh and the displaced mesh , we sample views around a pre-defined anchor view and render them using a differentiable renderer. For each view, , we render two 2D projections of the surface, for and for . Next, we draw a 2D augmentation and (details in Sec. 3.3). We apply , to the full view and to the uncolored view, and embed them into CLIP space. Finally, we average the embeddings across all views:
where and is the cosine similarity between and . We repeat the above with new sampled augmentations times for each iteration. We note that the terms using and update , and while the term using only updates and . The separation into a geometry-only loss and geometry-and-color loss is an effective tool for encouraging meaningful changes in geometry ( in Fig. 5).
3 Viewpoints and Augmentations
Given an input 3D mesh and target text, we first find an anchor view. We render the 3D mesh at uniform intervals around a sphere and obtain the CLIP similarity for each view and target text. We select the view with the highest (i.e. best) CLIP similarity as the anchor view. Often there are multiple high-scoring views around the object, and using any of them as the anchor will produce an effective and meaningful stylization. See Appendix C for details.
We render multiple views of the object from randomly sampled views using a Gaussian distribution centered around the anchor view (with /4). We average over the CLIP-embedded views prior to feeding them into our loss, which encourages the network to leverage view consistency. For all our experiments, (number of sampled views). We show in the Appendix C that setting beyond does not meaningfully impact the results.
The 2D augmentations generated using and are critical for our method to avoid degenerate solutions (see Sec. 4.2). involves a random perspective transformation and generates both a random perspective and a random crop that is 10% of the original image. Cropping allows the network to focus on localized regions when making fine grained adjustments to the surface geometry and color. (- in Fig. 5). Additional details are given in Appendix D.
Experiments
We examine our method across a diverse set of input source meshes and target text prompts. We consider a variety of sources including: COSEG , Thingi10K , Shapenet , Turbo Squid , and ModelNet . Our method requires no particular quality constraints or preprocessing of inputs, and the breadth of shapes we stylize in this paper and in our project webpage illustrates its ability to handle low-quality meshes. Meshes used in the main paper and the project webpage contain an average of 79,366 faces, 16% non-manifold edges, 0.2% non-manifold vertices, and 12% boundaries. Our method takes less than 25 minutes to train on a single GPU, and high quality results usually appear in less than 10 minutes.
In Sec. 4.1, we demonstrate the multiple control mechanisms enabled by our method. In Sec. 4.2, we conduct a series of ablations on the key priors in our method. We further explore the synergy between learning color and geometry in tandem. We introduce a user study in Sec. 4.3 where our stylization is compared to a baseline method. In Sec. 4.4, we show that our method can easily generalize to other target modalities beyond text, such as images or 3D shapes. Finally, we discuss limitations in Sec. 4.6.
Our method generates details with high granularity while still maintaining global semantics and preserving the underlying content. For example in Fig. 2, given a vase mesh and target text ‘colorful crochet’, the stylized output includes knit patterns with different colors, while preserving the structure of the vase. In Fig. 3, our method demonstrates a global semantic understanding of humans. Different body parts such legs, head and muscles are stylized appropriately in accordance with their semantic role, and these styles are blended seamlessly across the surface to form a cohesive texture. Moreover, our neural style field network generates structured textures which are aligned to sharp curves and features (see bricks in Figs. 1 and 2 and in the project webpage). We show in Fig. 6 and in the project webpage that our method styles the entire mesh in a consistent manner that is part-aware and exhibits natural variation in texture.
Fine Grained Controls. Our network leverages a positional encoding where the range of frequencies can be directly controlled by the standard deviation of the matrix Eq. 1. In Fig. 7, we show the results of three different frequency values when stylizing a source mesh of a torus towards the target text ‘stained glass donut’. Increasing the frequency value increases the frequency of style details on the mesh and produces sharper and more frequent displacements along the normal direction. We further demonstrate our method’s ability to successfully synthesize styles of varying levels of specificity. Fig. 8 displays styles of increasing detail and specificity for two input shapes. Note the retention of the style details from each level of target granularity to the next. Though the primary mode of style control is through the text prompt, we explore the way the network adapts to the geometry of the source shape. In Fig. 10, the target text prompt is fixed to ‘cactus’. We consider different input source spheres with increasing protrusion frequency. Observe that both the frequency and structure of the generated style changes to align with the pre-existing structure of the input surface. This shows that our method has the ability to preserve the content of the input mesh without compromising the quality of the stylization.
Meshes with corresponding connectivity can be used to morph between two surfaces . Thus, our ability to modify style while preserving the input mesh enables morphing (see Fig. 9). To morph between meshes, we apply linear interpolation between the style value (RGB and displacement) of every point on the mesh, for each instance of the stylized mesh.
2 Text2Mesh Priors
Our method incorporates a number of priors that allow us to perform stylization without a pre-trained GAN. We show an ablation where each prior is removed in Fig. 5. Removing the style field network (net), and instead directly optimizing the vertex colors and displacements, results in noisy and arbitrary displacements over the surface. In random 2D augmentations are necessary to generate meaningful CLIP-guided drawings. We observe the same phenomena in our method, whereby removing 2D augmentations results in a stylization completely unrelated to the target text prompt. Without Fourier feature encoding (FFN), the generated style loses all fine-grained details. With the cropping augmentation removed (crop), the output is similarly unable to synthesize the fine-grained style details that define the target. Removing the geometry-only component of (displ) hinders geometric refinement, and the network instead compensates by simulating geometry through shading (see also Fig. 11). Without a geometric prior (3D) there is no source mesh to impose global structure, thus, the 2D plane in 3D space is treated as an image canvas. For each result in Fig. 5, we report the CLIP similarity score, , as defined in Sec. 3. Our method obtains the highest score across different ablations, see Fig. 5. Ideally, there is a correlation between visual quality and CLIP scores. However, -3D manages to achieve a high CLIP similarity, despite its zero regard for global content semantics. This shows an example of how CLIP may naively prefer degenerate solutions, while our geometric prior steers our method away from these solutions.
Interplay of Geometry and Color. Our method utilizes the interplay between geometry and color for effective stylization, as shown in Fig. 11. Learning to predict only geometric manipulations produces inferior geometry compared to learning geometry and color together, as the network attempts to simulate shading by generating displacements for self shadowing. An extreme case of this can be seen with the “Batman” in Fig. 3, where the bat symbol on the chest is the result of a deep concavity formed through displacements alone. Similarly learning to predict only color results in the network attempting to hallucinate geometric detail through shading, leading to a flat and unrealistic texture that nonetheless is capable of achieving a relatively high CLIP score when projected to 2D. Fig. 11 illustrates this adversarial solution, where the “Color” mode achieves a similar CLIP score as our “Full” method.
3 Stylization Fidelity
Our method performs the task of general text-driven stylization of meshes. Given that no approaches exist for this task, we evaluate our method’s performance by extending VQGAN-CLIP . This baseline synthesizes color inside a binary 2D mask projected from the 3D source shape (without 3D deformations) guided by CLIP. Further, the baseline is initialized with a rendered view of the 3D source. We conduct a user study to evaluate the perceived quality of the generated outputs, the degree to which they preserve the source content, and how well they match the target style.
We had 57 users evaluate random source meshes and style text prompt combinations. For each combination, we display the target text and the stylized output in pairs. The users are then asked to assign a score (1-5) to three factors: (Q1) “How natural is the output depiction of {content} + {style}?” (Q2) “How well does the output match the original {content}?” (Q3) “How well does the output match the target {style}?” We report the mean opinion scores with standard deviations in parentheses for each factor averaged across all style outputs for our method and the baseline in Tab. 1. We include three control questions where the images and target text do not match, and obtain a mean control score of 1.16. Our method outperforms the VQGAN baseline across all questions, with a difference of , , and for Q1-Q3, respectively. Though VQGAN is somewhat effective at representing the natural content in our prompts, perhaps due to the implicit content signal it receives from the mask, it struggles to synthesize these representations with style in a meaningful way. Examples of our baseline outputs are provided in the Appendix E. Visual examples of generated styles and screenshots of the user study are also discussed in Appendix E.
4 Beyond Textual Stylization
Beyond text-based stylization, our method can be used to stylize a mesh toward different target modalities such as a 2D image or even a 3D object. For a target 2D image , in Eq. 5, represents the image-based CLIP embedding of . For a target mesh , is the average embedding, in CLIP space, of the 2D renderings of , where the views are the same as those sampled for the source mesh. Beyond different modalities, we can combine targets across different modalities by simply summing over each target. In Fig. 12 we consider a source mesh of a pig with different image targets. In Fig. 13(a-b), we consider stylization using a target mesh and in Fig. 13(c-d), we combine both a target mesh and a target text. Our method successfully adheres to the target style.
5 Incorporating Symmetries
We can make use of prior knowledge of the input shape symmetry to enforce style consistency across the axis of symmetry. Such symmetries can be introduced into our model by modifying the input to our positional encoding in Eq. 1. For instance, given a point and a shape with bilateral symmetry across the X-Y plane, one can apply a function prior to the the positional encoding such that . We show the effect of this symmetry prior on a UFO mesh in Fig. 14. This prior is effective even when the triangulation is not perfectly symmetrical, since the function is applied in Euclidean space. A full investigation into incorporating additional symmetries within positional encoding is an interesting direction for future work.
6 Limitations
Our method implicitly assumes there exists a synergy between the input 3D geometry and the target style prompt (see Fig. 15). However, stylizing a 3D mesh (e.g., dragon) towards an unrelated/unnatural prompt (e.g., stained glass) may result in a stylization that ignores the geometric prior and effectively erases the source shape content. Therefore, in order to preserve the original content when editing towards a mismatched target prompt, we simply include the object category in the text prompt (e.g., stained glass dragon) which adds a content preservation constraint into the target.
Conclusion
We present a novel framework for stylizing input meshes given a target text prompt. Our framework learns to predict colors and local geometric details using a neural stylization network. It can predict structured textures (e.g. bricks), without a directional field or mesh parameterization. Traditionally, the direction of texture patterns over 3D surfaces has been guided by 3D shape analysis techniques (as in ). In this work, the texture direction is driven by 2D rendered images, which capture the semantics of how textures appear in the real world.
Without relying on a pre-trained GAN network or a 3D dataset, we are able to manipulate a myriad of meshes to adhere to a wide variety of styles. Our system is capable of generating out-of-domain stylized outputs, e.g., a stained glass shoe or a cactus vase (Fig. 2). Our framework uses a pre-trained CLIP model, which has been shown to contain bias . We postulate that our proposed method can be used to visualize, understand, and interpret such model biases in a more direct and transparent way.
As future work, our framework could be used to manipulate 3D content as well. Instead of modifying a given input mesh, one could learn to generate meshes from scratch driven by a text prompt. Moreover, our NSF is tailored to a single 3D mesh. It may be possible to train a network to stylize a collection of meshes towards a target style in a feed-forward manner.
References
Appendix A Additional Results
Please refer to our project webpage additional results. We show multiple views of a chair mesh synthesized with a wood style in Fig 16 to demonstrate how our textures automatically align to the shape’s sharp features and curves.
Appendix B High Resolution Stylization
We learn to style high-resolution meshes, and thus are able to synthesize style with high fidelity. We show in Fig. 17 a 1670x2720 render of one of the stylized outputs we show in the main paper as a demonstration.
As mentioned in Sec. 3.1, our method is effective even on coarse inputs, and one can always increase the resolution of a mesh M to learn a neural field with finer granularity. In Fig. 18, we upsample the mesh by inserting a degree-3 vertex in the barycenter of each triangle face of the mesh. The network is able to synthesize a finer style by leveraging these additional vertices.
Appendix C Choice of anchor view.
As mentioned in Sec. 3.3 of the main text, we select the view with the highest (i.e. best) CLIP similarity to the content as the anchor. There are often many possible views that can be chosen as the anchor that will allow a high-quality stylization. We show in Fig. 19 a camel mesh where the vertices are colored according to the CLIP score of the view that passes from the vertex to the center of the mesh. The color range is shown where the minimum and maximum values in the range are 0 and 0.4, respectively. We show in Fig. 20 a view with one of the highest CLIP scores and a view with one of the lowest. The CLIP score exhibits a strong positive correlation with views that are semantically meaningful, and thus can be used for automatic anchor view selection, as described in the main paper. This metric is limited in expressiveness, however, as demonstrated by the constrained range that the scores fall within for all the views around the mesh. The highest score for any view of the camel is 0.35 whereas the lowest score is still 0.2.
As mentioned in Sec. 3.3, , the number of sampled views, is set to 5. We show in Fig. 21 that increasing the number of views beyond 5 does little to change the quality of the output stylization.
Appendix D Training and Implementation Details
D.2 Training
We use the Adam optimizer with an initial learning rate of , and decay the learning rate by a factor of every iterations. We train for 1500 iterations on a single Nvidia GeForce RTX2080Ti GPU, which takes around 25 minutes to complete. For augmentations , we use a random perspective transformation. For we randomly crop the image to 10% of its original size and then apply a random perspective transformation. Before encoding images with CLIP, we normalize per-channel by mean and standard deviation .
Appendix E Baseline Comparison and User Study
Examples of results for our VQGAN baseline, as described in Sec. 4 are shown in Fig. 22 and Fig. 23, alongside our results.
In addition, in Fig. 24, we provide screenshots of our user study, as shown for users.
Appendix F Societal Impact
Our framework utilizes a pre-trained CLIP embedding space, which has been shown to contain bias . Since our system is capable of synthesizing a style driven by a target text prompt, it enables visualizing such biases in a direct and transparent way. We’ve observed evidence of societal bias in some of our stylizations. For example, the nurse style in Fig. 25 is biased towards adding female features to the input male shape. Our method offers one of the first opportunities to directly observe the biases present in joint image-text embeddings through our stylization framework. An important future work may leverage our proposed system in helping create a datasheet for CLIP in addition to future image-text embedding models.