TEXTure: Text-Guided Texturing of 3D Shapes

Elad Richardson, Gal Metzer, Yuval Alaluf, Raja Giryes, Daniel Cohen-Or

Introduction

The ability to paint pictures with words has long been a sign of a master storyteller, and with recent advancements in text-to-image models, this has become a reality for us all. Given a textual description, these new models are able to generate highly detailed imagery that captures the essence and intent of the input text. Despite the rapid progress in text-to-image generation, painting 3D objects remains a significant challenge as it requires to consider the specific shape of the surface being painted. Recent works have begun making significant progress in painting and texturing 3D objects by using language-image models as guidance . Yet, these methods still fall short in terms of quality compared to their 2D counterparts.

In this paper, we focus on texturing 3D objects, and present TEXTure, a technique that leverages diffusion models to seamlessly paint a given 3D input mesh. Unlike previous texturing approaches that apply score distillation to indirectly utilize Stable Diffusion as a texturing prior, we opt to directly apply a full denoising process on rendered images using a depth-conditioned diffusion model .

At its core, our method iteratively renders the object from different viewpoints, applies a depth-based painting scheme, and projects it back to the mesh vertices or atlas. We show that our approach can result in a significant boost in both running time and generation quality. However, applying this process naïvely would result in highly inconsistent texturing with noticeable seams due to the stochastic nature of the generation process (see Figure 2 (A)).

To alleviate these inconsistencies, we introduce a dynamic partitioning of the rendered view to a trimap of “keep”, “refine”, and “generate” regions, which is estimated before each diffusion process. The “generate” regions are areas in the rendered viewpoint that are viewed for the first time and need to be painted. A “refine” region is an area that was already painted in previous iterations, but is now seen from a better angle and should be repainted. Finally, “keep” regions are painted regions that should not be repainted from the current view. We then propose a modified diffusion process that takes into account our trimap partitioning. By freezing “keep” regions during the diffusion process we attain more consistent outputs, but the newly generated regions still lack global consistency (see Figure 2 (B)). To encourage better global consistency in the “generate” regions, we further propose to incorporate both depth-guided and mask-guided diffusion models into the sampling process (see Figure 2 (C)). Finally, for “refine” regions, we design a novel process that repaints these regions but takes into account their existing texture. Together these techniques allow the generation of highly-realistic results in mere minutes (see Figure 2 (D) and Figure 1 for results).

Next, we show that our method can be used not only to texture meshes guided by a text prompt, but also based on an existing textures from some other colored mesh or even from a small set of images. Our method requires no surface-to-surface mapping or any intermediate reconstruction step. Instead, we propose to learn semantic tokens that represent a specific texture by building on Textual Inversion and DreamBooth , while extending them to depth-conditioned models and introducing learned viewpoint tokens. We show that we can successfully capture the essence of a texture even from a few unaligned images and use it to paint a 3D mesh based on its semantic texture.

Finally, in the spirit of diffusion-based image editing , we show that one can further refine and edit textures. We propose two editing techniques. First, we present a text-only refinement where an existing texture map is modified using a guiding prompt to better match the semantics of the new text. Second, we illustrate how users can directly apply an edit on a texture map, where we refine the texture to fuse the user-applied edits into the 3D shape.

We evaluate TEXTure and show its effectiveness for texture generation, transfer, and editing. We demonstrate that TEXTure offers a significant speedup compared to previous approaches, and more importantly, offers significantly higher-quality generated textures.

Related Work

The past year has seen the development of multiple large diffusion models capable of producing impressive images with pristine details guided by an input text prompt. The widely popular Stable Diffusion , is trained on a rich text-image dataset and is conditioned on CLIP’s frozen text encoder. Beyond simple text-conditioning, Stable Diffusion has multiple extensions that allow conditioning its denoising network on additional input modalities such as a depth map or an inpainting mask. Given a guiding prompt and an estimated depth image , the depth-conditioning model is tasked with generating images that follow the same depth values while being semantically faithful with respect to the text. Similarly, the inpainting model completes missing image regions, given a masked image.

Although current text-to-image models generate high-quality results when conditioned on a text prompt or depth map, editing an existing image or injecting objects specified by a few exemplar images remains challenging . To introduce a user-specific concept to a pre-trained text-to-image model, introduce Textual Inversion to map a few exemplar images into a learned pseudo-tokens in the embedding space of the frozen text-to-image model. DreamBooth further fine-tune the entire diffusion model on the set of input images to achieve more faithful compositions. The learned token or fine-tuned model can then be used to generate novel images using the custom token in new user-specified text prompts.

Texture and Content Transfer.

Early works focus on 2D texture synthesis through probabilistic models, while more recent works take a data-driven approach to generate textures using deep neural networks. Generating textures over 3D surfaces is a more challenging problem, as it requires attention to both colors and geometry. For geometric texture synthesis, applies a similar statistical method to while extended non-parametric sampling proposed by to 3D meshes. introduces a metric learning approach to transfer details from a source to a target shape, while use an internal learning technique to transfer geometric texture. For 3D color texture synthesis, given an exemplar colored mesh, use the relation between geometric features and color values to synthesize new textures on target shapes.

D Shape and Texture Generation.

Generating shapes and textures in 3D has recently gained significant interest. Text2Mesh , Tango , and CLIP-Mesh use CLIP-space similarities as an optimization objective to generate novel shapes and textures. CLIP-Mesh deforms an initial sphere with UV texture parameterization. Tango optimizes per-vertex colors and focuses on generating novel textures. Text2Mesh optimizes per-vertex color attributes, while allowing small geometric displacements. Get3D is trained to generate shape and texture through a DMTet mesh extractor and 2D adversarial losses.

Recently, DreamFusion introduced the use of pre-trained image diffusion models for generating 3D NeRF models conditioned on a text prompt. The key component in DreamFusion is the Score-Distillation loss which enables the use of a pretrained 2D diffusion model as a critique for optimizing the 3D NeRF scene. Recently, Latent-NeRF showed how the same Score-Distillation loss can be used in Stable Diffusion’s latent space to generate latent 3D NeRF models. In the context for texture generation, present Latent-Paint, a texture generation technique, where latent texture maps are painted using Score-Distillation and are then decoded to RGB for the final colorization output. Similarly, uses score-distillation to texture and refine a coarse initial shape. Both methods suffer from relatively slow convergence and less defined textures compared to our proposed approach.

Method

We first lay the foundation for our text-guided mesh texturing scheme, illustrated in Figure 3. Our TEXTure scheme performs an incremental texturing of a given 3D mesh, where at each iteration we paint the currently visible regions of the mesh as seen from a single viewpoint. To encourage both local and global consistency, we segment the mesh into a trimap of “keep”, “refine”, “generate” regions. A modified depth-to-image diffusion process is presented to incorporate this information into the denoising steps.

We then propose two extensions of TEXTure. First, we present a texture transfer scheme (Section 3.2) that transfers the texture of a given mesh to a new mesh, by learning a custom concept that represents the given texture. Finally, we present a texture editing technique that allows users to edit a given texture map, either through a guiding text prompt or a user-provided scribble (Section 3.3).

Our texture generation method relies on a pretrained depth-to-image diffusion model Mdepth\mathcal{M}_{depth} and a pretrained inpainting diffusion model Mpaint\mathcal{M}_{paint}, both based on Stable Diffusion and with a shared latent space. During the generation process, the texture is represented as an atlas through a UV mapping that is calculated using XAtlas .

We start from an arbitrary initial viewpoint v0=(r=1.25,ϕ0=0,θ=60)v_{0}=(r=1.25,\phi_{0}=0,\theta=60) where rr is the radius of the camera, ϕ\phi is the azimuth camera angle, and θ\theta is the camera elevation. We then use Mdepth\mathcal{M}_{depth} to generate an initial colored image I0I_{0} of the mesh as viewed from v0v_{0}, conditioned on the rendered depth map D0\mathcal{D}_{0}. The generated image I0I_{0} is then projected back to the texture atlas T0\mathcal{T}_{0} to color the shape’s visible parts from v0v_{0}. Following this initialization step, we begin a process of incremental colorization, illustrated in Figure 3, where we iterate through a fixed set of viewpoints. For each viewpoint, we then render the mesh using a renderer R\mathcal{R} to obtain Dt\mathcal{D}_{t} and QtQ_{t}, where QtQ_{t} is the rendering of the mesh as seen from the viewpoint vtv_{t} that considers all previous colorization steps. Finally, we generate the next image ItI_{t} and project ItI_{t} back to the updated texture atlas Tt\mathcal{T}_{t} while taking into account QtQ_{t}.

Once a single view has been painted, the generation task becomes more challenging due to the need for local and global consistency along the generated texture. Below we consider a single iteration tt of our incremental painting process and elaborate on our proposed techniques to handle these challenges.

Given a viewpoint vtv_{t}, we first apply a partitioning of the rendered image into three regions: “keep”, “refine”, and “generate”. The “generate” regions are rendered areas that are viewed for the first time and need to be painted to match the previously painted regions. The distinction between “keep” and “refine” regions is slightly more nuanced and is based on the fact that coloring a mesh from an oblique angle can result in high distortion. This is because the cross-section of a triangle with the screen is low, resulting in a low-resolution update to the mesh texture image Tt\mathcal{T}_{t}. Specifically, we measure the triangle’s cross-section as the zz component of the face normal nzn_{z} in the camera’s coordinate system.

Ideally, if the current view provides a better colorization angle for some of the previously-painted regions, we would like to “refine” their existing texture. Otherwise, we should “keep” the original texture and avoid modifying it to ensure consistency with previous views. To keep track of seen regions and the cross-section at which they were previously colored from, we use an additional meta-texture map N\mathcal{N} that is updated at every iteration. This additional map can be efficiently rendered together with the texture map at each iteration and is used to define the current trimap partitioning.

Masked Generation.

As the depth-to-image diffusion process was trained to generate an entire image, we must modify the sampling process to “keep” part of the image fixed. Following Blended Diffusion , in each denoising step we explicitly inject a noised versions of QtQ_{t}, i.e. zQtz_{Q_{t}}, at the “keep” regions into the diffusion sampling process, such that these areas are seamlessly blended into the final generated result. Specifically, the latent at the current sampling timestep ii is computed as

where the mask mblendedm_{blended} is defined in Equation 2. That is, for “keep” regions, we simply set ziz_{i} fixed according to their original values.

Consistent Texture Generation.

Injecting “keep” regions into the diffusion process results in better blending with “generate” regions. Still, when moving away from the “keep” boundary and deeper into the “generate” regions, the generated output is mostly governed by the sampled noise and is not consistent with previously painted regions. We first opt to use the same sampled noise from each viewpoint, this sometimes improves consistency, but is still very sensitive to the change in viewpoint. We observe that applying an inpainting diffusion model Mpaint\mathcal{M}_{paint} that was directly trained to complete masked regions, results in more consistent generations. However, this in turn deviates from the conditioning depth Dt\mathcal{D}_{t} and may generate new geometries. To benefit from the advantages of both models we introduce an interleaved process where we alternate between the two models during the initial sampling steps. Specifically, during sampling, the next noised latent zi−1z_{i-1} is computed as:

When applying Mdepth\mathcal{M}_{depth}, the noised latent is guided by the current depth Dt\mathcal{D}_{t} while when applying Mpaint\mathcal{M}_{paint}, the sampling process is tasked with completing the “generate” regions in a globally-consistent manner.

Refining Regions.

To handle “refine” regions we use another novel modification to the diffusion process that generates new textures while taking into account their previous values. Our key observation is that by using an alternating checkerboard-like mask in the first steps of the sampling process, we can guide the noise towards values that locally align with previous completions.

The granularity of this process can be controlled by changing the resolution of the checkerboard mask and the number of constrained steps. In practice, we apply the mask for the first 2525 sampling steps. Namely, the masked mblendedm_{blended} applied in Equation 1 is set as,

where a value of 11 indicates that this region should be painted and kept otherwise. Our blending mask is visualized in Figure 3.

Texture Projection.

To project ItI_{t} back to the texture atlas Tt\mathcal{T}_{t}, we apply gradient-based optimization for Lt\mathcal{L}_{t} over the values of Tt\mathcal{T}_{t} when rendered through the differential renderer R\mathcal{R}. That is,

To achieve smoother texture seams of the projections from different views, a soft mask msm_{s} is applied at the boundaries of the “refine” and “generate” region:

Additional Details.

Our texture is represented as a 1024×10241024\times 1024 atlas, where the rendering resolution is 1200×12001200\times 1200. For the diffusion process, we segment the inner region, resize it to 512×512512\times 512 and mat it onto a realistic background. All shapes are rendered with 88 viewpoints around the object, and two additional top/bottom views. We show that viewpoint order can also affect the end results.

2 Texture Transfer

Having successfully generated a new texture on a given 3D mesh, we now turn to describe how to transfer a given texture to a new, untextured target mesh. We show how to capture textures from either a painted mesh, or from a small set of input images. Our texture transfer approach builds on previous work on concept learning over diffusion models , by fine-tuning a pretrained diffusion model and learning a pseudo-token representing the generated texture. The fine-tuned model is then used for texturing a new geometry. To improve the generalization of the fine-tuned model to new geometries, we further propose a novel spectral augmentation technique, described next. We then discuss our concept learning scheme from meshes or images.

Since we are interested in learning a token representing the input texture and not the original input geometry itself, we should ideally learn a common token over a range of geometries containing the input texture. Doing so disentangles the texture from its specific geometry and improves the generalization of the fine-tuned diffusion model. Inspired by the concept of surface caricaturization , we propose a novel spectral augmentation technique. In our case, we propose random low-frequency geometric deformations to the textured source mesh, regularized by the mesh Laplacian’s spectrum .

Modulating random deformations over the spectral eigenbasis results in smooth deformations that keep the integrity of the input shape. Empirically, we choose to apply random inflations or deflations to the mesh, with a magnitude proportional to a randomly selected eigenfunction. We provide examples of such augmentations in Figure 4, and additional details in the supplementary materials.

Texture Learning.

Applying our spectral augmentation technique, we generate a large set of images with corresponding depth maps of the input shape. We render the images from several viewpoints (left, right, overhead, bottom, front, and back) and paste the rendered object onto a randomly colored background (see Figure 4).

Given the set of rendered images, we follow and optimize an embedding vector representing our texture using prompts of the form “a ⟨Dv⟩\langle D_{v}\rangle photo of a ⟨Stexture⟩\langle\mathcal{S}_{texture}\rangle” where ⟨Dv⟩\langle D_{v}\rangle is a learned token representing the view direction of the rendered image and ⟨Stexture⟩\langle\mathcal{S}_{texture}\rangle is the token representing our texture. Observe, that we have six learned directional tokens DvD_{v}, shared within images from the same view, and a single token Stexture\mathcal{S}_{texture} representing the texture, shared across all images. Additionally, to better capture the input texture we fine-tune the diffusion model itself as well, as done in . Our texture learning scheme is illustrated in Figure 4. After training, we use TEXTure (Section 3.1), to color the target shape, swapping the original Stable Diffusion model with our fine-tuned model.

Texture from Images.

Next, we explore the more challenging task of texture generation based on a small set of sample images. While we cannot expect the same quality given only a few images, we can still potentially learn concepts that represent different textures. Unlike standard textual inversion techniques our learned concepts represent mostly texture and not structure as they are trained on a depth-conditioned model. This potentially makes them more suitable for texturing other 3D shapes.

For this task, we segment the prominent object from the image using a pretrained saliency network , apply standard scale and crop augmentations, and paste the result onto a randomly-colored background. Our results show that one can successfully learn semantic concepts from images and apply them to 3D shapes without any explicit reconstruction stage in between. We believe this creates new opportunities for creating captivating textures inspired by real objects.

3 Texture-Editing

We show that our trimap-based TEXTuring can be used to easily extend 2D editing techniques to a full mesh. For text-based editing, we wish to alter an existing texture map guided by a textual prompt. To this end, we define the entire texture map as a “refine” region and apply our TEXTuring process to modify the texture to align with the new text prompt. We additionally provide scribble-based editing where a user can directly edit a given texture map (e.g. to define a new color scheme over a desired region). To allow this, we simply define the altered regions as “refine” regions during the TEXTuring process and “keep” the remaining texture fixed.

Experiments

We now turn to validate the robustness and effectiveness on our proposed method through a set of experiments.

We first demonstrate results achieved with TEXTure across several geometries and driving text prompts. Figure 5, Figure 1 and Figure 11 show highly-detailed realistic textures that are conditioned on a single text prompt. Observe, for example, how the generated textures nicely align with the geometry of the turtle and elephant shapes. Moreover, the generated textures are consistent both on a local scale (e.g., along the shell of the turtle) and a global scale (e.g., the shell is consistent across different views). Furthermore, TEXTure can successfully generate textures of an individual (e.g., of Albert Einstein in Figure 1). Observe how the generated texture captures fine details of Albert Einstein’s face. Finally, TEXTure can successfully handle challenging geometric shapes, including the non-orientable Klein bottle, shown in Figure 5.

Qualitative Comparisons.

In Figure 6 we compare our method to several state-of-the-art methods for 3D texturing. First, we observe that Clip-Mesh and Text2Mesh struggle in achieving globally-consistent results due to their heavy reliance CLIP-based guidance. See the ”90’s boombox” example at the top of Figure 6 where speakers are placed sporadically in competing methods. For Latent-Paint , the results are more plausible but still lack in terms of quality. For example, Latent-Paint often struggles in achieving visibly sharp textures such as those shown on “’a desktop iMac’. We attribute this shortcoming due to its reliance on score distillation , which tends to omit high-frequency details.

The last row in Figure 6, depicting a statue of Napoleon Bonaparte, showcases our method’s ability to produce fine details compared to alternative methods that struggle to produce matching quality. Notably, our results were achieved significantly faster than the alternative methods. Generating a single texture with TEXTure takes approximately 55 minutes compared to 1919 through 4545 minutes for alternative methods (See Table 1). We provide additional qualitative results in 11 and in the supplementary materials.

User Study.

Finally, we conduct a user study to analyze the fidelity and overall quality of the generated textures. We select 1010 text prompts and corresponding 3D meshes and texture the meshes using TEXTure and two baselines: Text2Mesh and Latent-Paint . For each prompt and method, we ask each respondent to evaluate the result with respect to two aspects: (1) its overall quality and (2) the level at which it reflects the text prompt, on a scale of 11 to 55. Results are presented in Table 1 where we show the average results across all prompts for each method. As can be seen, TEXTure outperforms both baselines in terms of both overall quality and text fidelity by a significant margin. Importantly, our method’s improved quality and fidelity are attained with a significant decrease in runtime. Specifically, TEXTure achieves a decrease of 6.4×6.4\times in running time compared to Text2Mesh and a decrease of 9.2×9.2\times relative to Latent-Paint.

In addition to the above evaluation setting, we ask respondents to rank the methods relative to each other. Specifically, for each of the 1010 prompts, we show the results of all methods side-by-side (in a random order) and ask respondents to rank the results. In Table 2 we present the average rank of each method, averaged across the 1010 prompts and across all responses. Note that a lower rank is preferable in this setting. As can be seen, TEXTure has a significantly better average rank relative to the two baselines, as desired. This further demonstrates the effectiveness of TEXTure in generating high-quality, semantically-accurate textures.

Ablation Study.

An ablation validating the different components of our TEXTure scheme is shown in Figure 7. One can see that each component is needed for achieving high-quality generations and improving our method’s robustness to the sensitivity of the generation process. Specifically, without differentiating between “keep” and “generate” regions the generated textures are unsatisfactory. For both the teddy bear and the sports car one can clearly see that the texture presented in A fails to achieve local and global consistency, with inconsistent patches visible across the texture. In contrast, thanks to our blending technique for “keep” regions, B achieves local consistency between views. Still, observe that the texture of the teddy bear legs in the top example, do not match the texture of its back. By incorporating our improved “generate” method, we achieve more consistent results across the entire shape, as shown in C. These results tend to have smeared regions as can be observed in the fur of the teddy bear, or the text written across the hood of the car. We attribute this to the fact that some regions are painted from oblique angles at early viewpoints, which are not ideal for texturing, and are not refined. By identifying “refine” regions and applying our full TEXTure scheme, we are able to effectively address these problems and produce sharper textures, see D.

2 Texture Capturing

We next validate our proposed texture capture and transfer technique that can be applied over both 3D meshes and images.

As mentioned in Section 3.2 we are able to capture a texture of a given mesh by combining concept learning techniques with view-specific directional tokens. We can then use the learned token with TEXTure to color new meshes accordingly. Figure 8 presents texture transfer results from two existing meshes onto various target geometries. While the training of the new texture token is performed using prompts of the form “A ⟨Dv⟩\langle D_{v}\rangle photo of ⟨Stexture⟩\langle\mathcal{S}_{texture}\rangle”, we can generate new textures by forming different texture prompts when transferring the learned texture. Specifically, the ”Exact” results in Figure 8 were generated using the specific text prompt used for fine-tuning, and the rest of the outputs were generated ”in the style of ⟨Stexture⟩\langle\mathcal{S}_{texture}\rangle”. Observe how a single eye is placed on Einstein’s face when the exact prompt is used, while two eyes are generated when only using a prompt ”a ⟨Dv⟩\langle D_{v}\rangle photo of Einstein that looks like ⟨Stexture⟩\langle\mathcal{S}_{texture}\rangle”. A key component of the texture-capturing technique is our novel spectral augmentations scheme. We refer the reader to the supplementary materials for an ablation study over this component.

Texture From Image

In practice, it is often more practical to learn a texture for a set of images rather than from a 3D mesh. As such, in Figure 9, we demonstrate the results of our transferring scheme where the texture is captured from a collection of images depicting the source object. One can see that even when given only several images, our method can successfully texture different shapes with a semantically similar texture. We find these results exciting, as it means one can easily paint different 3D shapes using textures derived from real-world data, even when the texture itself is only implicitly represented in the fine-tuned model and concept tokens.

3 Editing

Finally, in Figure 15 we show the editing results of existing textures. The top row of Figure 15 shows scribble-based results where a user manually edits the texture atlas image. Then we “refine” the texture atlas to seamlessly blend the edited region into the final texture. Observe how the manually annotated white spot on the bunny on the right turns into a realistic-looking patch of white fur.

The bottom of Figure 15 shows text-based editing results, where an existing texture is “refined” according to a new text prompt. The bottom right example illustrates a texture generated on a similar geometry with the same prompt from scratch. Observe that the texture generated from scratch significantly differs from the original texture. In contrast, when applying editing, the textures are able to remain semantically close to the input texture while generating novel details to match the target text.

Discussion, Limitations and Conclusions

This paper presents TEXTure, a novel method for text-guided generation, transfer, and editing of textures for 3D shapes. There have been many models for generating and editing high-quality images using diffusion models. However, leveraging these image-based models for generating seamless textures on 3D models is a challenging task, in particular for non-stochastic and non-stationary textures. Our work addresses this challenge by introducing an iterative painting scheme that leverages a pretrained depth-to-image diffusion model. Instead of using a computationally demanding score distillation approach, we propose a modified image-to-image diffusion process that is applied from a small set of viewpoints. This results in a fast process capable of generating high-quality textures in mere minutes. With all its benefits, there are still some limitations to our proposed scheme, which we discuss next.

While our painting technique is designed to be spatially coherent, it may sometimes result in inconsistencies on a global scale, caused by occluded information from other views. See Figure 10, where different looking eyes are added from different viewpoints. Another caveat is viewpoint selection. We use eight fixed viewpoints around the object, which may not fully cover adversarial geometries. This issue can possibly be solved by finding a dynamic set of viewpoints that maximize the coverage of the given mesh. Furthermore, the depth-guided model sometimes deviates from the input depth, and may generate images that are not consistent with the geometry (See Figure 10 left). This in turn may result in conflicting projections to the mesh, that cannot be fixed in later painting iterations.

With that being said, we believe that TEXTure takes an important step toward revolutionizing the field of graphic design and further opens new possibilities for 3D artists, game developers, and modelers who can use these tools to generate high-quality textures in a fraction of the time of existing techniques. Additionally, our trimap partitioning formulation provides a practical and useful framework that we hope will be utilized and “refined” in future studies.

Acknowledgements

We thank Dana Cohen and Or Patashnik for their early feedback and insightful comments. We would also like to thank Harel Richardson for generously allowing us to use his toys throughout this paper. The beautiful meshes throughout this paper are taken from .

References