AvatarCLIP: Zero-Shot Text-Driven Generation and Animation of 3D Avatars

Fangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai, Lei Yang, Ziwei Liu

Introduction

Creating digital avatars has become an important part in movie, game, and fashion industries. The whole process includes creating shapes of the character, drawing textures, rigging skeletons, and driving the avatar with the captured motions. Each step of the process requires many specialists that are familiar with professional software, large numbers of working hours, and expensive equipment, which are only affordable by large companies. The good news is that recent progress in academia, such as large-scale pre-trained models (Radford et al., 2021) and advanced human representations (Loper et al., 2015) are making it possible for this complicated work to be available to small studios and even reach out to the mass crowd. In this work, we take a step further and propose AvatarCLIP, which is capable of generating and animating 3D avatars solely from natural language descriptions as shown in Fig. 1.

Previously, there are several similar efforts, e.g. avatar generation, and motion synthesis, towards this vision. For avatar generation, several attempts have been made in 2D-GAN-based human generation (Sarkar et al., 2021b; Lewis et al., 2021) and 3D neural human generation (Grigorev et al., 2021; Zhang et al., 2021). However, most of them either compromise the generation quality and diversity or have limited control over the generation process. Not to mention, most generated avatars cannot be easily animated. As for motion synthesis, impressive progresses have been made in recent years. They are able to generate motions conditioned on action classes (Petrovich et al., 2021; Guo et al., 2020) or motion trajectories (Kania et al., 2021; Ling et al., 2020). However, they typically require paired data for fully supervised training, which not only limits the richness of the generated motions but also makes the generation process less flexible. There still exists a considerable gap between existing works and the vision of making digital avatar creation manageable for the mass crowd.

To skirt the complicated operations, natural languages could be used as a user-friendly control signal for convenient 3D avatars generation and animation. However, there are no existing high-quality avatar-text datasets to support supervised text-driven 3D avatar generation. As for animating 3D avatars, a few attempts (Ghosh et al., 2021; Tevet et al., 2022) have been made towards text-driven motion generation by leveraging a motion-text dataset. Nonetheless, restricted by the scarcity of paired motion-text data, those fully supervised methods have limited generalizability.

Fortunately, recent advances in vision-language models pave the way toward zero-shot text-driven generation. CLIP (Radford et al., 2021) is a vision-language pre-trained model trained with large-scale image-text pairs. Through direct supervision on images, CLIP shows great success in zero-shot text-guided image generation (Ramesh et al., 2021). Inspired by this thread of works, we choose to take advantage of the powerful CLIP to achieve zero-shot text-driven generation and animation of 3D avatars. However, neither the 3D avatar nor motion sequences can be directly supervised by CLIP. Major challenges exist in both static avatar generation and motion synthesis.

For static 3D avatar generation, the challenges lie in three aspects, namely texture generation, geometry generation, and the ability to be animated. Inspired by the recent advances in neural rendering (Tewari et al., 2021), the CLIP supervision can be applied to the rendered images to guide the generation of an implicit 3D avatar, which facilitates the generation of avatar textures. Moreover, to speed up the optimization, we propose to initialize the implicit function based on a template human mesh. At optimization time, we also use the template mesh constraint to control the overall shape of the implicit 3D avatar. To adapt to the modern graphics pipeline and, more importantly, to be able to be animated later, we need to extract meshes from generated implicit representations. Therefore, other than the texture, it is also desirable to generate high-quality geometry. To tackle the problem, the key insight is that when inspecting the geometry of 3D models on the computer, users usually turn off texture shading to get texture-less renderings, which can directly reveal the geometry. Therefore, we propose to randomly cast light on the surface of the implicit 3D avatar to get texture-less renderings, upon which the CLIP supervision is applied. Last but not least, to make the generated implicit 3D avatar animatable, we propose to leverage the recent achievements in parametric human models (Loper et al., 2015). Specifically, we align and register the generated 3D avatar to a SMPL mesh. So that it can be driven by the SMPL skeletons.

CLIP is only trained with static images and insensitive to sequential motions. Therefore, it is inherently challenging to generate reasonable motion sequences using only the supervision from CLIP. To tackle this problem, we divide the whole process into two stages: 1) generating candidate poses with the guidance of CLIP, and 2) synthesizing smooth and valid motions with the candidate poses as references. In the first stage, a code-book consisting of diverse poses is created by clustering. Poses that match motion descriptions are selected by CLIP from the code-book. These generated poses serve as important clues for the second stage. A motion VAE is utilized in the second stage to learn motion priors, which facilitates the reference-guided motion synthesis.

With careful designs of the whole pipeline, AvatarCLIP is capable of generating high-quality 3D avatars with reasonable motion sequences guided by input texts as shown in Fig. 1. To evaluate our framework quantitatively, comprehensive user studies are conducted in terms of both avatar generation and animation to show our superiority over existing solutions. Moreover, qualitative experiments are also performed to validate the effectiveness of each component in our framework.

To sum up, our contributions are listed below: 1) To the best of our knowledge, it is the first text-driven full avatar synthesis pipeline that includes the generation of shape, texture, and motion. 2) Incorporating the power of large-scale pre-trained models, the proposed framework demonstrates strong zero-shot generation ability. Our avatar generation pipeline is capable of robustly generating animation-ready 3D avatars with high-quality texture and geometry. 3) Benefiting from the motion VAE, a novel zero-shot text-guided reference-based motion synthesis approach is proposed. 4) Extensive qualitative and quantitative experiments show that the generated avatars and motions are of higher quality compared to existing methods and are highly consistent with the corresponding input natural languages.

Related Work

For its wide application in industries, human modeling has been thoroughly studied for decades. Driven by a large-scale human body dataset (Pishchulin et al., 2017), SMPL (Loper et al., 2015) and SMPL-X (Pavlakos et al., 2019) are proposed as a parametric human model. For its strong interpretability and compatibility with the modern graphics pipeline, we choose SMPL as the template of our avatar. However, they only provide the ability to model the human body without clothes. Many efforts (Bhatnagar et al., 2020; Jiang et al., 2020; Bhatnagar et al., 2019; Hong et al., 2021) have been made on the modeling of clothed humans.

Due to the complexity of clothes and accessories, non-parametric human modelings (Corona et al., 2021; Palafox et al., 2021; Saito et al., 2021; Burov et al., 2021; Mihajlovic et al., 2021; Habermann et al., 2021; Huang et al., 2020; Weng et al., 2019; Alldieck et al., 2018) are proposed to offer more flexibility on realistic human modeling. Moreover, inspired by recent advances in volume rendering, non-parametric human modeling methods based on the neural radiance field (i.e. NeRF (Mildenhall et al., 2020; Jain et al., 2021b)) have also been studied (Peng et al., 2021b; Zhao et al., 2021; Liu et al., 2021; Peng et al., 2021a; Habermann et al., 2021; Xu et al., 2021a). Combining advantages of volume rendering and SDF, NeuS (Wang et al., 2021b) is proposed recently to achieve high-quality geometry and color reconstruction. That justifies our choice of NeuS as base representations of avatars.

With the rapid development of deep learning in recent years, impressive progresses have been shown on 2D image generation. 2D face generation (Karras et al., 2019; Karras et al., 2020; Jiang et al., 2021; Fu et al., 2022) is now a very mature technology for the simplicity of the face structure and large-scale high-quality face datasets (Liu et al., 2015). Recent works (Han et al., 2018; Sarkar et al., 2021b; Lewis et al., 2021; Jiang et al., 2022) have also demonstrated wonderful results in terms of 2D human body generation and manipulation. Without the knowledge of 3D space, it has always been a difficult task to animate the 2D human body (Siarohin et al., 2019a, b; Yoon et al., 2021; Sarkar et al., 2021a).

The 3D human generation (Chen et al., 2022; Noguchi et al., 2021, 2022) is barely explored until the very recent. Combining the powerful SMPL-X and StyleGAN, StylePeople (Grigorev et al., 2021) proposes a data-driven animatable 3D avatar generation method. Inspired by the advancements in NeRF-GAN (Chan et al., 2021), a recent work (Zhang et al., 2021) proposes to bring in 3D awareness to traditional 2D generation pipelines.

Serving as one of the most significant parts of animation, motion synthesis has been the research focus of many researchers. Several large-scale datasets (Ionescu et al., 2013; von Marcard et al., 2018; Mehta et al., 2017; Mahmood et al., 2019; Varol et al., 2017; Cai et al., 2021, 2022) provide human motions as sequences of 3D keypoints or SMPL parameters. The rapid advancements of motion datasets stimulate researches on the motion synthesis. Early works apply classical machine learning to unconditional motion synthesis (Rose et al., 1998; Ikemoto et al., 2009; Mukai and Kuriyama, 2005). DVGANs (Xiao Lin, 2014), Text2Action (Ahn et al., 2018) and Language2Pose (Ahuja and Morency, 2019) generate motions conditioned on short texts using fully annotated data. Action2Motion (Guo et al., 2020) and Actor (Petrovich et al., 2021) condition the motion generation on pre-selected action classes. These methods require large amounts of data(Hong et al., 2022) with annotations of action classes or language descriptions, which limits their applications. On the contrary, our proposed AvatarCLIP can drive the human model by general natural languages without any paired data. Some other works (Li et al., 2021; Aggarwal and Parikh, 2021) focus on music-conditioned motion synthesis. Moreover, some works (Bergamin et al., 2019; Won and Lee, 2019) focus on the physics-based motion synthesis. They construct motion sequences with physical constrains for more realistic generation results.

Text-to-image synthesis (Mansimov et al., 2015) has long been studied. The ability to zero-shot generalize to unseen categories is first shown by (Reed et al., 2016). CLIP and DALL-E (Ramesh et al., 2021) further show the incredible text-to-image synthesis ability by excessively scale-up the size of training data. Benefiting from the zero-shot ability of CLIP, many amazing zero-shot text-driven applications (Frans et al., 2021; Xu et al., 2021b; Patashnik et al., 2021) are being developed. Combining CLIP with 3D representations like NeRF or mesh, zero-shot text-driven 3D object generation (Michel et al., 2021; Jain et al., 2021a; Sanghi et al., 2021; Jetchev, 2021) and manipulation (Wang et al., 2021a) have also come true in recent months.

Our Approach

Recall that our goal is zero-shot text-driven 3D avatar generation and animation, which can be formally defined as follows. The inputs are natural languages, text={tshape,tapp,tmotion}\textrm{text}=\{t_{\textrm{shape}},t_{\textrm{app}},t_{\textrm{motion}}\}. The three texts correspond to the descriptions of the desired body shape, appearance and motion. The output is two-part, including a) an animatable 3D avatar represented as a mesh M={V,F,C}M=\{V,F,C\}, where VV is the vertices, FF stands for faces, CC represents the vertex colors; b) a sequence of poses Θ={θi}i=1L\Theta=\{\theta_{i}\}_{i=1}^{L} comprising the desired motion, where LL is the length of the sequence.

(Radford et al., 2021) is a vision-language pre-trained model trained with large-scale image-text datasets. It consists of an image encoder EIE_{I} and a text encoder ETE_{T}. The encoders are trained in the way that the latent codes of paired images and texts are pulled together and unpaired ones are pushed apart. The resulting joint latent space of images and texts enables the zero-shot text-driven generation by encouraging the latent code of the generated content to align with the latent code of the text description. Formally, the CLIP-guided loss function is defined as

where (⋅)(\cdot) represents the cosine distance. By minimizing Lclip(I,T)\mathcal{L}_{\textrm{clip}}(I,T), the generated image II is encouraged to match the description of the text prompt TT.

(Wang et al., 2021b) proposes a novel volume rendering method that combines the advantages of SDF and NeRF (Mildenhall et al., 2020), leading to high-quality surface reconstruction as well as photo-realistic novel view rendering. For some viewing point oo and viewing direction vv, the color of the corresponding pixel is accumulated along the ray by

where p(t)=o+vtp(t)=o+vt is some point along the ray, c(p(t),v)c(p(t),v) output the color at point p(t)p(t), which is implemented by MLPs. w(t)w(t) is a weighting function for the point p(t)p(t). w(t)w(t) is designed to be unbiased and occlusion-aware, so that the optimization on multi-view images can lead to the accurate learning of a SDF representation. It is defined as

where ϕs\phi_{s} is the logistic density distribution, ff is an SDF network. NeuS is used in our work as the representation of the implicit 3D avatar.

2. Pipeline Overview

As shown in Fig. 2, the pipeline of our AvatarCLIP can be divided to two parts, i.e. static avatar generation and motion generation. For the first part, the natural language description of the shape tshapet_{\textrm{shape}} is used for the generation of a coarse body shape βt\beta_{t}. Together with a pre-defined standing pose θstand\theta_{\textrm{stand}}, we get the template mesh Mt={Vt,FSMPL}M_{t}=\{V_{t},F_{\textrm{SMPL}}\}, where vertices Vt=M(βt,θstand;Φ)V_{t}=M(\beta_{t},\theta_{\textrm{stand}};\Phi) and faces FSMPLF_{\textrm{SMPL}} are given by SMPL. MtM_{t} is then rendered to multi-view images for the training of a NeuS model NN, which is later used as the initialization of our implicit 3D avatar. Guided by the appearance description tappt_{\textrm{app}}, NN is further optimized by CLIP in a shape-preserving way for shape sculpting and texture generation. After that, the target static 3D avatar mesh M={V,F,C}M=\{V,F,C\} is extracted from N′N^{\prime} by the marching cube algorithm (Lorensen and Cline, 1987) and aligned with MtM_{t} to get ready for animation. For the second part, the natural language description of the motion tmotiont_{\textrm{motion}} is used to generate candidate poses from a pre-calculated code-book. Then the candidate poses are used as references for the optimization of a pre-trained motion VAE to get the desired motion sequence.

3. Static Avatar Generation

As the first step of generating avatars, we propose to generate a coarse shape as the template mesh MtM_{t} from the input description of the shape tshapet_{\textrm{shape}}. For this step, SMPL is used as the source of possible human body shapes. As shown in Fig. 3, we first construct a code-book by clustering body shapes. Then CLIP is used to extract features of the renderings of body shapes in the code-book and texts to get the best matching body shape. Although the process is straightforward, careful designs are required to make full use of both powerful tools.

For the CLIP-guided code-book query, it is important to design reasonable query scores. We observe that for some attributes like body heights, it is hard to be determined only by looking at the renderings of the body without any reference. Therefore, we propose to set a reference and let CLIP score the body shapes in a relative way, which is inspired by the usage of CLIP in 2D image editing (Patashnik et al., 2021). We define a neutral state body shape MnM_{n} and the corresponding neutral state text tnt_{n} as the reference. The scoring of each entry sis_{i} in the code-book is defined as

where ΔfI=EI(R(MB[i]))−EI(R(Mn))\Delta f_{I}=E_{I}(\mathcal{R}(M_{B}[i]))-E_{I}(\mathcal{R}(M_{n})), ΔfT=ET(tshape)−ET(tn)\Delta f_{T}=E_{T}(t_{\textrm{shape}})-E_{T}(t_{n}), R(⋅)\mathcal{R}(\cdot) denotes a mesh renderer, EIE_{I}, ETE_{T} represents CLIP image and text encoders. ΔfT\Delta f_{T} is the relative direction guided by the text, which is intended to be aligned with the visual relative shape differences ΔfI\Delta f_{I} in the CLIP latent space. Then the entry with the maximum score among the code-book is retrieved as the coarse shape Mt=MB[argmaxi(si)]M_{t}=M_{B}[\textrm{argmax}_{i}(s_{i})].

3.2. Shape Sculpting and Texture Generation

The generated template mesh MtM_{t} represents a desired coarse naked body shape. To generate high-quality 3D avatars, the shape and texture need to be further sculpted and generated to match the description of the appearance tappt_{\textrm{app}}. As discussed previously, we choose to use an implicit representation, i.e. NeuS, as the base 3D representation in this step for its advantages in both geometry and colors. In order to speed up the optimization and more importantly, control the general shape of the avatar for the convenience of animation, this step is designed as a two-stage optimization process.

The first stage creates an equivalent implicit representation of MtM_{t} by optimizing a randomly initialized NeuS NN with the multi-view renderings {IMti}\{I_{M_{t}}^{i}\} of the template mesh MtM_{t}. Specifically, as shown in Fig. 4, the NeuS N={f(p),c(p)}N=\{f(p),c(p)\} comprises of two sub-networks. The SDF network f(p)f(p) takes some point pp as input and outputs the signed distance to its nearest surface. The color network c(p)c(p) takes some point pp as input and outputs the color at that point. Both f(p)f(p) and c(p)c(p) are implemented using MLPs. It should be noted that compared to the original implementation of NeuS, we omit the color network’s dependency on the viewing direction. The viewing direction dependency is originally designed to model the shading that changes with the viewing angle. In our case, it is redundant to model such shading effects since our goal is to generate 3D avatars with consistent textures, which ideally should be albedo maps. Similar to the original design, NN is optimized by a three-part loss function

Lcolor\mathcal{L}_{\textrm{color}} is the reconstruction loss between the rendered images and the ground truth multi-view renderings {IMti}\{I_{M_{t}}^{i}\}. Lreg\mathcal{L}_{\textrm{reg}} is an Eikonal term (Gropp et al., 2020) to regularize the SDF f(p)f(p). Lmask\mathcal{L}_{\textrm{mask}} is a mask loss that encourages the network to only reconstruct the foreground object, i.e. the template mesh MtM_{t}. The resulting NeuS NN serves as an initialization for the second stage optimization.

The second stage stylizes the initialized NeuS NN from the first stage by the power of CLIP. At the same time, the coarse shape MtM_{t} should still be maintained, resulting in two possible solutions. The first one is to fix the weights of f(p)f(p), which ensures the final generated shape is the same as MtM_{t}. The color network c(p)c(p) is optimized to ‘colorize’ the fixed shape. However, as discussed before, MtM_{t} is generated from SMPL, which only provides the shape of a naked human body. Purely adding textures on the template shapes, although simple and tractable, is not considered the optimal solution. Because not only textures but also geometry is crucial in the process of avatar creation. To allow the fine-level shape sculpting, we need to fine-tune the weights of f(p)f(p). At the same time, to maintain the general shape of the template mesh, c(p)c(p) should maintain its original function of reconstructing the template mesh to allow the reconstruction loss during the optimization. This leads to the second solution, where we keep the original two sub-networks f(p)f(p) and c(p)c(p) and introduce an additional color network cc(p)c_{c}(p). Both color networks share the same SDF network f(p)f(p) as shown in Fig. 5. f(p)f(p) and c(p)c(p) are in charge of the reconstruction of the template mesh MtM_{t}. f(p)f(p) together with cc(p)c_{c}(p) are responsible for the stylizing part and comprises the targeting implicit 3D avatar. Formally, the new NeuS model N′={f(p),c(p),cc(p)}N^{\prime}=\{f(p),c(p),c_{c}(p)\} now consists of three sub-networks, where f(p)f(p) and c(p)c(p) are initialized by the pre-trained NN, and cc(p)c_{c}(p) is a randomly initialized color network. All three sub-networks are trainable. Similarly, we ignore the viewing direction dependency in the additional color network.

The second-stage optimization is supervised by a three-part loss function

where L1\mathcal{L}_{1} is the reconstruction loss over NeuS {f(p),c(p)}\{f(p),c(p)\} as defined in Eq. 5. Lclipc\mathcal{L}_{\textrm{clip}}^{c} and Lclipg\mathcal{L}_{\textrm{clip}}^{g} are CLIP-guided loss functions that guide the texture and geometry generation to match the description t1t_{1} using two different types of renderings, which are introduced as follows.

The first type of the rendering is the colored rendering IcI_{c} of the NeuS model {f(p),cc(p)}\{f(p),c_{c}(p)\}, which is calculated by applying Eq. 2 to each pixel of the image. The second type is the texture-less rendering IgI_{g} of the same NeuS model. The rendering algorithm used here is ambient and diffuse shading. For each ray {o,v}\{o,v\}, the normal direction n(o,v)n(o,v) of the first surface point it intersects with can be calculated by the accumulation of the gradient of the SDF function at each position along the ray, which is formulated as

where w(t)w(t) is the weighting function as defined in Eq. 3. To render the texture-less model, a random light direction needs to be sampled. To avoid the condition where the light and camera are at opposite sides of the model and the geometry details of the model cannot be revealed, the light direction is uniformly sampled in a small range around the camera direction. Formally, defining the camera direction in the spherical coordinate system as polar and azimuthal angles {θc,ϕc}\{\theta_{c},\phi_{c}\}, then the light direction ll is sampled from {θc+X1,ϕc+X2}\{\theta_{c}+X_{1},\phi_{c}+X_{2}\}, where X1,X2∼U(−π/4,π/4)X_{1},X_{2}\sim\mathcal{U}(-\pi/4,\pi/4). Since coloring is not needed for the texture-less rendering, we could simply calculate the gray level of the ray {o,v}\{o,v\} by

where A∼U(0,0.2)A\sim\mathcal{U}(0,0.2) is randomly sampled from a uniform distribution, D=1−AD=1-A is the diffusion. By applying Eq. 8 to each pixel of the image, we get the texture-less rendering IgI_{g}. In practice, we find that random shading on textured renderings IcI_{c} benefiting the uniformity of the generated textures, which is discussed later in the ablation study. The random shading on IcI_{c} is similar to the texture-less rendering, which is formally defined as

Then the two types of CLIP-guided loss functions are formally defined as Lclipc=Lclip(Ic,tapp)\mathcal{L}_{\textrm{clip}}^{c}=\mathcal{L}_{\textrm{clip}}(I_{c},t_{\textrm{app}}), Lclipg=Lclip(Ig,tapp)\mathcal{L}_{\textrm{clip}}^{g}=\mathcal{L}_{\textrm{clip}}(I_{g},t_{\textrm{app}}).

To render a H×WH\times W image, a total of H×W×QH\times W\times Q queries need to be performed given QQ query times for each ray. Due to the high memory footprint of volume rendering, the rendering resolution H×WH\times W is heavily constrained. In our experiment on one 32GB GPU, the maximum HmaxH_{\textrm{max}} and WmaxW_{\textrm{max}} are around 110110, which is far from the resolution upper bound of 224224 provided by CLIP. With the network structures and other hyper-parameters not modified, to increase the rendering resolution, we propose a dilated silhouettes-based rendering strategy based on the fact that the rays not encountering any surface do not contribute to the final rendering while consuming large amounts of memories. Specifically, we could get the rays that have high chances of encountering surfaces by calculating the silhouette of the rendering of the template mesh MtM_{t} with the given camera parameter. Moreover, to ensure a proper amount of shape sculpting space is allowed, we further dilate the silhouettes. The rays within the silhouettes are calculated and make contributions to the final rendering. Defining the ratio between the area of the dilated silhouette and the total area of the image as rsr_{s}, this rendering strategy dynamically increases the maximum resolution to Hmax′×Wmax′=Hmax×Wmax/rsH_{\textrm{max}}^{\prime}\times W_{\textrm{max}}^{\prime}=H_{\textrm{max}}\times W_{\textrm{max}}/r_{s}. Empirically, the resolution can be increased to around 1502150^{2}.

In order to further increase the robustness of the optimization process, three augmentation strategies are proposed: a) random background augmentation; b) random camera parameter sampling; c) semantic-aware prompt augmentation. They are introduced as follows.

Inspired by Dream Fields (Jain et al., 2021a), random background augmentation helps CLIP to focus more on the foreground object and prevents the generation of randomly floating volumes. As shown in Fig. 6, we randomly augment the backgrounds of the renderings to 1) pure black background; 2) pure white background; 3) Gaussian noises; 4) Gaussian blurred chessboard with random block sizes.

To prevent the network from finding short-cut solutions that only give reasonable renderings for several fixed camera positions, we randomly sample the camera extrinsic parameters for each optimization iteration in a manually-defined importance sampling way to get more coherent and smooth results. We choose to work with the ‘look at’ mode of the camera, which consists of a look-at point, a camera position, and an up direction. We set the up direction to always align with the up direction of the template body. The look at position is sampled from a Gaussian distribution X,Y,Z∼N(0,0.1)X,Y,Z\sim\mathcal{N}(0,0.1), which are then clipped between −0.3-0.3 and 0.30.3 to prevent the avatar from being out of the frame. The camera position is sampled using a spherical coordinate system with the radius RR sampled from a uniform distribution U(1,2)\mathcal{U}(1,2), the polar angle θc\theta_{c} sampled from a uniform distribution U(0,2π)\mathcal{U}(0,2\pi), the azimuthal angle ϕc\phi_{c} sampled from a Gaussian distribution N(0,π/3)\mathcal{N}(0,\pi/3) for the camera to point at the front side of the avatar in the most of the iterations.

So far, no human prior is introduced in the whole optimization process, which may result in generating textures at wrong body parts or not generating textures for important body parts, which will later be shown in the ablation study. To prevent such problems, we propose to explicitly bring in human prior in the optimization process by augmenting the prompts in a semantic-aware manner. For example, as shown in Fig. 7, if tapp=t_{\textrm{app}}=‘Steve Jobs’, we would augment tappt_{\textrm{app}} to two additional prompts tface=t_{\textrm{face}}=‘the face of Steve Jobs’ and tback=t_{\textrm{back}}=‘the back of Steve Jobs’. For every four iterations, the look-at point of the camera is set to be the center of the face to get renderings of the face, which will be supervised by the prompt tfacet_{\textrm{face}}. The ‘face augmentation’ directly supervises the generation of the face, which is important to the quality of the generated avatar since humans are more sensitive to faces. For the second augmented prompt, when the randomly sampled camera points at the back of the avatar, tbackt_{\textrm{back}} is used as the corresponding text to explicitly guide the generation of the back of the avatar.

3.3. Make the Static Avatar Animatable

With the generated implicit 3D avatar N′N^{\prime}, the marching cube algorithm is performed to extract its meshes M={V,F,C}M=\{V,F,C\} before making it animatable. Naturally, with the careful design of keeping the overall shape unchanged during the optimization, the generated mesh can be aligned with the initial template mesh. Firstly, the nearest neighbor for each vertex in VV is retrieved in the vertices of the template mesh MtM_{t}. Blend weights of each vertex in VV are copied from the nearest vertex in MtM_{t}. Secondly, an inverse LBS algorithm is used to bring the avatar’s standing pose θstand\theta_{\textrm{stand}} back to the zero pose θ0\theta_{0}. The vertices VV are transformed to Vθ0V_{\theta_{0}}. Finally, Vθ0V_{\theta_{0}} can be driven by any pose θ\theta using the LBS algorithm. Hence, for any pose, the animated avatar can be formally defined as M(θ)=(LBS(Vθ0,θ),F,C)M(\theta)=(\textrm{LBS}(V_{\theta_{0}},\theta),F,C).

4. Motion Generation

Empirically, CLIP is not capable of directly estimating the similarities between motion sequences and natural language descriptions. It also lacks the ability to assess the smoothness or rationality of motion sequences. These two limitations suggest that it is hard to generate motions only using CLIP supervision. Hence, we have to introduce other modules to provide motion priors. However, CLIP has the ability to value the similarity between a rendered human pose and a description. Furthermore, similar poses can be regarded as references for the expected motions. Based on the above observations, we propose a two-stage motion generation process: 1) candidate poses generation guided by CLIP. 2) motion sequences generation using motion priors with candidate poses as references. The details are illustrated as follows.

To generate poses consistent with the given description tmotiont_{\textrm{motion}}, an intuitive method is to directly optimize the parameter θ\theta in the SMPL model or the latent code of a pre-trained pose VAE (e.g. VPoser (Pavlakos et al., 2019)). However, they can hardly yield reasonable poses due to difficulties during optimization, which are later shown in the experiments. Hence, it is not a wise choice to directly optimize the poses.

Given the motion description tmotiont_{\textrm{motion}}, we calculate the similarity between tmotiont_{\textrm{motion}} and each pose Bθ[i]B_{\theta}[i] from the code-book BθB_{\theta}, which can be defined as

Top-k scores sis_{i} and their corresponding poses Bθ[i]B_{\theta}[i] are selected to construct the candidate pose set SS, which serves as references to generate a motion sequence in the next stage.

4.2. Reference-Based Animation with Motion Prior

We propose a two-fold method to generate a target motion sequence that matches the motion description tmotiont_{\textrm{motion}}. 1) A motion VAE is trained to capture human motion priors. 2) We optimize the latent code of the motion VAE using candidate poses SS as references. The details are introduced as follows.

The motion encoder EmotionE_{\textrm{motion}} contains a projection layer, a positional embedding layer, several transformer encoder layers and an output layer. Formally, EmotionE_{\textrm{motion}} can be defined as y=Emotion(Θ′)=Eo(TransEncoder⁡(ϕ(Ep(Θ′))))y=E_{\textrm{motion}}(\Theta^{\prime})=E_{o}(\operatorname{TransEncoder}(\phi(E_{p}(\Theta^{\prime})))), where EpE_{p} and EoE_{o} are fully connected layers, ϕ\phi represents positional embedding operation (Vaswani et al., 2017). TransEncoder⁡\operatorname{TransEncoder} includes multiple transformer encoder layers (Vaswani et al., 2017).

The reparameterization module yields a Gaussian distribution where μ=Rμ(y),σ=Rσ(y)\mu=R_{\mu}(y),\sigma=R_{\sigma}(y). Rμ,RσR_{\mu},R_{\sigma} are fully connected layers to calculate the mean μ\mu and the standard deviation σ\sigma of the distribution, respectively. Using the reparameterization trick (Kingma and Welling, 2013), a random latent code zmotionz_{\textrm{motion}} is sampled under the distribution N(μ,σ)\mathcal{N}(\mu,\sigma). The latent code is further decoded by DmotionD_{\textrm{motion}}, which can be formally defined as Θ∗=Dmotion(zmotion)=Do(TransDecoder⁡(Dp(zmotion),ϕ(0)))\Theta^{\ast}=D_{\textrm{motion}}(z_{\textrm{motion}})=D_{o}(\operatorname{TransDecoder}(D_{p}(z_{\textrm{motion}}),\phi(\mathbf{0}))), where DpD_{p} and DoD_{o} are fully connected layers, 0\mathbf{0} is a zero-vector. TransDecoder⁡\operatorname{TransDecoder} contains several transformer decoder layers. A loss function with two terms is proposed to train the motion VAE, which is defined as

where λ5\lambda_{5} is a hyper-parameter to balance two terms. LKL\mathcal{L}_{\textrm{KL}} computes the KL-divergence to enforce the distribution assumption. Lrecon\mathcal{L}_{\textrm{recon}} is the mean squared error between Θ+\Theta^{+} and Θ∗\Theta^{*} for the reconstruction of the motion sequences.

With the pre-trained motion VAE, we attempt to optimize its latent code ztz_{t} to synthesize a motion sequence Θ=Dmotion(zt)\Theta=D_{\textrm{motion}}(z_{t}). As shown in Fig. 10, three optimization constraints are proposed as

where λ6,λ7\lambda_{6},\lambda_{7} are hyper-parameters to balance the terms. Lpose\mathcal{L}_{\textrm{pose}} is the reconstruction term between the decoded motion sequence Θ\Theta and candidate poses SS. Ldelta\mathcal{L}_{\textrm{delta}} measures the range of the motion to prevent the motion from being overly-smoothed. Lclipm\mathcal{L}_{\textrm{clip}}^{m} encourages each single pose in the motion to match the input motion description. Details of three loss terms are introduced as follows.

Given reference poses S={θ1,θ2,…,θk}S=\{\theta_{1},\theta_{2},\dots,\theta_{k}\}, the target is to construct a motion sequence that is close enough to these poses. We propose to minimize the distance between θi\theta_{i} and its nearest frame Θj\Theta_{j}, where i∈{1,2,…,k}i\in\{1,2,\dots,k\}, and j∈{1,2,…,L}j\in\{1,2,\dots,L\}. Note that we assume θi\theta_{i} is less similar to tmotiont_{\textrm{motion}} with larger ii. Therefore, we use a coefficient λpose(i)=1−i−1k\lambda_{\textrm{pose}}(i)=1-\frac{i-1}{k} to focus more on candidate poses with higher similarities. Formally, the reconstruction loss is defined as

Empirically, only using the reconstruction loss Lpose\mathcal{L}_{\textrm{pose}}, the generated motion tends to be over-smoothed. To generate motions with larger motion ranges, we design a motion range term Ldelta\mathcal{L}_{\textrm{delta}} to measure the smoothness of adjacent poses,

which serves as a penalty term against over-smoothed motions. More intense motions will be generated when increasing λ6\lambda_{6}.

The matching scheme of the reconstruction term Lpose\mathcal{L}_{pose} does not guarantee the ordering of candidate poses. Lack of supervision on the pose orderings would lead to unstable generation results. Furthermore, candidate poses might only contribute to a small part of the final motion sequence, which would lead to unexpected motion pieces. To tackle these two problems, we design an additional CLIP-guided loss term

where sis_{i} is the similarity score between the pose Θi\Theta_{i} and text description tmotiont_{\textrm{motion}}, which is defined as

λclip(i)=iL\lambda_{\textrm{clip}}(i)=\frac{i}{L} is a monotonically increasing function so that Lclipm\mathcal{L}_{\textrm{clip}}^{m} gives higher penalty to the later poses in the sequence. With the above CLIP-guided term, the whole motion sequence will be more consistent with tmotiont_{\textrm{motion}}. Empirically, we find that we only need to sample a small part of poses in Θ\Theta for the calculation of this term, which will speed up the optimization without observable degradation in performance.

Experiments

The shape VAE (Sec. 3.3.1) uses a two-layer MLP with a 16 dimension latent space for the encoder and decoder. Then 40,960 latent codes are sampled from a Gaussian distribution and clustered to 2,048 centroids by K-Means which composite the code-book. For the shape sculpting and texture generation (Sec. 3.3.2), we adjust the original NeuS such that the SDF network uses a 6-layer MLP and the color network uses a 4-layer MLP. For each ray, we perform 32 uniform samplings and 32 importance samplings. The Adam algorithm (Kingma and Ba, 2014) with a learning rate of 5×10−45\times 10^{-4} is used for 30,000 iterations of optimization.

For candidate pose generation (Sec. 3.4.1), we directly use the pre-trained VPoser (Pavlakos et al., 2019) and use K-Means to acquire 4,096 cluster centroids from the AMASS dataset (Mahmood et al., 2019). We select top-5 poses for the next stage (Sec.3.4.2). Our proposed Motion VAE has a 256 dimension latent space. The length of the motion is 60. We train 100 epochs for motion VAE on AMASS dataset. An Adam optimizer with a 5×10−45\times 10^{-4} learning rate is used. During optimizing latent code for tmotiont_{\textrm{motion}}, Adam is used for 5,000 iterations of optimization with a 1×10−21\times 10^{-2} learning rate.

1.2. Baselines

Though it is the first work to generate and animate 3D avatars in a zero-shot text-driven manner, we design reasonable baseline methods for the evaluation of each part of AvatarCLIP. For the coarse shape generation, we design a baseline method where the shape parameters (i.e. the SMPL β\beta and latent code of the shape VAE) are directly optimized by CLIP-guided losses. For the shape sculpting and texture generation, we compare our design with Text2Mesh (Michel et al., 2021). Moreover, we introduce a NeRF-based baseline by adapting Dream Fields (Jain et al., 2021a), where our ‘additional color network’ design is added to its pipeline to constraint the general shape of the avatar.

For the candidate pose generation, three baseline methods are designed to compare with our method. To illustrate the difficulty of direct optimization, we set two baselines that directly optimize on SMPL parameter θ\theta and latent code zz in VPoser. Moreover, inspired by CLIP-Forge (Sanghi et al., 2021), we use Real NVP (Dinh et al., 2016) to get a bi-projection between the normal distribution and latent space distribution of VPoser. This normalization flow network is conditioned on the CLIP features. This method does not need paired pose-text data. As for the second part of the motion generation, we design two baseline methods to compare with: Baseline (i) first sorts candidate poses SS by their similarity scores sis_{i}. Then, direct interpolations over the latent codes between each pair of adjacent poses are performed to generate the motion sequence. Baseline (ii) uses the motion VAE to introduce motion priors into the generative pipeline. But (ii) directly calculates the reconstruction loss without using the re-weighting technique.

2. Overall Results

Overall results of the whole pipeline of AvatarCLIP are shown in Fig. 11. Avatars with diverse body shapes along with varied appearances are generated with high quality. They are driven by generated motion sequences that are reasonable and consistent with the input descriptions. In a zero-shot style, AvatarCLIP is capable of generating animatable avatars and motions, making use of the strong prior in pre-trained models. The whole process of avatar generation and animation, which originally requires expert knowledge of professional software, can now be simply driven by natural languages with the help of our proposed AvatarCLIP.

3. Experiments on Avatar Generation

To validate the effectiveness of various designs in the avatar generation module, we perform extensive ablation studies. We ablate the designs of 1) background augmentation; 2) supervision on texture-less renderings; 3) random shading on the textured renderings; 4) semantic-aware prompt augmentation. The ablation settings shown in Fig. 12 are formed by subsequently adding the above four designs to a baseline method where only textured renderings are supervised by CLIP. As shown in the first two columns of Fig. 12, background augmentation has a great influence on the texture generation, without which the textures tend to be very dark. Comparing the second and third columns, adding the supervision on texture-less renderings improves the geometry quality by a large margin. The geometry of ‘Ablation 2’ has lots of random bumps, which make the surfaces noisy. While the geometry of ‘Ablation 3’ is smooth and has detailed wrinkles of garments. As shown by the ‘Ablation 3’ and ‘Ablation 4’, adding random shadings on textured renderings helps the generation of more uniform textures. For example, the ‘Donald Trump’ of ‘Ablation 3’ has a brighter upper body than the lower one, which is improved in ‘Ablation 4’. Without the awareness of human body semantics, the previous four settings cannot generate correct faces for the avatars. The last column, which uses the semantic-aware prompt augmentation, has the best results in terms of the face generation.

3.2. Qualitative Results of Coarse Shape Generation.

For this part, we design two intuitive baseline methods where direct CLIP supervision is back-propagated to the shape parameters. As shown in Fig. 13 (a), both optimization methods fail to generate body shapes consistent with description texts. Even opposite text guidance (e.g. ‘skinny’ and ‘overweight’) leads to the same optimization direction. In comparison, our method robustly generates reasonable body shapes agreeing with input texts. More diverse qualitative results of our method are shown in Fig. 13 (b).

3.3. Qualitative Results of Shape Sculpting and Texture Generation.

Throughout extensive experiments, our method is capable of generating avatars from a wide range of appearances descriptions including three types: 1) celebrities, 2) fictional characters and 3) general words that describe people, as shown in Fig. 17. As shown in Fig. 17 (a), given celebrity names as the appearance description, the most iconic outfit of the celebrity is generated. Thanks to our design of semantic-aware prompt augmentation, the faces are also generated correctly. For the fictional character generation as illustrated in Fig. 17 (b), avatars of most text descriptions can be correctly generated. Interestingly, for the characters that have accessories with complex geometry (e.g. helmets of ‘Batman’, the dress of ‘Elsa’), the optimization process has the tendency of ‘growing’ new structures out of the template human body. As for the general descriptions, our method can handle very broad ranges including common job names (e.g. ‘Professor’, ‘Doctor’), words that describe people at a certain age (e.g. ‘Teenager’, ‘Senior Citizen’) and other fantasy professions (e.g. ‘Witch’, ‘Wizard’). It can be observed that other than the avatars themselves, their respective iconic objects can also be generated. For example, the ‘Gardener’ grasps flowers and grasses in his hands.

Other than the overall descriptions of the appearance as shown above, our method is also capable of zero-shot controlling at a more detailed level. As shown in Fig. 18, we can control the faces generated on the avatars, e.g. Bill Gates wearing an Iron Man suit, by tuning the semantic-aware prompt augmentation. Moreover, we can control the clothing of the avatar by direct text guidance, e.g. ‘Steve Jobs in white shirt’.

Inspired by DALL-E (Ramesh et al., 2021) and Dream Fields (Jain et al., 2021a), one of the most exciting applications of CLIP-driven generation is concept mixing. As shown in Fig. 14, we demonstrate examples of mixing fantasy elements with celebrities. While maintaining the recognizable identities, the fantasy elements blend in with the generated avatars naturally.

The critical design of texture-less rendering supervision mainly contributes to the geometry generation. We mainly compare our AvatarCLIP with the adapted Dream Field, which is based on NeRF. As shown in Fig. 19, our method consistently outperforms Dream Fields in terms of geometry quality. Detailed muscle shapes, armor curves, and cloth wrinkles can be generated.

Other than the generation quality, we also investigate the robustness of our algorithm compared with the baseline method Text2Mesh. For each method, we use the same five random seeds for five independent runs for the same prompt. As shown in Fig. 20, our method manages to output results with high quality and consistency with the input text. Text2Mesh fails most runs, which shows that the representations of meshes are unstable for optimization, especially with weak supervision.

Although a wide range of generation results are experimented with and demonstrated above, there exist failure cases in generating loose garments and accessories. For example, as shown in Fig. 17, the dress of ‘Elsa’ and the cloak of ‘Doctor Strange’ are not generated. The geometry of exaggerated hair and beard (e.g. the breaded Forrest Gump) is also challenging to be generated correctly. This is caused by the reconstruction loss in the optimization process. The ‘growing’ is discouraged and very limited changes of the geometry are allowed.

3.4. Quantitative Results

To quantitatively evaluate the results of our avatar generation method, we ask 2222 volunteers to perform a user study in terms of 1) the consistency with input texts, 2) texture quality, and 3) geometry quality. We randomly select 88 input texts that describe appearances. They are used for avatar generation by three methods, i.e. Dream Field, Text2Mesh and our AvatarCLIP. For each sample, the volunteers are asked to score the results of three methods from 11 to 55 in terms of the above three aspects. As illustrated in Fig. 15, our method consistently outperforms the other two baseline methods in all three aspects. Moreover, the standard deviations of our method are the lowest among the three methods, which also demonstrates the stable quality of our method.

4. Experiments on Motion Generation

To evaluate the effectiveness of our design choices in the reference-based animation module (i.e. the proposed three constraint terms, the usage of motion VAE as a motion prior), we compare our method with two baseline methods as shown in Fig. 21. For ‘Brush Teeth’, (i) generates unordered and unrelated pose sequences. (ii) also fails to generate reasonable motions, which is caused by its blind focus on reconstructing the unordered candidate poses. By introducing a re-weighting mechanism, our method not only focuses on reconstruction but also considers the rationality of the generated motion. For ‘Kick Soccer’, the motion sequences from (i) (ii) have no drastic changes when the leg kicks out. Ldelta\mathcal{L}_{\textrm{delta}} plays a significant role here to control the intensity of motions. As for ‘Raise Both Arms’, it is supposed to generate a motion sequence from a neutral pose to a pose with raised arms. However, (i) generates a motion sequence that is contrary to the expected result. (ii) introduces several unrelated actions. With the help of Lclipm\mathcal{L}_{\textrm{clip}}^{m}, our method is capable of generating motions with the correct ordering and better consistency with the descriptions.

We also demonstrate some failure cases in Fig. 23. Throughout experiments, we find it hard to precisely control the body parts. For example, we cannot specifically control left or right hands, as shown in ‘raising left arm’ case. Moreover, limited by the diversity of candidate poses, it is challenging to generate more complex motions like ‘hugging’, ‘playing piano’ and ‘dribbling’.

4.2. Qualitative Results of Candidate Pose Generation

Comparisons between different candidate pose generation methods are shown in Fig. 22 (a). Both direct optimization methods (i) (ii) fail to generate reasonable poses, let alone poses that are consistent with the given description. These results suggest that direct optimization on human poses parameters is intractable. Compared with (i) and (ii), conditioned Real NVP (iii) can yield rational poses. However, compared to the proposed solution based on the code-book, the generated poses from (iii) are less reasonable and of lower quality.

To demonstrate the zero-shot ability of the pose generation method, we experiment with four categories of motion descriptions: 1) abstract emotion descriptions (e.g. ‘tired’ and ‘sad’); 2) common action descriptions (e.g. ‘walking’ and ‘squatting’); 3) descriptions of motions related to body parts (e.g. ‘raising both arms’, ‘washing hands’) 4) descriptions of motions that involve interaction with objects (e.g. ‘shooting basketball’). They are shown in Fig.22 (b).

4.3. Quantitative Results

Quantitatively, we evaluate the candidate pose generation and reference-based animation separately. For the candidate pose generation, we ask 5858 volunteers to select the candidate poses that are the most consistent with the given text inputs for 1515 randomly selected samples. The percentage of the selected times of each method for each text input is used as the scores for counting. As shown in Fig. 16 (a), compared with the baseline methods introduced in Sec. 4.4.2, our method outperforms them by a large margin. For the reference-based animation, we ask 2020 volunteers to score the results from 1 to 5 in terms of the consistency with the input texts and the overall quality for 1010 randomly selected samples. As shown in Fig. 16 (b), compared with the two baseline methods introduced in Sec. 4.4.1, our method outperforms them by large margins in both consistency and quality.

Discussion

In this work, by proposing AvatarCLIP, we make the originally complex and demanding 3D avatar creation accessible to layman users in a text-driven style. It is made possible with the powerful priors provided by the pre-trained models including the shape/ motion VAE and the large-scale vision-language pre-trained model CLIP. Extensive experiments are conducted to validate the effectiveness of careful designs of our methods.

For the avatar generation, limited by the weak supervision and low resolution of CLIP, the results are not perfect if zoomed in. Besides, it is hard to generate avatars with large variations given the same prompt. For the same prompt, the CLIP text feature is always the same. Therefore, the optimization directions are the same, which leads to similar results across different runs. For the motion synthesis, limited by the code-book design, it is hard to generate out-of-distribution candidate poses, which limits the ability to generate complex motions. Moreover, due to the lack of video CLIP, it is difficult to generate stylized motion.

The usage of pre-trained models might induce ethical issues. For example, if we let tappt_{\textrm{app}} = ‘doctor’, the generated avatar is male. If tappt_{\textrm{app}} = ‘nurse’, the generated avatar is female, which demonstrates the gender bias. We think the problem originates from the large-scale internet data used for CLIP training, which might be biased if not carefully reviewed. Future works regarding the ethical issues of large-scale pre-trained models are required for the zero-shot techniques safe to be used. Moreover, with the democratization of producing avatars and animations, users can easily produce fake videos of celebrities, which might be misused and cause negative social impacts.

References