MotionCLIP: Exposing Human Motion Generation to CLIP Space
Guy Tevet, Brian Gordon, Amir Hertz, Amit H. Bermano, Daniel Cohen-Or
Introduction
Human motion generation includes the intuitive description, editing, and generation of 3D sequences of human poses. It is relevant to many applications that require virtual or robotic characters. Motion generation is, however, a challenging task. Perhaps the most challenging aspect is the limited availability of data, which is expensive to acquire and to label. Recent years have brought larger sets of motion capture acquisitions [Mahmood et al., 2019], sometimes sorted by classes [Liu et al., 2019; Ji et al., 2018] or even labeled with free text [Punnakkal et al., 2021; Plappert et al., 2016]. Yet, it seems that while this data may span a significant part of human motion, it is not enough for machine learning algorithms to understand the semantics of the motion manifold, and it is definitely not descriptive enough for natural language usage. Hence, neural models trained using labeled motion data [Ahuja and Morency, 2019; Lin et al., 2018; Yamada et al., 2018; Petrovich et al., 2021; Maheshwari et al., 2022] do not generalize well to the full richness of the human motion manifold, nor to the natural language describing it.
In this work, we introduce MotionCLIP, a 3D motion auto-encoder that induces a latent embedding that is disentangled, well behaved, and supports highly semantic and elaborate descriptions. To this end, we employ CLIP [Radford et al., 2021], a large scale visual-textual embedding model. Our key insight is that even though CLIP has not been trained on the motion domain what-so-ever, we can inherit much of its latent space’s virtue by enforcing its powerful and semantic structure onto the motion domain. To do this, we train a transformer-based [Vaswani et al., 2017] auto-encoder that is aligned to the latent space of CLIP, using existing motion textual labels. In other words, we train an encoder to find the proper embedding of an input sequence in CLIP space, and a decoder that generates the most fitting motion to a given CLIP space latent code. To further improve the alignment with CLIP-space, we also leverage CLIP’s visual encoder, and synthetically render frames to guide the alignment in a self-supervised manner (see Figure 2). As we demonstrate, this step is crucial for out-of-domain generalization, since it allows finer-grained description of the motion, unattainable using text.
The merit of aligning the human motion manifold to CLIP space is two-fold: First, combining the geometric motion domain with lingual semantics benefits the semantic description of motion. As we show, this benefits tasks such as text-to-motion and motion style transfer. More importantly however, we show that this alignment benefits the motion latent space itself, infusing it with semantic knowledge and inherited disentanglement. Indeed, our latent space demonstrates unprecedented compositionality of independent actions, semantic interpolation between actions, and even natural and linear latent-space based editing.
As mentioned above, the textual and visual CLIP encoders offer the semantic description of motion. In this aspect, our model demonstrates never-before-seen capabilities for the field of motion generation. For example, motion can be specified using arbitrary natural language, through abstract scene or intent descriptions instead of the motion directly, or even through pop-culture references. For example, the CLIP embedding for the phrase “wings” is decoded into a flapping motion like a bird, and “Williams sisters” into a tennis serve, since these terms are encoded close to motion seen during training, thanks to CLIP’s semantic understanding. Through the compositionality induced by the latent space, the aforementioned process also yields clearly unseen motions, such as the iconic web-swinging gesture that is produced for the input ”Spiderman” (see this and other culture references in Figure 1). Our model also naturally extents to other downstream tasks. In this aspect, we depict motion interpolation to depict latent smoothness, editing to demonstrate disentanglement, and action recognition to point out the semantic structure of our latent space. For all these applications, we show comparable or preferable results either through metrics or a user study, even though each task is compared against a method that was designed especially for it. Using the action recognition benchmark, we also justify our design choices with an ablation study.
Related Work
One means to guide motion generation is to condition on another domain. An immediate, but limited, choice is conditioning on action classes. ACTOR [Petrovich et al., 2021] and Action2Motion [Guo et al., 2020] suggested learning this multi-modal distribution from existing action recognition datasets using Conditional Variational-Autoencoder(CVAE) [Sohn et al., 2015] architectures. MUGL [Maheshwari et al., 2022] model followed with elaborated Conditional Gaussian-Mixture-VAE [Dilokthanakul et al., 2016] that supports up to classes and multi-person generation, based on the NTU-RGBD-120 dataset [Liu et al., 2019].
Motion can be conditioned on other domains. For example, recent works [Li et al., 2021; Aristidou et al., 2021] generated dance moves conditioned on music and the motion prefix. Edwards et al. generated facial expressions to fit a speaking audio sequence.
A more straightforward approach to control motion is using another motion. In particular, for style transfer applications. Holden et al. suggested to code style using the latent code’s Gram matrix, inspired by Gatys et al. . Aberman et al. injected style attributes using a dedicated temporal-invariant AdaIN layer [Huang and Belongie, 2017]. Recently, Wen et al. encoded style in the latent code of Normalizing Flow generative model [Dinh et al., 2014]. We show that MotionCLIP also encodes style in its latent representation, without making any preliminary assumptions or using a dedicated architecture.
2. Text-to-Motion
The KIT dataset[Plappert et al., 2016] provides about 11 hours of motion capture sequences, each sequence paired with a sentence explicitly describing the action performed. KIT sentences describe the action type, direction and sometimes speed, but lacks details about the style of the motion, and not including abstract descriptions of motion. Current text-to-motion research is heavily based on KIT. Plappert et al. learned text-to-motion and motion-to-text using seq2seq RNN-based architecture. Yamada et al. learned those two mappings by simultaneously training text and motion auto-encoders while binding their latent spaces using text and motion pairs. Lin et al. further improved trajectory prediction by adding a dedicated layer. Ahuja et al. introduced JL2P model, which got improved results with respect to nuanced concepts of the text, namely velocity, trajectory and action type. They learned joint motion-text latent space and apply training curriculum to ease optimization.
More recently, BABEL dataset [Punnakkal et al., 2021] provided per-frame textual labels ordered in classes to the larger AMASS dataset [Mahmood et al., 2019], including about 40 hours of motion capture. Although providing explicit description of the action, often lacking any details besides the action type, this data spans a larger variety of human motion. MotionCLIP overcomes the data limitations by leveraging out-of-domain knowledge using CLIP [Radford et al., 2021].
3. CLIP aided Methods
Neural networks have successfully learned powerful latent representations coupling natural images with natural language describing it [He and Peng, 2017; Ramesh et al., 2021]. A recent example is CLIP[Radford et al., 2021], a model coupling images and text in deep latent space using a constructive objective[Hadsell et al., 2006; Chen et al., 2020]. By training over hundred millions of images and their captions, CLIP gained a reach semantic latent representation for visual content. This expressive representation enables high quality image generation and editing, controlled by natural language [Patashnik et al., 2021; Gal et al., 2021; Frans et al., 2021]. Even more so, this model has shown that connecting the visual and textual worlds also benefits purely visual tasks [Vinker et al., 2022], simply by providing a well-behaved, semantically structured, latent space.
Closer to our method are works that utilize the richness of CLIP outside the imagery domain. In the 3D domain, CLIP’s latent space provides a useful objective that enables semantic manipulation [Sanghi et al., 2021; Michel et al., 2021; Wang et al., 2021a] where the domain gap is closed by a neural rendering. CLIP is even adopted in temporal domains [Guzhov et al., 2021; Luo et al., 2021; Fang et al., 2021] that utilize large datasets of video sequences that are paired with text and audio. Unlike these works that focus on classification and retrieval, we introduce a generative approach that utilizes limited amount of human motion sequences that are paired with text.
Method
Our goal is learning a semantic and disentangled motion representation that will serve as a basis for generation and editing tasks. To this end, we need to learn not only the mapping to this representation (encoding), but also the mapping back to explicit motion (decoding).
Our training process is illustrated in Figure 2. We train a transformer-based motion auto-encoder, while aligning the latent motion manifold to CLIP joint representation. We do so using (i) a Text Loss, connecting motion representations to the CLIP embedding of their text labels, and (ii) an Image Loss, connecting motion representations to CLIP embedding of rendered images that depict the motion visually.
At inference time, semantic editing applications can be performed in latent space. For example, to perform style transfer, we find a latent vector representing the style, and simply add it to the content motion representation and decode the result back into motion. Similarly, to classify an action, we can simply encode it into the latent space, and see to which of the class text embedding it is closest. Furthermore, we use the CLIP text encoder to perform text-to-motion - An input text is decoded using the text encoder then directly decoded by our motion decoder. The implementation of these and other applications is detailed in Section 4.
To project the motion manifold into the latent space, we learn a transformer-based auto-encoder [Vaswani et al., 2017], adapted to the motion domain [Petrovich et al., 2021; Wang et al., 2021b; Li et al., 2021]. MotionCLIP’s architecture is detailed in Figure 3.
Transformer Encoder. , Maps a motion sequence to its latent representation . The sequence is embedded into the encoder’s dimension by applying linear projection for each frame separately, then adding standard positional embedding. The embedded sequence is the input to the transformer encoder, together with additional learned prefix token . The latent representation, is the first output (the rest of the sequence is dropped out). Explicitly, .
Transformer Decoder. , predicts a motion sequence given a latent representation . This representation is fed to the transformer as key and value, while the query sequence is simply the positional encoding of . The transformer outputs a representation for each frame, which is then mapped to pose space using a linear projection. Explicitly, . We further use a differentiable SMPL layer to get the mesh vertices locations, .
Losses. This auto-encoder is trained to represent motion via reconstruction losses on joint orientations, joint velocities and vertices locations. Explicitly,
Given text-motion and image-motion pairs, , correspondingly, we attach the motion representation to the text and image representations using cosine distance,
The motion-text pairs can be derived from labeled motion dataset, whereas the images can be achieved by rendering a single pose from a motion sequence, to a synthetic image , in an unsupervised manner (More details in Section 4).
Overall, the loss objective of MotionCLIP is defined,
Results
To evaluate MotionCLIP, we consider its two main advantages. In Section 4.2, we inspect MotionCLIP’s ability to convert text into motion. Since the motion’s latent space is aligned to that of CLIP, we use CLIP’s pretrained text encoder to process input text, and convert the resulting latent embedding into motion using MotionCLIP’s decoder. We compare our results to the state-of-the-art and report clear preference for both seen and unseen generation. We also show comparable performance to state-of-the-art style transfer work simply by adding the style as a word to the text prompt. Lastly, we exploit CLIP expert lingual understanding to convert abstract text into corresponding, and sometimes unexpected, motion.
In Section 4.3 we focus on the resulting auto-encoder, and the properties of its latent-space. We inspect its smoothness and disentanglement. Smoothness is shown through well-behaved interpolations, even between distant motion. Disentanglement is demonstrated using latent space arithmetic; by adding and subtracting various motion embeddings, we achieve compositionality and semantic editing. Lastly, we leverage our latent structure to perform action recognition over the trained encoder. The latter setting is also used for ablation study. In the following, we first lay out the data used, and other general settings.
We train our model on the BABEL dataset [Punnakkal et al., 2021]. It comprises about 40 hours of motion capture data, represented with the SMPL body model [Loper et al., 2015]. Each frame is annotated with per-frame textual labels, and is categorized into one of 260 action classes. We down sample the data to frames per-second and cut it into sequences of length . We get a single textual label per sequence by listing all actions in a given sequence, then concatenating them to a single string. Finally, we choose for each motion sequence a random frame to be rendered using the Blender software and the SMPL-X add-on [Pavlakos et al., 2019] (See Figure 4). This process outputs triplets of (motion, text, synthetic image) which are used for training.
We train a transformer auto-encoder with layers for each encoder and decoder as described in Section 3. We align it with the CLIP-ViT-B/32 frozen model. Out of the data triplets, the text-motion pairs are used for the text loss and image-motion pairs for the image loss. Both values are set to throughout our experiments. https://github.com/GuyTevet/MotionCLIP
2. Text-to-Motion
Text-to-motion is performed at inference time, using the CLIP text encoder and MotionCLIP decoder, without any further training. Even though not directly trained for this task, MotionCLIP shows unprecedented performance in text-to-motion, dealing with explicit descriptions, subtle nuances and abstract language.
Actions. We start by demonstrating the capabilities of MotionCLIP to generate explicit actions - both seen and unseen in training. We compare our model to JL2P [Ahuja and Morency, 2019]. Since the two models were trained on different datasets, we define a new common ground for evaluation. We define two new sets of samples for a user study: (1)The in-domain set comprises actions with textual labels that appear in at least of the labels of both datasets, and (2) the Out-of-domain set includes textual labels that do not appear in any of the labels of both datasets, hence, unseen for both models. For fairness, we construct this set from the list of Olympic sports (both summer and winter) that are disjoint to both datasets. We conduct a user study, comparing the generation of each model conditioned on a given textual label. For each example, we then ask users to choose which of the two motions best fits the label. Table 1 shows that MotionCLIP was clearly preferred by the users for both sets. Figure 5 demonstrates a variety of sports performed by MotionCLIP, as used in the user-study. Note how even though this is not a curated list, the motion created according to all 30 depicted text prompts resembles the requested actions.
Styles. We investigate MotionCLIP’s ability to represent motion style, without being explicitly trained for it. We compare the results produced by MotionCLIP to the style transfer model by Aberman et al. . The latter receives two input motion sequences, one indicating content and the other style, and combines them through a dedicated architecture, explicitly trained to disentangle style and content from a single sequence. In contrast, we simply feed MotionCLIP with the action and style textual names (e.g.“walk proud”). We show to users the outputs of the two models side-by-side and ask them to choose which one presents both style and/or action better (See Figure 6). Even though Aberman et al. was trained specifically for this task and gets the actual motions as an input, rather then text, Table 2 shows comparable results for the two models, with an expected favor toward Aberman et al.. This, of course, also means that MotionCLIP allows expressing style with free text, and does not require an exemplar motion to describe it. Such novel free text style augmentations are demonstrated in Figure 7.
Abstract language. One of the most exciting capabilities of MotionCLIP is generating motion given text that doesn’t explicitly describe motion. This includes obvious linguistic connections, such as the act of sitting down, produced from the input text ”couch”. Other, more surprising examples include mimicking the signature moves of famous real and fictional figures, like Usain Bolt and The Karate Kid, and other cultural references like the famous ballet performance of Swan Lake and the YMCA dance (Figures 1 and 8). These results include motions definitely not seen during training (e.g., Spiderman in Figure 1), which strongly indicates how well the motion manifold is aligned to CLIP space.
3. Motion Manifold Applications
It is already well established that the CLIP space is smooth and expressive. We demonstrate its merits also exist in the aligned motion manifold, through the following experiments.
Interpolation As can be seen in Figure 9, linear interpolation between two latent codes yields semantic transitions between motions in both time and space. This is a strong indication to the smoothness of this representation. Here, the source and target motions (top and bottom respectively) were sampled from the validation set, and between them are three transitions evenly sampled from the linear trajectory between the two motion representations, then decoded by MotionCLIP.
Latent-Based Editing To demonstrate how disentangled and uniform MotionCLIP latent space is, we experiment with latent-space arithmetic to edit motion (see Figure 10). As can be seen, these linear operations allow motion compositionality - the upper body action can be decomposed from the lower body one, and recomposed with another lower body performance. In addition, Style can be added by simply adding the vector of the style name embedding. These two properties potentially enable intuitive and semantic editing even for novice users.
Action Recognition Finally, we further demonstrate how well our latent spaces is semantically structured. We show how combined with the CLIP text encoder, MotionCLIP encoder can be used for action recognition. We follow BABEL -classes benchmark and train the model with BABEL class names instead of the raw text. At inference, we measure the cosine distance of a given motion sequence to all class name encodings and apply softmax, as suggested originally for image classification [Radford et al., 2021]. In table 3, we compare Top-1 and Top-5 accuracy of MotionCLIP classifier to 2s-AGCN classifier [Shi et al., 2019], as reported by Punnakkal et al. . As can be seen, this is another example where our framework performs similarly to dedicated state-of-the-art methods, even though MotionCLIP was not designed for it.
Conclusions
We have presented a motion generation network that leverages the knowledge encapsulated in CLIP, allowing intuitive operations, such as text conditioned motion generation and editing. As demonstrated, training an auto-encoder on the available motion data alone struggles to generalize well, possibly due to data quality or the complexity of the domain. Non the less, we see that the same auto-encoder with the same data can lead to a significantly better understanding of the motion manifold and its semantics, merely by aligning it to a well-behaved knowledge-rich latent space.
We restress the fascinating fact that even though CLIP has never seen anything from the motion domain, or any other temporal signal, its latent structure naturally induces semantics and disentanglement. This succeeds even though the connection between CLIP’s latent space and the motion manifold is through sparse and inaccurate textual labeling. In essence, the alignment scheme transfers semantics by encouraging the encoder to place semantically similar samples closer together. Similarly, it induces the disentanglement built into the CLIP space, as can be seen, for example, in our latent-space arithmetic experiments.
Of course, MotionCLIP has its limitations, opening several novel research opportunities. It struggles to understand directions, (e.g. left, right and counter-clockwise), to capture some styles (such as heavy and proud), and is of course not consistent for out-of-domain cultural reference exapmles (e.g, it fails to produce Cristiano Ronaldo’s goal celebration, and Superman’s signature pose). Nonetheless, we believe MotionCLIP is an important step toward intuitive motion generation. Knowledge-rich disentangled latent spaces have already proven themselves as a flexible tool to novice users in other fields, such as facial images. In the future, we would like to further explore how powerful large-scale latent spaces could be leveraged to benefit additional domains. We would also like to explore more elaborate architectures and domain adaptation schemes for the main generation part, and to deepen our investigation into downstream tasks that could benefit from this powerful backbone.