AI Choreographer: Music Conditioned 3D Dance Generation with AIST++

Ruilong Li, Shan Yang, David A. Ross, Angjoo Kanazawa

Introduction

The ability to dance by composing movement patterns that align to musical beats is a fundamental aspect of human behavior. Dancing is an universal language found in all cultures , and today, many people express themselves through dance on contemporary online media platforms. The most watched videos on YouTube are dance-centric music videos such as “Baby Shark Dance”, and “Gangnam Style” , making dance a more and more powerful tool to spread messages across the internet. However, dancing is a form of art that requires practice—even for humans, professional training is required to equip a dancer with a rich repertoire of dance motions to create an expressive choreography. Computationally, this is even more challenging as the task requires the ability to generate a continuous motion with high kinematic complexity that captures the non-linear relationship with the accompanying music.

In this work, we address these challenges by presenting a novel Full Attention Cross-modal Transformer (FACT) network, which can robustly generate realistic 3D dance motion from music, along with a large-scale multi-modal 3D dance motion dataset, AIST++, to train such a model. Specifically, given a piece of music and a short (2 seconds) seed motion, our model is able to generate a long sequence of realistic 3D dance motions. Our model effectively learns the music-motion correlation and can generate dance sequences that varies for different input music. We represent dance as a 3D motion sequence that consists of joint rotation and global translation, which enables easy transfer of our output for applications such as motion retargeting as shown in Figure 1.

In order to generate 3D dance motion from music, we propose a novel Full Attention Cross-modal Transformer (FACT) model, which employs an audio transformer and seed motion transformer to encode the inputs, which are then fused by a cross-modal transformer that models the distribution between audio and motion. This model is trained to predict NN future motion sequences and at test time is applied in an auto-regressive manner to generate continuous motion. The success of our model relies on three key design choices: 1) the use of full-attention in an auto-regressive model, 2) future-N supervision, and 3) early fusion of two modalities. The combination of these choices is critical for training a model that can generate a long realistic dance motion that is attuned to the music. Although prior work has explored using transformers for motion generation , we find that naively applying transformers to the 3D dance generation problem without these key choices does not lead to a very effective model.

In particular, we notice that because the context window in the motion domain is significantly smaller than that of language models, it is possible to apply full-attention transformers in an auto-regressive manner, which leads to a more powerful model. It is also critical that the full-attention transformer is trained to predict NN possible future motions instead of one. These two design choices are key for preventing 3D motion from freezing or drifting after several auto-regressive steps as reported in prior works on 3D motion generation . Our model is trained to predict 20 future frames, but it is able to produce realistic 3D dance motion for over 1200 frames at test time. We also show that fusing the two modalities early, resulting in a deep cross-modal transformer, is important for training a model that generates different dance sequences for different music.

In order to train the proposed model, we also address the problem of data. While there are a few motion capture datasets of dancers dancing to music, collecting mocap data requires heavily instrumented environments making these datasets severely limited in the number of available dance sequences, dancer and music diversity. In this work, we propose a new dataset called AIST++, which we build from the existing multi-view dance video database called AIST . We use the multi-view videos to recover reliable 3D motion from this data. We will release code and this dataset for research purposes, where AIST++ can be a new benchmark for the task of 3D dance generation conditioned on music.

In summary, our contributions are as follows:

We propose Full Attention Cross-Modal Transformer model, FACT, which can generate a long sequence of realistic 3D dance motion that is well correlated with the input music.

We introduce AIST++ dataset containing 5.2 hours of 3D dance motions accompanied with music and multi-view images, which to our knowledge is the largest dataset of such kind.

We provide extensive evaluations validating our design choices and show that they are critical for high quality, multi-modal, long motion sequence generation.

Related Work

The problem of generating realistic and controllable 3D human motion sequences has long been studied. Earlier works employ statistical models such as kernel-based probability distribution to synthesize motion, but abstract away motion details. Motion graphs address this problem by generating motions in a non-parametric manner. Motion graph is a directed graph constructed on a corpus of motion capture data, where each node is a pose and the edges represent the transition between poses. Motion is generated by a random walk on this graph. A challenge in motion graph is in generating plausible transition that some approaches address via parameterizing the transition . With the development in deep learning, many approaches explore the applicability of neural networks to generate 3D motion by training on a large-scale motion capture dataset, where network architectures such as CNNs , GANs , RBMs , RNNs and Transformers have been explored. Auto-regressive models like RNNs and vanilla Transformers are capable of generating unbounded motion in theory, but in practice suffer from regression to the mean where motion “freezes” after several iterations, or drift to unnatural motions . Some works propose to ease this problem by periodically using the network’s own outputs as inputs during training. Phase-functioned neural networks and it’s variations address this issue via conditioning the network weights on phase, however, they do not scale well to represent a wide variety of motion.

Audio to motion generation has been studied in 2D pose context either in optimization based approach , or learning based approaches where 2D pose skeletons are generated from a conditioning audio. Training data for 2D pose and audio is abundant thanks to the high reliability of 2D pose detectors . However, predicting motion in 2D is limited in its expressiveness and potential for downstream applications. For 3D dance generation, earlier approaches explore matching existing 3D motion to music using motion graph based approach . More recent approach employ LSTMs , GANs , transformer encoder with RNN decoder or convolutional sequence-to-sequence models. Concurrent to our work, Chen et al. proposed a method that is based on motion graphs with learned embedding space. Many prior works solve this problem by predicting future motion deterministically from audio without seed motion. When the same audio has multiple corresponding motions, which often occurs in dance data, these methods collapse to predicting a mean pose. In contrast, we formulate the problem with seed motion as in , which allows generation of multiple motion from the same audio even with a deterministic model.

Closest to our work is that of Li et al. , which also employ transformer based architecture but only on audio and motion. Furthermore, their approach discretize the output joint space in order to account for multi-modality, which generates unrealistic motion. In this work we introduce a novel full-attention based cross-modal transformer (FACT model) for audio and motion, which can not only preserve the correlation between music and 3D motion better, but also generate more realistic long 3D human motion with global translation. One of the biggest bottleneck in 3D dance generation approaches is that of data. Recent work of Li et al. reconstruct 3D motion from dance videos on the Internet, however the data is not public. Further, using 3D motion reconstructed from monocular videos may not be reliable and lack accurate global 3D translation information. In this work we also reconstruct the 3D motion from 2D dance video, but from multi-view video sequences, which addresses these issues. While there are many large scale 3D motion capture datasets , mocap dataset of 3D dance is quite limited as it requires heavy instrumentation and expert dancers for capture. As such, many of these previous works operate on either small-scale or private motion capture datasets . We compare our proposed dataset with these public datasets in Table 1.

Beyond of the scope of human motion generation, our work is closely related to the research of using neural network on cross-modal sequence to sequence generation task. In natural language processing and computer vision, tasks like text to speech (TTS) and speech to gesture , image/video captioning (pixels to text) involve solving the cross-modal sequence to sequence generation problem. Initially, combination of CNNs and RNNs were prominent in approaching this problem. More recently, with the development of attention mechanism , transformer based networks achieve top performance for visual-text , visual-audio cross-modal sequence to sequence generation task. Our work explores audio to 3D motion in a transformer based architecture. While all cross-modal problems induce its own challenges, the problem of music to 3D dance is uniquely challenging in that there are many ways to dance to the same music and that the same dance choreography may be used for multiple music. We hope the proposed AIST++ dataset advances research in this relatively under-explored problem.

AIST++ Dataset

Data Collection We generate the proposed 3D motion dataset from an existing database called AIST Dance Database . AIST is only a collection of videos without any 3D information. Although it contains multi-view videos of dancers, these cameras are not calibrated, making 3D reconstruction of dancers a non-trivial effort. We recover the camera calibration parameters and the 3D human motion in terms of SMPL parameters. Please find the details of this algorithm in the Appendix. Although we adopt the best practices in reconstructing this data, no code base exist for this particular problem setup and running this pipeline on a large-scale video dataset requires non-trivial amount of compute and effort. We will make the 3D data and camera parameters publicly available, which allows the community to benchmark on this dataset on an equal footing.

Resulting AIST++ is a large-scale 3D human dance motion dataset that contains a wide variety of 3D motion paired with music. It has the following extra annotations for each frame:

99 views of camera intrinsic and extrinsic parameters;

1717 COCO-format human joint locations in both 2D and 3D;

2424 SMPL pose parameters along with the global scaling and translation.

Besides the above properties, AIST++ dataset also contains multi-view synchronized image data unlike prior 3D dance dataset, making it useful for other research directions such as 2D/3D pose estimation. To our knowledge, AIST++ is the largest 3D human dance dataset with 1408\mathbf{1408} sequences, 30\mathbf{30} subjects and 10\mathbf{10} dance genres with basic and advanced choreographies. See Table. 1 for comparison with other 3D motion and dance datasets. AIST++ is a complementary dataset to existing 3D motion dataset such as AMASS , which contains only 17.817.8 minutes of dance motions with no accompanying music.

Owing to the richness of AIST, AIST++ contains 10 dance genres: Old School (Break, Pop, Lock and Waack) and New School (Middle Hip-hop, LA-style Hip-hop, House, Krump, Street Jazz and Ballet Jazz). Please see the Appendix for more details and statistics. The motions are equally distributed among all dance genres, covering wide variety of music tempos denoted as beat per minute (BPM). Each genre of dance motions contains 85%85\% of basic choreographies and 15%15\% of advanced choreographies, in which the former ones are those basic short dancing movements while the latter ones are longer movements freely designed by the dancers. However, note that AIST is an instructional database and records multiple dancers dancing the same choreography for different music with varying BPM, a common practice in dance. This posits a unique challenge in cross-modal sequence-to-sequence generation. We carefully construct non-overlapping train and val subsets on AIST++ to make sure neither choreography nor music is shared across the subsets.

Music Conditioned 3D Dance Generation

Here we describe our approach towards the problem of music conditioned 3D dance generation. Specifically, given a 22-second seed sample of motion represented as X=(x1,…,xT)\mathbf{X}=(x_{1},\dots,x_{T}) and a longer conditioning music sequence represented as Y=(y1,…,yT′)\mathbf{Y}=({y}_{1},\dots,{y}_{T^{\prime}}), the problem is to generate a sequence of future motion X′=(xT+1,…,xT′){\mathbf{X^{\prime}}}=({x}_{T+1},\dots,{x}_{T^{\prime}}) from time step T+1T+1 to T′T^{\prime}, where T′≫TT^{\prime}\gg T.

Transformer is an attention based network widely applied in natural language processing. A basic transformer building block (shown in of Figure 3 (a)) has multiple layers with each layer composed of a multi-head attention-layer (Attn) followed by a feed forward layer (FF). The multi-head attention-layer embeds input sequence X\mathbf{X} into an internal representation often referred to as the context vector C\mathbf{C}. Specifically, the output of the attention layer, the context vector C\mathbf{C} is computed using the query vector Q{\mathbf{Q}} and the key K{\mathbf{K}} value V{\mathbf{V}} pair from input with or without a mask M{\mathbf{M}} via,

where DD is the number of channels in the attention layer and W\mathbf{W} are trainable weights. The design of the mask function is a key parameter in a transformer. In natural language generation, causal models such as GPT uses an upper triangular look-ahead mask M\mathbf{M} to enable causal attention where each token can only look at past inputs. This allows efficient inference at test time, since intermediate context vectors do not need to be recomputed, especially given the large context window in these models (2048). On the other hand, models like BERT employ full-attention for feature learning, but rarely are these models employed in an auto-regressive manner, due to its inefficiency at test time.

1 Full Attention Cross-Modal Transformer

We propose Full Attention Cross-Modal Transformer (FACT) model for the task of 3D dance motion generation. Given the seed motion X{\mathbf{X}} and audio features Y\mathbf{Y}, FACT first encodes these inputs using a motion transformer fmotf_{\text{mot}} and audio transformer faudiof_{\text{audio}} into motion and audio embeddings hx1:T{\mathbf{h}^{x}}_{1:T} and hy1:T′{\mathbf{h}^{y}}_{1:T^{\prime}} respectively. These are then concatenated and sent to a cross-modal transformer fcrossf_{\text{cross}}, which learns the correspondence between both modalities and generates NN future motion sequences X′\mathbf{X^{\prime}}, which is used to train the model in a self-supervised manner. All three transformers are jointly learned in an end-to-end manner. This process is illustrated in Figure 2. At test time, we apply this model in an auto-regressive framework, where we take the first predicted motion as the input of the next generation step and shift all conditioning by one.

FACT involves three key design choices that are critical for producing realistic 3D dance motion from music. First, all of the transformers use full-attention mask. We can still apply this model efficiently in an auto-regressive framework at test time, since our context window is not prohibitively large (240). The full-attention model is more expressive than the causal model because internal tokens have access to all inputs. Due to this full-attention design, we train our model to only predict the unseen future after the context window. In particular, we train our model to predict NN futures beyond the current input instead of just 11 future motion. This encourages the network to pay more attention to the temporal context, and we experimentally validate that this is a key factor training a model that does not suffer from motion freezing or diverging after a few generation steps. This attention design is in contrast to prior work that employ transformers for the task of 3D motion or dance generation , which applies GPT style causal transformer trained to predict the immediate next future token. We illustrate this difference in Figure 3 (b).

Lastly, we fuse the two embeddings early and employ a deep 12-layer cross-modal transformer module. This is in contrast to prior work that used a single MLP to combine the audio and motion embeddings , and we find that deep cross-modal module is essential for training a model that actually pays attention to the input music. This is particularly important as in dance, similar choreography can be used for multiple music. This also happens in AIST dataset, and we find that without a deep cross-modal module, the network is prone to ignoring the conditioning music. We experimentally validate this in Section 5.2.3.

Experiments

We first carefully validate the quality of our 3D motion reconstruction. Possible error sources that may affect the quality of our 3D reconstruction include inaccurate 2D keypoints detection and the estimated camera parameters. As there is no 3D ground-truth for AIST dataset, our validation here is based-on the observation that the re-projected 2D keypoints should be consistent with the predicted 2D keypoints which have high prediction confidence in each image. We use the 2D mean per joint position error MPJPE-2D, commonly used for 3D reconstruction quality measurement ) to evaluate the consistency between the predicted 2D keypoints and the reconstructed 3D keypoints along with the estimated camera parameters. Note we only consider 2D keypoints with prediction confidence over 0.5 to avoid noise. The MPJPE-2D of our entire dataset is 6.26.2 pixels on the 1920×10801920\times 1080 image resolution, and over 86%86\% of those has less than 1010 pixels of error. Besides, we also calculate the PCKh metric introduced in on our AIST++. The PCKh@0.5 on the whole set is 98.7%98.7\%, meaning the reconstructed 3D keypoints are highly consistent with the predicted 2D keypoints. Please refer to the Appendix for detailed analysis of MPJPE-2D and PCKh on AIST++.

2 Music Conditioned 3D Motion Generation

All the experiments in this paper are conducted on our AIST++ dataset, which to our knowledge is the largest dataset of this kind. We split AIST++ into train and test set, and report the performance on the test set only. We carefully split the dataset to make sure that the music and dance motion in the test set does not overlap with that in the train set. To build the test set, we first select one music piece from each of the 10 genres. Then for each music piece, we randomly select two dancers, each with two different choreographies paired with that music, resulting in total 4040 unique choreographies in the test set. The train set is built by excluding all test musics and test choreographies from AIST++, resulting in total 329329 unique choreographies in the train set. Note that in the test set we intentionally pick music pieces with different BPMs so that it covers all kinds of BPMs ranging from 8080 to 135135 in AIST++.

2.2 Quantitative Evaluation

In this section, we evaluate our proposed model FACT on the following aspects: (1) motion quality, (2) generation diversity and (3) motion-music correlation. Experiments results (shown in Table 2) show that our model out-performs state-of-the-art methods , on those criteria.

We also evaluate our model’s ability to generate diverse dance motions when given various input music compared with the baseline methods. Similar to the prior work , we calculate the average Euclidean distance in the feature space across 4040 generated motions on the AIST++ test set to measure the diversity. The motion diversity in the geometric feature space and in the kinetic feature space are noted as Distm\text{Dist}_{m}and Distk\text{Dist}_{k}, respectively. Table 2 shows that our method generates more diverse dance motions comparing to the baselines except Li et al. , which discretizes the motion, leading to discontinuous outputs that results in high Distk\text{Dist}_{k}. Our generated diverse motions are visualized in Figure 4.

Further, we evaluate how much the generated 3D motion correlates to the input music. As there is no well-designed metric to measure this property, we propose a novel metric, Beat Alignment Score (BeatAlign), to evaluate the motion-music correlation in terms of the similarity between the kinematic beats and music beats. The music beats are extracted using librosa and the kinematic beats are computed as the local minima of the kinetic velocity, as shown in Figure 5. The Beat Alignment Score is then defined as the average distance between every kinematic beat and its nearest music beat. Specifically, our Beat Alignment Score is defined as:

where Bx={tix}B^{x}=\{t_{i}^{x}\} is the kinematic beats, By={tjy}B^{y}=\{t_{j}^{y}\} is the music beats and σ\sigma is a parameter to normalize sequences with different FPS. We set σ=3\sigma=3 in all our experiments as the FPS of all our experiments sequences is 60. A similar metric Beat Hit Rate was introduced in , but this metric requires a dataset dependent handcrafted threshold to decide the alignment (“hit”) while ours directly measure the distances. This metric is explicitly designed to be uni-directional as dance motion does not necessarily have to match with every music beat. On the other hand, every kinetic beat is expected to have a corresponding music beat. To calibrate the results, we compute the correlation metrics on the entire AIST++ dataset (upper bound) and on the random-paired data (lower bound). As shown in Table 2, our generated motion is better correlated with the input music compared to the baselines. We also show one example in Figure 5 that the kinematic beats of our generated motion align well with the music beats. However, when comparing to the real data, all four methods including ours have a large space for improvement. This reflects that music-motion correlation is still a challenging problem.

2.3 Ablation Study

We conduct the following ablation experiments to study the effectiveness of our key design choices: Full-Attention Future-N supervision, and early cross-modal fusion. Please refer to our supplemental video for qualitative comparison. The effectiveness of different model architectures is measured quantitatively using the motion quality (FIDk\text{FID}_{k}, FIDg\text{FID}_{g}) and the music-motion correlation (BeatAlign) metrics, as shown in Table 4 and Table 3.

Here we dive deep into the attention mechanism and our future-N supervision scheme. We set up four different settings: causal-attention shift-by-1 supervision, and full-attention with future-{1,10,20}\{1,10,20\} supervision. Qualitatively, we find that the motion generated by the causal-attention with shift-by-1 supervision (as done in ) starts to freeze after several seconds (please see the supplemental video). Similar problem was reported in the results of . Quantitatively (shown in the Table 3), when using causal-attention shift-by-1 supervision, the FIDs are large meaning that the difference between generated and ground-truth motion sequences is substantial. For the full-attention with future-1 supervision setting, the results rapidly drift during long-range generation. However, when the model is supervised with 1010 or 2020 future frames, it pays more attention to the temporal context. Thus, it learns to generate good quality (non-freezing, non-drifting) long-range motion.

Here we investigate when to fuse the two input modalities. We conduct experiments in three settings, (1) No-Fusion: 14-layer motion transformer only; (2) Late-Fusion: 13-layer motion/audio transformer with 1-layer cross-modal transformer; (3) Early-Fusion: 2-layer motion/audio transformer with 12-layer cross-modal transformer. For fair comparison, we change the number of attention layers in the motion/audio transformer and the cross-modal transformer but keep the total number of the attention layers fixed. Table 4 shows that the early fusion between two input modalities is critical to generate motions that are well correlated with input music. Also we show in Figure 6 that Early-Fusion allows the cross-model transformer pays more attention to the music, while Late-Fusion tend to ignore the conditioning music. This also aligns with our intuition that the two modalities need to be fully fused for better cross-modal learning, as contrast to prior work that uses a single MLP to combine the audio and motion .

2.4 User Study

Finally, we perceptually evaluate the motion-music correlation with a user study to compare our method with the three baseline methods and the “random” baseline, which randomly combines AIST++ motion-music. (Refer to the Appendix for user study details.) In this study, each user is asked to watch 10 videos showing one of our results and one random counterpart, and answer the question “which person is dancing more to the music? LEFT or RIGHT” for each video. For user study on each of the four baselines, we invite 30 participants, ranging from professional dancers to people who rarely dance. We analyze the feedback and the results are: (1) 81%81\% of our generated dance motion is better than Li et al. ; (2) 71%71\% of our generated dance motion is better than Dancenet ; (3) 77%77\% of our generated dance motion is better than DanceRevolution ; (4) 75%75\% of the unpaired AIST++ dance motion is better than ours. Clearly we surpass the baselines in the user study. But because the “random” baseline consists of real advanced dance motions that are extremely expressive, participants are biased to prefer it over ours. However, quantitative metrics show that our generated dance is more aligned with music.

Conclusion and Discussion

In this paper, we present a cross-modal transformer-based neural network architecture that can not only learn the audio-motion correspondence but also can generate non-freezing high quality 3D motion sequences conditioned on music. We also construct the largest 3D human dance dataset: AIST++. This proposed, multi-view, multi-genre, cross-modal 3D motion dataset can not only help research in the conditional 3D motion generation research but also human understanding research in general. While our results shows a promising direction in this problem of music conditioned 3D motion generation, there are more to be explored. First, our approach is kinematic based and we do not reason about physical interactions between the dancer and the floor. Therefore the global translation can lead to artifacts such as foot sliding and floating. Second, our model is currently deterministic. Exploring how to generate multiple realistic dance per music is an exciting direction.

Acknowledgement

We thank Chen Sun, Austin Myers, Bryan Seybold and Abhijit Kundu for helpful discussions. We thank Emre Aksan and Jiaman Li for sharing their code. We also thank Kevin Murphy for the early attempts on this direction, as well as Peggy Chi and Pan Chen for the help on user study experiments.

References

Appendix

Appendix A AIST++ Dataset Details

We show the detailed statistics of our AIST++ dataset in Table 5. Thanks to the AIST Dance Video Database , our dataset contains in total 5.25.2-hour (1.11.1M frame, 14081408 sequences) of 3D dance motion accompanied with music. The dataset covers 10 dance genre (shown in Figure 8) and 6060 pieces of music. For each genre, there are 66 different pieces of music, ranging from 2929 seconds to 5454 seconds long, and from 8080 BPM to 130130 BPM (except for House genre which is 110110 BPM to 135135 BPM). Among those motion sequences for each genre, 120120 (8585%) of them are basic choreographies and 2121 (1515%) of them are advanced. Advanced choreographies are longer and more complicated dances improvised by the dancers. Note for the basic dance motion, dancers are asked to perform the same choreography on all the 66 pieces of music with different speed to follow different music BPMs. So the total unique choreographies in for each genre is 120/6+21=41120/6+21=41. In our experiments we split the AIST++ dataset such that there is no overlap between train and test for both music and choreographies (see Sec. 5.2.1 in the paper).

As described in Sec. 5.1 in the paper, we validate the quality of our reconstructed 3D motion by calculating the overall MPJPE-2D (in pixel) between the re-projected 2D keypoints and the detected 2D keypoints with high confidence (>0.5>0.5). We provide here the distribution of MPJPE-2D among all video sequences (Figure 9). Moreover, we also analyze the PCKh metric with various thresholds on the AIST++, which measures the consistency between the re-projected and detected 2D keypoints. Averaged PCKh@0.5 is 98.4% on all joints shows that our reconstructed 3D keypoints are highly consistent with the detected 2D keypoints.

Appendix B User Study Details

As mentioned in Sec. 5.2.5 in the main paper, we qualitatively compare our generated results with several baselines in a user study. Here we describe the details of this user study. Figure 11 shows the interface that we developed for this user study. We visualize the dance motion using stick-man and conduct side-by-side comparison between our generated results and the baseline methods. The left-right order is randomly shuffled for each video to make sure that the participants have absolutely no idea which is ours. Each video is 1010-second long, accompanied with the music. The question we ask each participant is “which person is dancing more to the music? LEFT or RIGHT”, and the answers are collected through a Google Form. At the end of this user study, we also have an exit survey to ask for the dance experience of the participants. There are two questions: “How many years have you been dancing?”, and “How often do you watch dance videos?”. Figure 10 shows that our participants ranges from professional dancers to people rarely dance, with majority with at least 1 year of dance experience.