Towards Diverse and Natural Scene-aware 3D Human Motion Synthesis
Jingbo Wang, Yu Rong, Jingyuan Liu, Sijie Yan, Dahua Lin, Bo Dai
Introduction
The capability of synthesizing long human motion sequences is essential for a number of real-world applications, such as virtual reality and robotics. Beyond early attempts that consider body movement synthesis in isolation , recent works begin to explore the influences of surrounding scenes on human motion synthesis for different actions. Limited by the 2D representation of scene context or the reliance on manually assigned interacting targets , these approaches mainly focus on modeling the body movements and fail to comprehensively investigate the inherent diversity of scene-aware human motions. In order to synthesize long-term human motions guided by the scene context and the target action sequence, we propose to model the inherent motion diversity across different granularities, each contributing to different aspects of human motion.
As shown in Figure LABEL:fig:teaser, the diversity of scene-aware human motions can be factorized into three levels, given the target action sequence (e.g. A man lies first. Then he sits in different places. At last, he stands somewhere.). Firstly, given the surrounding scene context and the target action sequence, there exists a distribution of valid locations to realize the actual human-scene interactions for each of these actions (e.g. We can sit on any chairs or beds and stand on the ground). Different locations can be sampled from the distribution and serve as the anchors of the whole synthesized motion sequence. Based on those anchors, we can then follow various paths to bridge them one by one. Finally, our body poses also differ from case to case when we move along the paths to connect all anchors. We demonstrate these three levels of diversity in Figure LABEL:fig:teaser. Existing attempts for scene-aware human motion synthesis only emphasize the last level of diversity (e.g. walking to the pre-defined object or position in the scene) via manually assigning the interaction locations and motion paths. Consequently, the importance of the scene semantics is substantially muted, as it mainly affects the distribution of valid interaction anchors and the distribution of valid motion paths. To faithfully capture the diversity of scene-aware human motions, we propose a novel three-stage motion synthesis framework, each stage of which is responsible for modeling one level of the aforementioned diversity.
For diverse human-scene interaction anchors, we design our pose placing framework for the given action sequence. Different from which only consider the influence of scene context, we first synthesize scene-agnostic poses according to the target action via a conditional VAE (CVAE) . Then we follow the practice of POSA to place these poses into the scene. To be specific, the 3D scene is uniformly split into a set of non-overlapping grids, each of which is associated with a validity score that measures its compatibility as a candidate for placing the poses. We make two modifications to the original placing method used by POSA First, we introduce the position relationship between poses with the same action label to enhance the placing diversity by avoiding them being placed to the nearby positions. Furthermore, we leverage another CVAE model as the placing refiner to produce diverse offsets for each discrete grid. Examples of generated anchors are depicted in Figure LABEL:fig:teaser (a).
To produce diverse obstacle-free motion paths following the sampled anchors, we employ an adapted A∗ algorithm over the discrete 3D grids as the path planner. The standard A∗ algorithm used by previous works only generates deterministic paths as they only consider collision between objects and distances to the target locations. To model the inherent diversity of motion paths, we amend the original algorithm with a trainable stochastic module learned in a data-driven manner. The new module, named Neural Mapper, can provide dynamic scene-conditioned probabilistic guidance to the A∗ algorithm, so that the algorithm can automatically produce diverse yet natural paths given the deterministic scenes and location anchors. We show several examples of generated diverse paths given the same start and end locations in Figure LABEL:fig:teaser (b).
Lastly, we propose a novel Transformer-based CVAE, called motion completion network, to synthesize diverse body movements guided by the paths generated in the previous step. Inspired by , we leverage Transformer as the basic architecture for synthesizing continuous and smooth motions. Differently, we focus on diverse motion completion of poses with long-term distance and different actions, rather than synthesize motions for the single action . Therefore, this motion completion network first generates diverse moving trajectories, loosely following the paths sampled by the aforementioned A∗ algorithm. The body poses are then produced by taking the scene contexts, action labels, human-scene interaction anchors, and synthesized trajectories as inputs.
To summarize our contributions: 1) We analyze the inherent diversity of the human motion and decompose it into three components, namely the diversity on human-scene interaction anchors, paths, and body poses. 2) We propose a novel three-stage framework to faithfully capture the diversities of scene-aware human motions. This framework can automatically synthesize human motions following these diversities with the condition action labels. Qualitative and quantitative results on datasets such as PROX demonstrate that our method significantly surpasses previous approaches in terms of diversity and naturalness. 3) In the proposed framework, we make several technique contributions for this task, including the action conditioned pose placing framework for generating diverse human-scene interaction anchors, Neural Mapper for planning diverse paths, and motion completion network for producing diverse and continuous motions. With our decomposition on motion diversity, these technique contributions can achieve our goal efficiently and effectively.
Related Works
Early works focus on synthesizing natural body poses and neglect the influences of other factors such as action and environments. Recent studies begin to explore the relationship between human motions with actions and scene contexts. Recent works generates human pose sequences with a CVAE model based on the given action labels. ACTOR builds up a transformer based on CAVE to synthesize human motion sequence directly from the given action label. Cao et al. propose a three-stage motion prediction method that can predict different human motions with different destinations. Wang et al. extend CSGN to explore the influence of 2D scene contexts on human motion synthesis. Wang et al. build up a framework to synthesize human motions in the 3D scene controlled by the given pairs of begin-end points. SAMP extends to use 3D oriented objects to facilitate the synthesis of human motions with specific action labels. Besides, a planning module is incorporated into their framework to find obstacle-free paths.
The limitations of previous works mainly lie in their reliance on predefined objects or positions, which constrain their ability to explore the inherent interaction diversity of synthesized scene-aware human motions. In this work, we aim to overcome the limitations of the previous works and synthesize diverse motions guided by target action sequences in the given scenes. To achieve this, we first synthesize diverse human-interaction anchors, which interacts with different objects in the scene. Then we plan diverse paths and complete diverse body movements between these anchors.
Motion prediction is closely related to our problem. Different from the motion synthesis, the goal of this task is to predict human dynamics in the future with the given moving orientations or previous motions. Martinez et al. and ERD proposed motion prediction framework based on the Seq2Seq model . Ac-LSTM mixes synthesized frames and observed frames to enhance the capability of LSTM during the training stage. The graph convolution network is widely used in recent motion prediction . These methods model dynamic spatial and temporal relationships between the obvious frames and the future frames. Different from these works, our goal is to synthesize motions without prior knowledge of the previous motions.
Methodology
The overview of our framework is depicted in Figure 1. We aim to solve this challenging problem in a hierarchical manner via exploiting the inherent properties of the scene-aware human motions. Our framework first generates diverse human-scene interaction anchors for the given actions. In this step, the framework first produces scene-agnostic poses corresponding to the action labels and then places these poses into the scene considering the compatibility between the synthesized poses and the scene. In the next step, we leverage a path planning module to produce diverse obstacle-free paths under the guidance of the synthesized anchors from the first step. Finally, a motion completion module is adopted to synthesize diverse body movements that fill in the missing motions between consecutive anchors while roughly following the planned paths from the second step. In the following, we introduce our modules in detail.
2 Human-Scene Interaction Anchor Synthesis
We first synthesize human-scene interaction anchors. Unlike previous works that only condition human motion synthesis on the scene context, we use action labels describing interaction types as an additional condition. To be specific, we first synthesize scene-agnostic poses corresponding to the action labels. Then we follow the practice of POSA with several modifications to diversely place the synthesized poses into the scene. This design affords us more control over the final synthesized motions.
As shown in Figure 2 (a), we follow the standard CVAE framework to synthesize scene-agnostic poses with the target action . To be specific, we first sample noises from the prior Gaussian distribution and encode them with a fully-connected layer. Then, we use the one-hot vector to represent the action condition and encode it with another fully-connected layer. These two features are added up and then served as additional input besides the noise. The model outputs the synthesized pose , which is directly used as the body pose for the anchor for -th anchor.
In this step, we place the scene-agnostic poses into the given scene. There are two aspects to be taken into consideration in this step. The first one is how to place poses to locations with compatible scene structure and interaction semantics. The other one is how to efficiently find multiple reasonable locations given a pose.
Therefore, we first select our placing candidates following the practice of POSA . Specifically, each candidate consists of a translation parameter and an orientation parameter for the anchor . We split the given scene into uniform non-overlapping discrete girds as translation candidates. For each discrete gird, we then uniformly sample eight different orientations that are parallel with the ground plane to build orientation candidates. Each translation candidate is paired with one of its associated orientations to form one placing candidate. For each scene-agnostic pose , we then rank all the placing candidates by their compatibility scores with the pose, which is proposed by that considers both the affordance and penetration. An intuitive idea is to select the candidate with the best score. However, our empirical study shows that candidates with the same action labels tend to be located close to each other since the same action usually shares similar physical and semantic structures, as shown in the first row of Figure 6. To increase the placing diversity, we introduce an additional penalty on the locations that have been occupied by anchors with the same action labels. As shown in Figure 6, this new penalty helps produce more diverse placing candidates for similar poses. In this way, we can sample an initial placing candidate for each pose . The initial anchor is then constructed subsequently.
In practice, we further adopt another sub-module called Place Refiner to improve the micro diversity of the placing candidates. Place Refiner is implemented as a CVAE model that takes the noise of , the scene context encoded by the PointNet and the initial anchor as the input. It outputs the offset to the sampled position and orientation . The final position and orientation are obtained as and . The framework of Place Refiner is depicted in Figure 2 (b).
3 Diverse Path Planning
In this step, we discuss how to generate diverse obstacle-free paths from human-scene interaction anchors. Previous works such as SMAP often use standard searching for this purpose. The algorithm tends to generate deterministic shortest path for practice. However, humans usually move stochastically in the given scene. To reflect the diversity of human path planning, we incorporate the standard algorithm with scene-aware random information concerning the diversity of human motion.
To begin with, we first discuss how to apply the standard algorithm into our scenario. We first divide the whole 3D scene into the same set of non-overlapping discrete grids as in Section 3.2. We then define and calculate the cost function for each grid in the algorithm as:
where measures the cost for moving from the beginning point to grid , and measures the cost between grid and the target grid, during searching points as the next step for in the neigbourhood . To ensure obstacle-free paths, we further filter out inaccessible grids that might have collisions with the human body. The collisions are detected via placing a cylinder model that approximates the volume of a human at each grid. We show an example in the right of the Figure 3 (a), where red stands for valid and blue stands for invalid. After calculating for each grid and excluding invalid grids, an obstacle-free path connecting two human-scene interaction anchors can thus be obtained using the standard algorithm. It is worth noting that the path obtained in this manner is deterministic and fixed for the same pair of two human-scene interaction anchors.
An intuitive solution to incorporate diversity in path planning is appending the cost function defined in Equation (2) with a random noise term. This strategy sounds feasible but fails to generate reasonable paths, which is demonstrated by the examples shown in the top two rows of Figure 4. To this end, we replace the random noise term with a controllable signal produced by another CVAE, referred as Neural Mapper. For each grid , Neural Mapper takes sampled latent code and the local scene context feature obtained via BPS as the input and outputs the feasibility score for each neighbor grid . Based on the Neural Mapper, the cost function is updated as:
The score of indicates the feasibility of moving from the current grid to this adjacent one so that we can build the cost as to reflect the moving guidance by our Neural Mapper. The Neural Mapper is trained in a data-driven manner thus it can help the algorithm to generate diverse and reasonable paths. We show several examples produced by Neural Mapper in the bottom row of Figure 4.
Without complex manually designed conditions and constraints, the proposed Neural Mapper equips the algorithm with the ability to find diverse obstacle-free paths in a flexible and generalizable way. In Neural Mapper, we can easily change the characteristics of sampled paths by restricting the latent codes, without hurting their naturalness and coherency.
4 Motion Completion
With the obstacle-free path obtained from path planning, we are now ready to complete the missing motions between consecutive human-scene interaction anchors. As shown in Figure 5, our motion completion network consists of two components, namely Path Refiner and Motion Synthesizer. Although paths for human-scene interaction anchors are planned in Section 3.3, this Path Refiner accounts for the gap between diverse real human motions and the path formed by straight lines between the discrete grids. Both the Path Refiner and the Motion Synthesizer follow the CVAE framework. Specially, we apply Transformer as the basic architecture for both the encoder and decoder of these two networks to synthesize continuous and smooth motions. Our motion completion network simultaneously synthesizes frame paths and body poses as , instead of one-by-one in an auto-regressive manner .
For Path Refiner, we take the scene context encoded by PointNet to synthesize the refined path. The refined path is composed of pairs of the translation and orientation sequence. Following , we introduce the positional encoding formed from sinusoidal functions which take time steps as input to ensure the continuity and smoothness of the refined path. Moreover, we leverage one more positional encoding obtained from the planned path by encoding each step of the planned path in Section 3.3 by a fully connected layer, to ensure the refined path is still in the obstacle-free regions.The effectiveness of this additional positional encoding is illustrated in Section 4.2, where our Path Refiner further improves the diversity of synthesized motion.
The motion sequence with body poses is completed by our Motion Synthesizer. Same as the Path Refiner, we take the scene context encoded by the PointNet as the condition to complete these scene-aware motions. The completed motions should fulfill two requirements, namely matching the paths produced by the Path Refiner and naturally transforming between the given human-scene interaction anchors. To achieve this, the Motion Synthesizer at first takes the refined path as additional position encoding to guide motion synthesis, similar to the practice of Path Refiner. For the motion transformation, we need to model the relationship between the given two human-scene interaction anchors and the potential motions that could be completed in our Motion Synthesizer. Inspired by the practice of action token in , which helps the transformer decoder to build up the relationship between synthesized motions and the given action, we encode the action labels and poses of human-scene interaction anchors by additional fully connected layers as learnable tokens and add them to the beginning and ending of the positional encoding respectively. With these tokens, our Motion Synthesizer can directly build up this relationship between and synthesize reasonable and smooth motions. Following these two steps, the motion completion network can generate natural motions for the given human-scene interaction anchors following the planned path.
Experiments
In this section, we first illustrate our experiment settings and metrics for evaluation. Then we discuss the effectiveness of the proposed framework. At last, we demonstrate the qualitative results in different scenes.
Following , we train our framework on PROX dataset . We manually label the motions in PROX with action labels (i.e. sit, lie, stand, walk, and squat) as the action condition. We do not conduct experiments on GTA-IM and SAMP as they do not provide reconstructed 3D real-world scenes. For the fair comparison, we follow the split of the train and test set as and synthesize human motions on the unseen scenes during training. To demonstrate the generalization ability of the proposed framework, we further evaluate it on Matterport3D , which provides large-scale reconstructed 3D scenes. Please be noted that our framework does not leverage Matterport3D for training.
We measure the diversity on synthesized human motions in three aspects, namely human-scene interaction anchors, planned paths, and completed motions. To evaluate the diversity of the human-scene interacting anchors, we preform K-Means () clustering on the synthesized human-scene interaction anchors, following . To be specific, we consider two types of the clusters. The first one considers all parameters . The second one only considers translation and orientation . The diversity is measured as the entropy of the cluster sizes and the average distances between the clusters center and the samples belonging to it. We evaluate path diversity by the standard deviation (STD) of distances between the paths from Neural Mapper and the ones from the standard . To fairly compare with previous works that manually assign anchors or target objects , we evaluate the diversity of the synthesized human motions with the fixed human-scene interaction anchors. To measure the ability of our motion completion network in generating diverse results, we do not introduce the diverse sampling strategies as . Following , we calculate the Average Pairwise Distance (APD) on the SMPL-X parameters of synthesized motions to measure its diversity.
We evaluate the naturalness of synthesized motions via user study and the physical plausibility. We ask users to compare our results against other methods and score them from 1 to 5 (the higher the better) as the results. Besides, we involve the non-collision score and contact score to measure the physical plausibility of synthesized motions between the 3D scenes.
To evaluate the quality of the whole synthesized motions, we follow to calculate the Frechét Distance (FD) between synthesized motions and ground-truth motions. This distance is computed using the parameters of each frame.
2 Experimental Results
In this section, we show quantitative results on PROX dataset. The quantitative results on Matterport3D dataset are included in our supplementary materials. We also show qualitative results on these two datasets in this section.
We first show the diversity of synthesized human-scene interaction anchors in Table 5. For this evaluation, we sample 100 poses for each action and employ the placing strategy in POSA as our baseline. As shown in the table, the position related sampling process and the Pose Refiner can both improve the diversity of interaction anchors. We further show the effectiveness of these two process in Figure 6. Examples shown in this figure are all generated from the action label “sit”. (a) shows the first placed poses. Without the position related sampling, (a) and (b), which have the same action label, are placed close to each other. (c) and (d) are the generated anchors using our position related sampling. It is revealed that they interact with different objects in the scene. The second row demonstrates the result pairs ((d) V.S. (e) and (f) V.S. (g)), which are produced by our Place Refiner works and optimization post-process . Using the diverse translations and orientations as initialization states, the optimization algorithm can produce diverse optimal solutions. In the supplemental material, we show comparison with previous works that are extended to synthesize specific actions using our action condition.
We compare the diversity and naturalness of the planned path. For the evaluation of diversity, We compute the standard deviation of the distances between sampled paths and the paths produced by standard . In practice, we calculate distances between the discrete points on the paths, which are set as 1/6, 1/3, 1/2, 2/3, and 5/6 of the sample paths. We also show the results of the user study to reflect the naturalness of the planned paths. The number of samples is set to be for each method. The evaluation results are shown in Table 2. Standard algorithm, which is used in only produces the deterministic path for practice while the proposed Neural Mapper can generate diverse paths with similar naturalness as the standard A does. On the other hand, two methods using random noises cannot produce natural results, although they generate more diverse paths than ours. Similar results are also demonstrated in Figure 4. Compared with the methods using random noises, Neural Mapper can provide consistent and reasonable guidance for the similar local scene context to avoid unnatural moving. Moreover, Neural Mapper can cope with other manual constraints such as avoiding passing a certain region. We will discuss it in our supplementary materials.
In this subsection, we compare with other advances on scene-aware motion synthesis . We use the official model of trained on PROX dataset and extend and to PROX for fair comparison. In Table 3, we first compare against these methods using FD and APD for the motion quality and diversity. Firstly, for the comparison of FD, we sample 500 motion sequences which begin with the same action, as well as 500 motions which finish the same action. Besides, for the comparison of APD, we sample pairs of human-scene interaction anchors and synthesize motions for each pair. It is revealed that our method achieves the best results against other methods. All the comparison results show that our method can synthesize more diverse and natural motions than other methods do. In the supplementary material, we first compare more naturalness results between these methods, e.g. physical compatibility and user study. Then we further discuss our design choice on the motion completion network, including the effectiveness of the Path Refiner and the positional encoding based on planned paths.
We show more qualitative results of the proposed method on the PROX and Matterport3D in Figure 7. We show all three aspects of the synthesized scene-aware motions, including the human-scene interaction anchors, diverse planned paths, and completed motions. These results demonstrate that our framework can synthesize diverse human motions in the specific scene contexts for the given target action sequence. More qualitative results are included in the supplemental material.
Conclusion
In this paper, we focus on synthesizing diverse and natural human motions in the given scene environment guided by target action sequence. We decompose the diversity of scene-aware human motions into three levels, namely the diversity of action-conditioned human-scene interactions, the diversity of obstacle-free paths, and the diversity of body movements. To comprehensively leverage the inherent diversity of human motions, we propose a novel hierarchy framework with each component accounting for each level of the diversity. Thanks to the effective decomposition of diversity and elaborated designed modules, our framework is able to produce various vivid human motions in the scene across all three levels with improved efficiency and generality. Furthermore, the factorized design of our framework make it can be easily incorporated into other human motion synthesizing frameworks.
Supplementary Materials
In this section, we first discuss the detailed structures of the CVAE models used in our framework. Then we discuss the training and inference details of these models.
1.2 Training and Inference Details.
Firstly, we show how to train the CVAE model for scene-agnostic pose synthesis in Section 3.2. As the standard VAE model, the training objective consists of two parts. The first one is the reconstruction loss between the reconstructed human poses and the input human pose. The other objective is Kullback-Leibler (KL) Divergence between the Gaussian Distribution , where and are predicted by the encoder, and the standard Gaussian Distribution .
The Place Refiner takes the placed body poses, and the scene contexts encoded by the PointNet as inputs and predict the offset for the sampled . To train this network, we first build up the discrete candidates following the same procedure of scene-conditioned anchor placing in Section 3.2 for each scene in our training set. For practice, we split each scene into non-overlapping discrete grids uniformly as translation candidates and then uniformly sample eight different orientations paralleling with the ground plane as the orientation candidate. Each pose in the given scene is neighbor to four-position candidates. We assign each pose in the training set to one randomly sampled neighbor position candidate and one orientation candidate of this position candidate. Then we predict the offset from these candidates to the original translation and orientation for this pose. Similar as , the training objective is the reconstruction loss and the KL-Divergence.
We train our Path Refiner and Motion Synthesizer together in an end-to-end manner. Both two models synthesize frames of paths or motions. Similarly, the training objective consists of the reconstruction loss on the synthesized paths or motions, as well as the KL-Divergence. Specially, we do not use the planned path obtained from Section 3.3 in training the Path Refiner. Instead, we use the directions pointing from the beginning to the ending point of the motion as the planned path for practice. The reason mainly lies in two aspects. The first one is that repeat running of path planning module to obtain planned paths is not efficient in training. The second one is that the shortest path from the planning module is similar to the straight line in short-term motions.
During the inference stage, we find that directly synthesizing motions from two consecutive anchors lead to unstable results. We believe it is majorly caused by the variance of the lengths of the planned paths. To resolve this, we first split the planned paths into several pieces with equal length. Each split point is then assigned with an intermediate status action label. Using the new action labels, the intermediate anchors can be produced following the same method as placing human-scene interaction anchors in Section 3.2. Given the new anchors and the split paths, our motion completion network synthesizes human motions for each piece and then connects them together as the integrated motion. For practice, we insert the motions with random poses conditioned on “walking”, “standing”, and “ squatting” action.
1.3 Optimization.
2 Experiments
Then we compare the naturalness of these methods in Table 6. For the physical plausibility, we use the same motion as the comparison in Table.3 of our paper. For user study, we randomly sample motions with , , and different target actions. All the comparison results show that our method can synthesize more natural motions than other methods do. Especially, our method achieves better results without the optimization post-process, because of the guidance from planned obstacle-free paths. The comparison between the last two rows shows that the guidance of the proposed Path Refiner is advantageous in synthesizing natural motions.
We show the quantitative results on Matterport3D dataset . The sampling strategies are the same as our experiments on PROX dataset. We first perform K-Means () and evaluate the obtained human-scene interaction anchors on Matterport3D with the entropy of cluster sizes and the average distance between the cluster center and the samples belong to it.
As shown in Table 5, our method enhance the diversity of the anchors for motion synthesis. Besides, we evaluate the synthesized motion on Matterport3D via the APD, Non-Collision score and Contact score. The results are listed in Table 6. It is revealed that our method can synthesize better results than previous methods with better diversity and physical plausibility. Besides, the Path Refiner still can improve the diversity and naturalness on this dataset.
3 Further Discussion
In this section, we first discuss the reason for using the human-centric paradigm, for human-scene interaction anchors. The human-centric paradigm means we place the sampled poses to the positions which match the physical structure of these poses. Then we show how to use our Neural Mapper to work with other manually set constraints. At last, we show the influence of the planned path on motion synthesis.
Previous works of synthesizing human-scene interaction anchors aim to explore the influence of scene context to place human pose in the given scene and neglect the action labels. Intuitively, we can incorporate these action labels as an additional condition and incorporate them into their frameworks to synthesize poses. However, as shown in Figure 8, simply extending the previous works cannot guarantee to synthesize the physically plausible poses with the given actions and scene contexts. We believe it is due to the reason that these methods do not build up the relationship between the action and the scene context explicitly. For example, method directly uses the pooled 2D image features as the condition and ignores the relationship between the spatial information and the action. Another method first samples different positions to build up the BPS and then synthesize different poses. However, the poses for each action have their specific physical structure and match different scene structures. It is difficult to find suitable places for the poses conditioned on the given action label, as shown in the Figure 8. Instead, the human-centric paradigm proposed by us can effectively leverage the explicit relationship between the synthesized 3D human and the scene structure (e.g. physical and semantic structures) and thus makes the whole placing process more controllable.
Several failure cases generated from randomly sampling intermediate points are included in Figure 9. In the first row, when sampled intermediate points and the ending points are obstructed, the original algorithm can not find paths for these points. We adjust by allowing to search paths in the obstacle regions , and only produces impractical paths crossing the table as the first row of Figure 9. In the second row, random sampled points can also lead to unnatural zigzag paths. One may argue that these failures can be avoided via complex constraints used in previous methods . However, the proposed Neural Mapper provides an automatic and data-driven way to embed semantic information into natural and diverse path planning, without complex constraints. Besides, our Neural Mapper also can work with manually set constraints, such as avoiding passing a certain region. We show the planned results in Figure 10. It is revealed that our method can still produce natural and diverse paths under such constraints.
As shown in Figure 11, we show the effectiveness of using planned paths in the procedure of motion synthesis. Firstly, without the additional positional encoding from the planned path, the synthesized motion can not follow the planned path and penetrate to the table, as shown in Figure 11 (a). Besides, we find that both the translation and orientation for the planned path are also crucial for motion synthesis. As shown in Figure 11 (c), the Path Refiner synthesizes unnatural orientations for human motions without encoding the orientation of the planned path into the positional encoding as Section 3.4. Instead, as shown in Figure 11 (b) and (d), our method can synthesize natural human motion with the translation and orientation of planned path as the additional positional encoding for our Path Refiner.
Acknowledgement
This study is supported under the General Research Fund (GRF) of Hong Kong (No.,14205719), the RIE2020 Industry Alignment Fund–Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s).