COUCH: Towards Controllable Human-Chair Interactions
Xiaohan Zhang, Bharat Lal Bhatnagar, Vladimir Guzov, Sebastian Starke, Gerard Pons-Moll
Introduction
To synthesize realistic virtual humans which can achieve goals and act upon the environment, reasoning about the interactions and in turn contacts, is necessary. Reaching a goal, like sitting on a chair, is often preceded by intentional contact with the hand to support the body. In this work we investigate a motion synthesis method which exploits predictable contact to achieve more control and diversity over the animations.
Although most applications in VR/AR, digital content creation and robotics require synthesizing motion within the environment, it is not considered in the majority of works in human motion synthesis . Recent work does take the environment into account but is limited to synthesizing static poses . Synthesising dynamic human motion coherent with the environment is a substantially harder task and recent works show promising results. However, these methods do not reason about intentional contacts with the environment, and can not be controlled with user provided contacts.
Thus, in this work, we investigate a new problem: synthesizing human motion conditioned on contact positions on the object to allow for controllable movement variations. As a testbed to investigate this new problem, we focus on human-chair interactions as one of the most common actions, which are of crucial importance for ergonomics, avatars in virtual reality or video game animation. Contact-driven motion synthesis is a more challenging learning problem compared to conditioning only on coarse object geometry . First, the human needs to approach the chair differently depending on the contacts, regardless of the starting position, walking around it if necessary. Second, a chair can be approached and contacted in many different ways; we can directly sit without using our hands, or we can first support the body weight using the left/right or both hands with different parts of the chair. Furthermore, individual styled free-interactions can be modelled such as leaning back, stretching legs, using hands to support the head, and so on.
Contact driven motion allows for providing more detailed instructions to the virtual human such as approaching to sit on the chair, while supporting the body with the left hand and placing it on the armrest, as illustrated in Figure 1. Given the contact and the goal, the full-body needs to coordinate at run-time to achieve a plausible sequence of pose transitions. Intuitively, this emulates our planning of motion as real humans: we plan in terms of goals and intermediate object contacts to reach; the full-body then moves to follow such desired trajectories.
To this end, we propose COUCH, a method for controllable contact driven human-chair interactions, which is composed of two core components: 1) ControlNet is responsible for motion planning by predicting the future control signal of the body limbs which guides the future body movement. Our spatial-temporal control signal consists of dynamic trajectories of the hands towards the contact points and the local phase, an auxiliary continuous variable that encodes the temporal information of a limb during a particular movement (e.g. the left hand reaching an armrest). 2) PoseNet conditions on the predicted control signal to synthesise motion that follows the dynamic trajectories, ensuring the contact point is reached. At runtime, COUCH can operate in two modes. First, in an interactive mode where the user specifies the desired contact points on the target object. Second, in a generative mode where COUCH can automatically sample diverse intentional contacts on the object with a novel neural network called ContactNet. Training and evaluating COUCH calls for a dataset of rich and accurate human chair interactions. Existing interaction 3D datasets are captured with Inertial Sensors, and hence do not capture the real body motion and the contacts with the real chair geometry – instead synthetic chairs are fit to the avatar as post-process in those works. Hence, to jointly capture real human-chair interactions with fine-grained contacts, we fit the SMPL model and scanned chair models to data obtained from multiple Kinects and IMUs. The dataset (the COUCH dataset, Table 4) consists of 3 hours (over 500 sequences) of motion capture (MoCap) on human-chair interactions. Moreover it features multiple subjects, accurately captured contacts with registered chairs, and annotation on the type of hand contact. Our experiments demonstrate that COUCH runs in real-time at 30 fps, COUCH generalizes across chairs of varied geometry, and different starting positions relative to the chair. Compared to SoTA models (trained on the same data) adapted to incorporate contacts, our method significantly outperforms them in terms of control by improving the average contact distance by 55.
The contributions of our work can be summarized as follows:
We propose COUCH, the first method for synthesizing controllable contact-based human-chair interactions. Given the same input control, the COUCH model can achieve diverse sitting motions. By specifying different control signals, the user enables control over the style of interaction with the object. Results show our method outperforms the state of the art both qualitatively and quantitatively.
To train COUCH, we captured a large-scale MoCap dataset consisting of 3 hours (over 500 sequences) of human interacting with chairs different styles of sitting and free movements. The dataset features multiple subject, real chair geometry, accurately annotated hand contacts, and RGB-D images.
To stimulate further research in controllable synthesis of human motion, we will release the COUCH model and dataset.
Related Work
Synthesizing realistic human motion has drawn much attention from the computer vision and graphics communities. However, many methods do not take the scene into account. Existing methods on short (1 sec) and long (1 min) term 3D human motion prediction aim to produce realistic-looking human motion (typically walking and its variants). There also exists work on conditional motion generation based on music . These methods have two major limitations, i) except for work that use generative models , these methods are largely deterministic and cannot be used to generate diverse motions and ii) this body of work is unfortunately agnostic to scene geometry , which is critical to model human scene interactions. Our method on the other hand can generate realistic motion and interactions, taking into account the 3D scene.
Affordance and Static Scene Interactions.
Although the focus of our work is to model human-scene interactions over time, we find the works predicting static affordances in a 3D scene relevant. This branch of work aims at predicting static humans in a scene that satisfies the scene constraints. More recently there have been attempts to model fine-grained interactions (contacts) between the hand and objects .
The aforementioned methods focus on predicting static humans / human poses that satisfy the scene constraints in case of affordances or grasping an object in case of hand-object interactions. But these methods cannot produce a full sequence of human motion and interaction with the scene. Ours is the first approach that can model fine-grained interactions (contacts) between an object in the scene and the human.
Dynamic Scene Interactions.
Although various algorithms have been proposed for scene-agnostic motion prediction, affordance prediction as well as the synthesis of static human-scene interaction, generating dynamic human-scene interactions is less explored. Recent advances include predicting human motion from scene images , and using a semantic graph and RNN to predict human and object movements . More recently, Wang et al. introduce a hierarchical framework that generates ‘in-between’ locations and poses on the scene and interpolates between the goal poses. However, it requires a carefully tuned post-optimization step over the full motion synthesis to solve the discontinuity of motion between sub-goals and to achieve robust foot contacts with the scene. Alternatively, Chao et al. use a reinforcement learning based approach by training a meta controller to coordinate sub-controllers to complete a sitting task. An important category of human-scene interaction involves performing locomotion on uneven terrains. The Phase-functioned Neural Network first introduced the use of external phase variables to represent the state of the motion cycle. Zhang et al. applies same concept for quadruped motion and further incorporates a gating network that segments the locomotion modes based on foot velocities. Both works show impressive results thanks to the mixture of experts styled architectures.
The most relevant work to us, are the Neural State Machine (NSM) and SAMP . While NSM is a powerful method and models human-scene interactions such as sitting, carrying boxes and opening doors, it does not generate motion variations for the same task and object geometry. SAMP predicts diverse goal locations in the scene for the virtual human, which is then used to drive the motion generation. Our work takes inspiration from these works, but it is demonstrated qualitatively and quantitatively from our experiments that neither of the work enables control over the style of interaction (Section 5.2). Our work focus on controllable, fine-grained interactions given on contacts on the object. To the best of our knowledge, no previous work has tackled the problem of generating controllable human-chair interactions.
The COUCH Dataset
Large scale datasets of human motion such as AMASS and H3.6m have allowed us to build models of human motion. Unfortunately these datasets only contain sequences of 3D poses but no information about the 3D environment, making these datasets unsuitable for learning human interactions. On the other hand, datasets containing human-object interactions are either restricted to just hands , contain only static human poses without any motion, or have little variation in motion .
We present a multi-subject dataset of human chair interactions. Our dataset consists of 6 different subjects interacting with chairs with over 500 motion sequences. We collect our dataset using 17 wearable Inertial Measurement Units (XSens) , from which we obtain high-quality pose sequences in SMPL format using Unity . The total capture length is 3 hours.
Motion capture with marker-based capture systems is restrictive to capturing human-object interactions because markers often get occluded during the interactions leading to inaccurate tracking. IMU-based systems are prevalent for large-scale motion capture, however, the error from its calibration can lower the accuracy of the motion. We propose to combine IMUs with Kinect-based capture system as an efficient trade-off between scalability and accuracy. Our capture system is lightweight and can be generalized to capture many human scene interactions. We use the SMPL registration method similar to to obtain SMPL fits for our data. The dataset is captured in four different indoor scenes. The average fitting error for the SMPL human model, and the chair scans to the point clouds from the Kinects are 3.12 cm and 1.70 cm, respectively (in Chamfer distance). More details about data capture can be found in supp. mat.
Diversity on Starting Points and Styles. We capture people approaching the chairs from different starting points surrounding the chairs. Each subject then performs different styles of interactions with the chairs during sitting. This includes, one hand touching the arm of the chair, both hands touching the armrests of the chair, one hand touching the sitting plane of the chair before sitting down, and no hand contacts. It also includes free interactions such as crossing legs or leaning forward and backward on the chairs. To ensure the naturalness of motion, each subject is only provided with high-level instruction before capturing each sequence and was asked to perform their styles freely. Annotations of the direction of the starting points relative to the chair as well as the type of hand contact are included in the dataset.
Objects. Our dataset contains three different chair models that vary in terms of their shapes, as well as a sofa. The objects are 3D scanned before registering into the Kinect captured point clouds. To generalize the synthesized motion to unseen objects, we perform data augmentation as in .
Contacts. Studying contact-conditioned interaction calls for accurate contacts to be annotated in the dataset. Since we capture both the body motion and the object pose, it is possible to capture contacts between the body and the object. We detect the contacts of five key joints of the virtual human skeleton, which are the pelvis, hands, and feet. We then augment our data by randomly switching or scaling the object at each frame. The data augmentation is performed on 30 instances from ShapeNet over categories of straight chairs, chairs with arms, and sofas. At every frame, we project the contacts detected from the ground truth data to the new object, and apply full-body inverse kinematics to recompute the pose such that the contacts are consistent, keeping the original context of the motion.
Method
We address the problem of synthesising 3D human motion that is natural and satisfies environmental geometry constraints and user-defined contacts with the chair. COUCH allows fine-grained control over how the human interacts with the chair. At run-time, our model operates in two modes. First, a generative mode where COUCH can automatically sample diverse intentional contacts on the object with our proposed generative model. Second, an interactive mode where the user specifies the desired contact points on the target object. The input to our method is the current character pose, the target chair geometry as well the target contacts for the hands that need to be met. Our method takes these inputs and predicts the future poses that satisfy the desired contacts auto-regressively.
Synthesising natural human motion subject to environmental constraints is a challenging task , particularly when also satisfying a set of desired contacts. To this end, we first divide our motion synthesis task into motion planning and motion prediction. We derive our intuition from the way humans execute complex interactions e.g., to sit on the chair, we first prepare a mental model of how we will sit (place a hand on the arm-rest and sit, place a hand on the sitting plane and sit or just sit without using the hands etc.) and then we move our bodies accordingly. We propose two neural networks ControlNet , and PoseNet , for motion planning and motion prediction respectively. Furthermore, we observe that it is useful to perform detailed hand motion planning only when we are close to the chair right before sitting and not when we are far off. Thus, we decompose the motion synthesis into approaching and sitting. The approaching motion can be generated directly with PoseNet but both networks are required for sitting, ControlNet and PoseNet, for generating the sitting motion that satisfies the given contacts.
2 Motion Planning with ControlNet
ControlNet is the core of our method and plays an important role in motion planning, that is predicting the future control signals of the key joints which are used to guide the body motions. At a high level, the contact-aware control signal contains the local phases and the future locations of the key joints (in our case, the two hands). The local phase is an auxiliary variable that encodes the temporal alignment of motion for each of the hands and prepares for a future contact to be made. When the virtual human is ready to make contact with the chair, and at the beginning of the hand movement, the local phase is equal to 0, and it gradually reaches the value 1 as the hand comes closer to the contact. The hand trajectory, on the other hand, encodes the spatial relationship between the hand joint and the given contact location.
The ControlNet is trained to minimize the following MSE loss on the future hand trajectories and the local phase, which is formulated as follows:
3 Motion Synthesis with PoseNet
ControlNet generates important signals that guide the motion of the person such that user-defined contacts are satisfied. To this end, we train PoseNet , that takes as input the control signals predicted by the ControlNet along with the 3D scene and motion in the past and predicts full body motion.
4 Contact Generation with ContactNet
From the user’s perspective, it is useful to automatically generate plausible contact points on any given chairs. To this reason, we propose ContactNet. The network adopts a conditional variational auto-encoder architecture (cVAE) which encodes the chair geometry introduced in Section 4.3 and the contact positions to a latent vector . The decoder of the network then reconstructs the hand contacts . Note, the position of each voxel in the scene representation in this case is computed relative to the center of the chair instead of the character’s root. During training, the network is trained to minimize the following loss,
5 Decomposition of Approaching and Sitting Motion
Detailed hand motion planning is only required when the human model is close enough to the chair right before sitting as sitting requires synthesizing more precise full-body motion, especially for the hands, such that the person makes the desired contacts and sits on the chair. For this reason, we decompose our synthesis into approaching and sitting by only activating the ControlNet during the sitting. When the ControlNet is deactivated the control signal or when a “no contact” signal is present the control signal for the corresponding hand is zeroed.
Evaluation
Studying contact-conditioned interaction with chairs requires accurately labelled contacts and a diverse range of styled interactions. The COUCH dataset is captured to meet such needs. We evaluate our contact constrained motion synthesis method on the COUCH dataset qualitatively and quantitatively. Our method is the first approach that allows the user to explicitly define how the person should contact the chair and we generate natural and diverse motions satisfying these contacts. As such we evaluate our method on three axis, (i) accuracy in reaching the contacts, (ii) diversity and (iii) naturalness of the synthesised motion. For qualitative results, we highly encourage the readers to see our supplementary video. It can be seen that our method can generate diverse and natural motions while reaching the user-specified contacts. We quantitatively evaluate the accuracy of contacts and motion diversity on a total of 120 testing sequences on six subject-specific models trained on corresponding subsets of our COUCH dataset. Note that we evaluate raw synthesized motion without post-processing.
To our best knowledge, the most related work to ours are the Neural State Machine (NSM) and the SAMP since they both synthesize human-scene interactions. However, neither of the methods allows the use of fine-grained control over how the interaction should take place. We adapt these baselines for our task by additionally conditioning on the contact positions and refer to these new baselines as NSM+Control and SAMP+Control. Quantitative results are reported for both the original baselines and their adapted version. For each of the methods, we train subject-specific models with the corresponding subset of our COUCH dataset using the code provided by the authors. Our experiments, detailed below, show that naively providing contacts as input to existing motion synthesis approaches does not ensure that the generated motion satisfies the contacts. Our method, on the other hand, does not suffer from this limitation.
2 Evaluation on Control
In order to evaluate how well our synthesised motion meets the given contacts, we report the average contact error (ACE) as the mean squared error between the predicted hand contact and the corresponding given contact. We use the closest position of the predicted hand motion to the given contact as our hand contact. Since ACE might be susceptible to outliers and inspired by the literature on object detection , we also report average contact precision (AP@k), where we consider a contact as correctly predicted if it is closer than cm.
We compare our method with NSM+Control and SAMP+Control in Table 2. It can be observed that COUCH outperforms prior methods by a significant margin. Prior methods are trained to condition on the contact positions, however it is found (Figure 4) to be not sufficient as the contact input can be easily ignored during auto-regressive prediction. As a result the contact constraints are often not met. This highlights the importance of motion planning in form of trajectory predictors in order to reach the desired contacts. Our ControlNet provides valuable information on how to synthesise motion such that the given contacts are satisfied. Our motion prediction network PoseNet uses these control signals to generate contact constrained motions.
3 Evaluation on Motion Diversity
Diversity is an essential element for our motion synthesis, since a chair can be approached and interacted with in different ways. To quantify diversity, we evaluate using the Average Pairwise Distance (APD) on the synthesized pose features of the virtual human . defined as:
where is the total number of frames in all the testing sequences. Note that for evaluation, the virtual human is initialized at different starting points and is instructed to approach and sit on randomly selected chairs with randomly sampled contact points from the dataset, and motion is synthesized for 16 seconds for each sequence. We compare the diversity of synthesized motion in Table 3 and it can be seen that using explicit contacts allows our method to generate more varied motion.
4 Controlling with a series of Contacts.
A useful application of our approach is to automatically generate a motion sequence with a series of desired contacts in the context of animation, character control, when executing a set of complex actions. For instance, the person can be instructed to first sit with their hands on the armrest, then lift the arms to support the head before bringing the hands back to the armrest, see Figure 5) and the supplementary video. Our approach can be adapted for this task by iteratively providing the new goal locations for the hands as input after the present locations are reached.
5 Contact Prediction on Novel Shapes
Apart from user-specified contacts, we can additionally generate the contacts on the surface of a given chair using our proposed ContactNet. This allows us to generate fully automatic and diverse motions for sitting. To measure the diversity of the generated contacts from ContactNet, we compute the Average Pairwise Distance (APD) among the generated hand contact positions with unseen chair shapes. A total number of 200 unseen chairs are chosen, and each 10 contact positions are predicted for both hands.
is the number of objects and is the number of contacts generated per object. The APD on contact positions is 11.82 cm which is comparable to the ground truth dataset which has an APD of 14.07 cm. As shown qualitatively in Figure 6, we can generate diverse and plausible contact positions on chairs, which can generalize to unseen shapes.
Conclusion
We propose COUCH, the first method for synthesising controllable contact-driven human-chair interactions. Given initial conditions and the contacts on the chair, our model plans the motion of the hands, which drives the full body poses to satisfy contacts. In addition to the model, we contribute the COUCHdataset for human chair interactions which includes a wide variety of sitting motions approaching and contacting the chair in different ways. It consists of 3 hours of motion capture with 6 subjects interacting with registered 3D chair models, captured in high quality with IMUs and Kinects. Experiments demonstrate that our method consistently outperforms the SoTA by improving the average contact accuracy by 55 to better satisfy contact constraints. In addition to better control, it can been seen in the supplementary video that our approach generates more natural motion compared to the baseline methods. In the future, we want to extend our dataset to new activities and train a multi-activity contact driven model. In the supplementary, we discuss further future directions in this new problem of fine-grained controlled motion synthesis. Our dataset and code will be released to foster further work in this new research direction.
Acknowledgement
We would like to thank Xianghui Xie for helping with data processing, and we are very grateful for all the participants who took part in the data capture.
References
APPENDIX
In this appendix, we provide additional information about the dataset, implementation details, post-processing techniques. We also discuss on the current limitation as future research perspectives.
Dataset
Table 4 shows a break down of our dataset in terms of different types of interactions. Our dataset consists of 3 hours of MoCap with over 500 motion sequences.
2 Data Processing
SMPL Fitting. We segment the human in captured RGB images by running Detectron V2 followed by manual correction with on the segmentation masks. These masks are then used to segment multi-view depth maps and lift human point clouds from 2D to 3D. We use FrankMocap to initialize the SMPL pose from the images and then apply instance specific optimization to fit the SMPL model to the segmented human point cloud. For more accurate fitting, we additionally obtain the SMPL shape parameters of each subject from 3D scans using . Synchronization with the IMUs. The fitted SMPL model provides us with accurate contacts with the scene, however, the fitted motion sequence is prone to occlusion and drastic body movements, as a result, the fitted motion can be jittery at times. On the other hand, the pose captured with the IMUs is smooth over time, but it might not accurately capture the contacts. To this reason, we synchronize the Kinect captured data with the body sensors by incorporating the SMPL fitted poses into the IMU pose sequences. After synchronization, we optimize the joint rotations to achieve temporal smoothness via the objective
where represents the acceleration of the body joints in frame approximated by central difference.
where represents the foot joint positions at frame . The resulting motion sequence is temporally smooth and has accurate contacts registered with the chair models. Object Processing. To obtain object segmentation, we pre-scan objects using a 3D scanner . We then use multi-view object keypoints, marked by manual annotators on the images, to fit the pre-scanned chair meshes to the given frame. The segmentation masks are then obtained by projecting fitted object meshes to the images. Since the chairs remain static during the capture, we average over the 6D pose of the fitted chair model during each capture session to obtain the final transformation of the chair.
Training Details
As shown in Figure 7, the contact network is a two-layer LSTM architecture. Each layer has a hidden dimension of 512. The pose and the control signals (hand trajectories, and the local phases) are each encoded through a two-layer fully connected network with of shape {128, 128} before passing through the LSTM. We apply scheduled sampling on hand trajectories for better model performance. For the local phases, we always use the ground truth. Each of our training samples is in a sequence of 60 frames. The ControlNet is trained for 150 epochs with an Adam optimizer. The initial learning rate is 1e-3 and a cosine learning rate scheduler was used to decay the learning rate gradually to 5e-6. The full training of a subject-specific model takes approximately 1 hour on an NVIDIA V100 GPU.
2 PoseNet
The PoseNet adopts the mixture-of-expert structure . It consists of different feature encoders of structures shown in Table 5. The gating network and the prediction networks are both three-layer fully-connected networks, with hidden dimensions of 128 and 512 respectively. The number of experts is set to 10. The PoseNet is trained for 150 epochs with an Adam optimizer. The initial learning rate is 1e-4 and a cosine learning rate scheduler was used to decay the learning rate gradually to 5e-6. The full training of a subject-specific model takes approximately 6 hours on an NVIDIA V100 GPU.
3 ContactNet
The ContactNet encodes the scene through a three-layer fully connected network of shape {512, 512, 64}. The latent vector of the VAE is of size 6. The weight of the Kullback-Leibler divergence is 0.1. We use the Adam optimizer with a learning rate of 1e-3 and train ContactNet for 150 epochs. The full training of a subject-specific model takes approximately 10 minutes on an NVIDIA V100 GPU.
Contact Projection and Trajectory Fitting
To ensure the ContactNet predicts contacts that land exactly on the surface of the object, we perform a post-processing step, when the distance of the network predicted contact to the surface is less than a set threshold of 10 cm, we simply project the contact onto the nearest point on the chair surface. When the distance is greater than 10 cm, we simply neglect the predicted contact. The ControlNet predicts the future hand trajectories, and it would be possible to fit the predicted pose to the predicted hand position from the hand trajectories at each frame to further improve the satisfaction of the contact constraints. Note, in the evaluation of the main paper we do not apply such fitting technique.
Limitations and Future Direction
We observe the synthesized motion can slightly intersect with the chair. A solution to this problem would be to apply a post-processing step to avoid such collision. In order to generalize to more different chair shapes, it would be useful to investigate better ways of encoding the scene geometry while trying to avoid over-fitting.
Different shaped person can intersect with the same object very differently even when performing the same motion. The COUCH dataset captures human interaction with different body shapes. With the dataset, it is possible to study how to build subject-variant motion synthesis model and how to effectively condition on the body shapes. These are challenges in motion synthesis that have not been tackled.
Our work on controllable human-chair interaction. It would be useful to extend the scope of interacted objects, especially considering the cases when the objects are non-static, when performing motions such as lifting a box, or opening a door. Another possible direction would be to further apply contact-based control in these interactions.