Learning by Watching: Physical Imitation of Manipulation Skills from Human Videos
Haoyu Xiong, Quanzhou Li, Yun-Chun Chen, Homanga Bharadhwaj, Samarth Sinha, Animesh Garg
I Introduction
Robotic Imitation Learning, also known as Learning from Demonstration (LfD), allows robots to acquire manipulation skills performed by expert demonstrations through learning algorithms . While progress has been made by existing methods, collecting expert demonstrations remains expensive and challenging as it assumes access to both observations and actions via kinesthetic teaching , teleoperation , or crowdsourcing platform . In contrast, humans have the ability to imitate manipulation skills by watching third-person performances. Motivated by this, recent methods resort to endowing robots with the ability to learn manipulation skills via physical imitation from human videos .
Unlike conventional LfD methods , which assume access to both expert obsevations and actions, approaches based on imitation from human videos relax the dependencies, requiring only human videos as supervision . One of the main challenges of these imitation learning methods is how to minimize the domain gap between humans and robots. For instance, human arms may have different morphologies than those of robot arms. To overcome the morphology mismatch issue, existing imitation learning methods typically leverage image-to-image translation models (e.g., CycleGAN ) to translate videos from the human domain to the robot domain. However, simply adopting vanilla image-to-image translation models still does not solve the imitation from human videos task, since the image-to-image translation models often capture only the macro features at the expense of neglecting the details in salient regions that are crucial for downstream tasks .
In this paper, we present Learning by Watching (LbW), a framework for physical imitation from human videos for learning robot manipulation skills. As shown in Figure 1, our framework is composed of a perception module and a policy learning module for physical imitation. The perception module aims at minimizing the domain gap between the human domain and the robot domain as well as capturing the details of salient regions that are crucial for downstream tasks. To achieve this, our perception module learns to translate the input human video to the robot domain with an unsupervised image-to-image translation model, followed by performing unsupervised keypoint detection on the translated robot video. The detected keypoints then serve as a structured representation that contains semantically meaningful information and can be used as input to the downstream policy learning module.
To learn manipulation skills, we cast this as a reinforcement learning (RL) problem, where we aim to enable the robot to perform physically viable learning with the objective to imitate similar behavior as demonstrated in the translated robot video under context-specific constraints. We evaluate the effectiveness of our LbW framework on five robot manipulation tasks, including reaching, pushing, sliding, coffee making, and drawer closing in two simulation environments (i.e., the Fetch-Robot manipulation in OpenAI gym and meta-world ). Extensive experimental results show that our algorithm compares favorably against the state-of-the-art approaches.
The main contributions are summarized as follows:
We present a framework for physical imitation from human videos for learning robot manipulation skills.
Our method learns structured representations based on unsupervised keypoint detection that can be used directly for computing task reward and policy learning.
Extensive experimental results show that our LbW framework achieves the state of the art on five robot manipulation tasks.
II Related Work
Imitation from human videos. Existing imitation learning approaches collect demonstrations by kinesthetic teaching , teleoperation , or through crowdsourcing platform , and assume access to both expert observations and expert actions at every time step. Recent progress in deep representation learning has accelerated the development of imitation from videos . While applying image-to-image translation models to achieve imitation from human videos has been explored , the dependency on paired human-robot training data makes these methods hard to scale.
Among them, AVID is closely related to our work which translates human demonstrations to robot domain via CycleGAN in an unpaired data setting. However, directly encoding the translated images using a feature extractor for deriving state representations may suffer from visual artifacts generated by image-to-image translation models, leading to suboptimal performance on downstream tasks.
Different from methods based on image-to-image translation models, Maximilian et al. leverage 3D detection to minimize the visual gap between the human domain and the robot domain. SFV enables humanoid characters to learn skills from videos based on deep pose estimation. Our method shares a similar reward computing scheme as these approaches . The difference is that these methods require additional label data, whereas our framework is learned in an unsupervised fashion.
Cycle consistency. The idea of exploiting cycle consistency constraints has been widely applied in the context of image-to-image translation. CycleGAN learns to translate images in an unpaired data setting by exploiting the idea of cycle consistency. UNIT achieves image-to-image translation by assuming a shared latent space between the two domains. Other methods explore translating images across multiple domains or learning to generate diverse outputs . Recently, the idea of cycle consistency is also applied to address various problems such as domain adaptation and policy learning . In our work, our LbW framework employs a MUNIT model to perform human to robot translation for achieving physical imitation from human videos. We note that other unpaired image-to-image translation models are also applicable in our task. We leave the discussion on the effect of different image-to-image translation models as future work.
Unsupervised keypoint detection. Detecting keypoints from images without supervision has been studied in the literature . In the context of computer vision, existing methods typically infer keypoints by assuming access to the temporal transformation between video frames or employing a differentiable keypoint bottleneck network without access to frame transition information . Other approaches estimate keypoints based on the access to known image transformations and dense correspondences between features .
Apart from the aforementioned approaches, some recent methods focus on learning keypoint detection for image-based control tasks . In our method, we adopt Transporter to detect keypoints from the translated robot video in an unsupervised manner. We note that while other unsupervised keypoint detection methods can also be used in our framework, the focus of our paper lies in learning structured representations that are semantically meaningful and can be used directly for downstream policy learning. We leave the development of unsupervised keypoint detection methods as future work.
III Preliminaries
To achieve physical imitation from human videos, we decompose the problem into a series of tasks: 1) human to robot translation, 2) unsupervised keypoint-based representation learning, and 3) physical imitation with RL. Here, we review the first two tasks, in which our method builds upon existing algorithms.
III-B Unsupervised Keypoint Detection
To perform control tasks, existing approaches typically resort to learning state representations based on image observations . However, the image observations generated by image-to-image translation models often capture only macro features while neglecting the details in salient regions that are crucial for downstream tasks. Deriving state representations by encoding the translated image observations using a feature encoder would lead to suboptimal performance. On the other hand, existing methods may also suffer from visual artifacts generated by the image-to-image translation models. In contrast to these approaches, we leverage Transporter to detect the keypoints in each translated video frame in an unsupervised fashion. The detected keypoints form a structured representation that captures the robot arm pose and the location of the interacting object, providing semantically meaningful information for downstream control tasks while avoiding the negative impact of visual artifacts caused by the imperfect image-to-image translation.
To realize the learning of unsupervised keypoint detection, Transporter leverages object motion between a pair of video frames to transform a video frame into the other by transporting features at the detected keypoint locations. Given two video frames and , Transporter first extracts feature maps and for both video frames using a feature encoder and detects -dimensional keypoint locations and for both video frames using a keypoint detector . Transporter then synthesizes the feature map by suppressing the feature map of around each keypoint location in and and incorporating the feature map of around each keypoint location in :
where is a Gaussian heat map with peaks centered at each keypoint location in .
In the next section, we leverage the Transporter model to detect keypoints for each translated video frame. The detected keypoints are then used as a structured representation for defining the reward function and as the input of the policy network to predict an action that is used to interact with the environment.
IV Proposed Method
In this section, we first provide an overview of our approach. We then describe the unsupervised domain transfer with keypoint-based representations module. Finally, we describe the details of physical imitation with RL.
IV-B Unsupervised Domain Transfer with Keypoints
To achieve physical imitation from human videos, we develop a perception module that consists of a MUNIT model for human to robot translation and a Transporter network for keypoint detection as shown in Figure 3. To train the MUNIT model, we first collect the training data for the source domain (i.e., human domain) and the target domain (i.e., robot domain). The source domain contains the human demonstration video that we want the robot to learn from. To increase the diversity of the training data in the source domain for facilitating the MUNIT model training, we follow AVID and collect a few random data by asking the human to randomly move the hands above the table without performing the task. As for the target domain training data, we collect a number of robot videos generated by having the robot perform a number of actions that are randomly sampled from the action space. As such, the collection of the robot videos does not require human expertise and effort.
As mentioned in Section III-B, we aim to learn keypoint-based representations from the translated robot video in an unsupervised fashion. To achieve this, we leverage Transporter to detect the keypoints in each translated robot video frame in an unsupervised fashion, as there are no ground-truth keypoint annotations available.
IV-C Physical Imitation with RL
To control the robot, we use RL to learn a policy from image-based observations that maximize the cumulative values of a learned reward function. In our method, we decouple the policy learning phase from the keypoint-based representation learning phase. Given the keypoints trajectory of the translated robot demonstration video and the keypoint-based representation of the current observation , our policy network outputs an action which is executed in the environment to obtain the next observation . To achieve physical imitation, we aim to minimize the distance between the keypoints trajectory of the agent and that of the translated robot demonstration video. Specifically, we define the reward as
where and are hyperparameters that balance the importance between the two terms, and the aforementioned goal is imposed on and , which are defined by the following equations:
where , , aims to minimize the distance between the keypoint-based representation of the current observation and the most similar (closest) keypoint-based representation in the keypoints trajectory of the translated robot demonstration video , and is the first-order difference equation of .
We add the tuple to a replay buffer. Then, the policy network can be trained with any RL algorithms in principle. We make use of Soft-Actor Critic (SAC) as the RL algorithm for policy learning in our experiments.
V Experiments
In this section, we describe the experimental settings and report results with comparisons to state-of-the-art methods on five robot manipulation tasks. Through experiments, we aim to investigate the following questions:
How accurate is our perception module in handling the human-robot domain gap and in detecting keypoints?
How does LbW compare with state-of-the-art baselines in terms of performance on robot manipulation tasks?
We perform experimental evaluations in two simulation environments, i.e., the Fetch-Robot manipulation in OpenAI gym and meta-world . We evaluate on five tasks: reaching, pushing, sliding, coffee making, and drawer closing. Figure 4 presents the overview of each task, including the task scenes and one sample human video frame for each task. The goal of each task is described as follows.
For the reaching task, the robot has to move its end-effector to reach the target.
For the pushing task, a puck is placed on the table in front of the robot, and the goal is to move the puck to the target location.
For the sliding task, a puck is placed on a long slippery table and the target location is beyond the reach of the robot. The goal is to apply an appropriate force to the puck so that the puck slides and stops at the target location due to friction.
For the coffee making task, a cup is placed on the table in front of the robot, and the goal is to move the cup to the location right below the coffee machine outlet. The moving distance of coffee making task is longer than the one in the pushing task.
For the drawer closing task, the robot has to move its end-effector to close the drawer.
In the policy learning phase, the robot receives only an RGB image of size as the observation. The robot arm is controlled by an Operational Space Controller in end-effector positions. As each of the tasks is described by a single human video, we set the initial locations of the object and the target to a fixed configuration.
V-B Comparison to Baseline Methods
To evaluate the effectiveness of our perception module, we implement two baseline methods using the same control model as LbW, which is adopted from SACAE , but with different reward learning methods.
Classifier-reward. We implement a classifier-based reward learning method in a similar way as VICE . For each task, given robot demonstration videos, instead of the human videos, the CNN classifier is pre-trained on ground-truth goal images with positive labels and the remaining images with negative labels. To learn a policy in the environment, we adopt the implementation from SACAE, where we use the classifier-based reward to train the agent.
AVID-m. Since AVID is the state-of-the-art method that outperforms prior approaches, including BCO and TCN , we focus on comparing our method with AVID. For a fair comparison, we reproduce the reward learning method of AVID and replace the control module with SACAE. We denote this method as AVID-m. For each task, given human demonstration videos, we first translate the human demonstration videos to the robot domain using the CycleGAN model. Then the CNN classifier is pre-trained on the translated goal images with positive labels and the remaining translated images with negative labels. For RL training, we adopt the implementation from SACAE .
V-C Dataset Collection and Statistics
We decouple the training phase of the perception module from that of the policy learning module.
Dataset for perception module training. To train our perception module and the CycleGAN method, we collect human expert videos and videos of a human performing random actions without performing the tasks for the human domain. For the robot domain, we first constrain the action space of the robot such that unexpected robot poses will not occur (i.e., robot arms are constrained to move above the table), and then run a random policy to collect robot videos. Note that we do not use robot expert videos for training the perception module. Table I presents the dataset statistics for each task for training the perception module.
Dataset for policy learning. For policy learning, we use only one single human expert video to train our policy network. The AVID-m method uses human expert videos, while the classifier-reward approach uses robot expert videos.
V-D Performance Evaluations
Following AVID , we use success rate as the evaluation metric. At test time, the task is considered to be a success if the robot is able to complete the task within a specified number of time steps (i.e., time steps for reaching and pushing, and time steps for sliding, coffee making, and drawer closing). The results are evaluated by test episodes for each task. Table II reports the success rates of our method and the two baseline approaches on all five tasks. We find that for the reaching task, all three methods achieve a success rate of . For the sliding, drawer closing, and coffee making tasks, our LbW performs favorably against the two competing approaches.
The difference between the AVID-m method and the classifier-reward approach is that AVID-m leverages CycleGAN for human to robot translation, while the classifier-reward method using ground-truth robot images directly. As shown in Figure 5, the translated images of AVID-m have clear visual artifacts. For instance, the red cube disappears and the robot poses in the translated images do not match those in the human video frames. The comparisons between AVID-m and the classifier-reward method and the visual results of AVID-m in Figure 5 show that using image-to-image translation models alone for minimizing the human-robot domain gap will have negative impact to the performance of the downstream tasks. Our perception module learns unsupervised human to robot translation as well as unsupervised keypoint detection on the translated robot videos. The learned keypoint-based representation provides semantically meaningful information for the robot, allowing our LbW framework compares favorably two competing approaches. More results, videos, performance comparisons, and implementation details are available at pair.toronto.edu/lbw-kp/.
V-E Discussion of Limitations
While results on five tasks demonstrate the effectiveness of our LbW framework, there are two limitations. First, existing imitation learning methods that are based on image-to-image translation require the pose of the human arms and that of the robot arms to be similar. As a result, these methods may not perform well on human demonstration videos that have larger pose variations or with more natural poses. Our LbW framework also leverages an image-to-image translation model, thus suffering from the same limitation as these methods. Second, in our method, learning from only a single human video limits the model generalization to new scenes.
VI Conclusions
We introduced LbW, a framework for physical imitation from human videos. Our core technical novelty lies in the design of the perception module that minimizes the domain gap between the human domain and the robot domain followed by keypoint detection on the translated robot video frames in an unsupervised manner. The resulting keypoint-based representations capture semantically meaningful information that guide the robot to learn manipulation skills through physical imitation. We defined a reward function with a distance metric that encourages the trajectory of the agent to be as close to that of the translated robot demonstration video as possible. Extensive experimental results on five robot manipulation tasks demonstrate the effectiveness of our approach and the advantage of learning keypoint-based representations over conventional state representation learning approaches.
Acknowledgement. Animesh Garg is supported by CIFAR AI chair, and we would like to acknowledge Vector institute for computation support.