Keypoints into the Future: Self-Supervised Correspondence in Model-Based Reinforcement Learning
Lucas Manuelli, Yunzhu Li, Pete Florence, Russ Tedrake
Introduction
It has been argued that one of the hallmarks of human-level learning is the ability to construct and leverage causal models of the world . In the area of manipulation, this manifests itself in the ability to approximately predict how an object will move if we grasp or push it. Traditional model-based robotics has successfully leveraged such predictive models, oftentimes derived from first principles, to solve challenging planning and control problems . In the area of practical vision-based robotic manipulation, however, it has been particularly hard to leverage such predictive models, due to varied and novel objects and the high-dimensional observation spaces involved (e.g., RGB or RGBD images). Alternative approaches, such as imitation learning or model-free reinforcement learning, sidestep the task of building a predictive model and directly learn a policy. Although this can be appealing, model-based techniques offer several benefits. They can be sample efficient compared to model-free methods and, in contrast to behavior cloning techniques, can leverage off-policy non-expert data. Once a model has been acquired, it can be used together with a planner to achieve a wide variety of tasks and goals. One of the main challenges for model learning applied to robotic manipulation is determining the state representation on which the dynamics model should be learned. Prior work has used approaches ranging from full image space dynamics to a variety of autoencoder formulations .
In this paper we propose to use object keypoints, which are tracked over time, as the latent state in which to learn the dynamics. These keypoints anchor our model-based predictions, and provide various advantages over alternatives such as abstract latent states: (i) the output is interpretable, which enables the ability to analyze the performance of the visual model separately from the predictive model. (ii) The representation is 3D and hence can naturally handle changing general off-axis 3D camera positions. (iii) As demonstrated in the visual models we use, Dense Object Nets, have demonstrated reliable performance in a variety of real-world settings and are able to generalize at the category level. We found that autoencoder approaches particularly struggled with category-level generalization. And as discussed in prior works , keypoints and dense correspondences provide advantages over using 6D object poses: they can apply to deformable objects and represent category-level generalization.
We show that this formulation enables reliable, sample-efficient learning capable of precise visual-feedback-based manipulation in the real world – and is trained with nothing other than a small amount of interaction data (10 minutes) and a single demonstration for goal specification. In our approach to acquiring keypoints, we extract them as descriptors which are tracked from a dense descriptor model – while multiple approaches could be used to acquire keypoints, this route can be entirely self-supervised. As opposed to which uses keypoints from a dense-correspondence model in an imitation learning framework, the use of keypoints as input to a model-based RL system presents several unique challenges. In particular the keypoints need to be both informative for the task at hand and be able to be tracked accurately. In this paper we explore these challenges and propose solutions.
Contributions: Our primary contribution is (i) a novel formulation of predictive model-learning using learned dense visual descriptors as the state representation. We use this approach to perform closed-loop visual feedback control via model-predictive-control (MPC). (ii) Using simulated manipulation experiments we demonstrate that this approach offers performance benefits over a variety of baselines, and (iii) we validate our approach in real-world robot experiments.
Related Work
We focus our related work on methods that target robotic manipulation with learned predictive models. The Introduction discussed alternative approaches to synthesizing closed-loop feedback controllers without predictive models, via imitation learning or model-free reinforcement learning.
Model-based RL in Robotic Manipulation. These methods can be classified by whether they use first-principles-based or data-driven models, and whether they consume raw visual inputs (such as RGBD images) or they consume ground truth state information (from a simulator or an external perception system). First-principles-based models, e.g., , rely on known object models and thus don’t generalize to novel or unknown objects, and also rely on external vision systems. Given this, the tasks we consider are out of scope for these approaches. In the area of methods that use external vision systems but learn data-driven models, learns a dynamics model for a closed-loop-controlled planar pushing task, however, the approach is tailored to the specifics of the pusher-slider task and doesn’t readily generalize to other tasks, or to novel objects within the same task. learns deep dynamics models for a variety of different simulated dexterous manipulation tasks, using ground truth object states, and one task on real hardware, using an external camera-based 3D object position tracker. Although achieve impressive results, their reliance on ground truth object state and/or specialized visual trackers limits their general applicability in more diversified real-world manipulation tasks.
There also exists a large literature on model-based RL for robotic manipulation that operates in the more challenging problem class of directly consuming image observations. Methods can be broadly categorized into whether they predict the image-space dynamics of the entire image , or predict the dynamics of a low-dimensional latent state . Although image-space dynamics approaches are general, they require large training datasets. For latent-space dynamics approaches, to avoid a trivial solution where all observations get mapped to a constant vector, a regularization strategy is needed for the latent state . use an autoencoder architecture and regularize the latent state using a reconstruction loss. regularize the latent space by simultaneously training both forward and inverse dynamics models while uses a contrastive loss on the latent state. In contrast to these approaches we use visual-correspondence pretraining to produce a latent state which is physically grounded and interpretable as the 3D locations of keypoints on the object(s).
Visual Representation Learning For more related-work in self-supervised visual learning for robotics, we refer the reader to . Approaches that use autoencoders or full image-space dynamics rely on image reconstruction as their source of visual supervision. is perhaps most related to ours, in that they first learn a visual model which is then used for a downstream task, and they show that freezing the visual model and using a keypoint-type representation as an input to model-free RL algorithms improves performance on Atari ALE . Our approach is distinct in that (i) we use a predictive model-based method rather than a model-free method, (ii) we use a correspondence-based training loss, while uses an image-reconstruction-based loss with a specialized pixel-space transport mechanism, and (iii) we demonstrate results with real-world hardware. As a baseline, we try using their vision model as an input to the same model-based RL algorithm used by our own method.
Formulation: Self-Supervised Correspondence in Model-Based RL
This section describes our approach to model-based RL using visual observations. The goal is to learn a dynamics model that can then be used to perform online planning for closed-loop control.
Our setting consists of an environment with states , observations , actions and transition dynamics . The task is specified by a reward function and the goal is to choose actions to maximize the expected reward over a trajectory. We approach this problem by first learning an approximate dynamics model , which is trained to minize the dynamics prediction error on the observed data . The learned model is then used to perform online planning to obtain a feedback controller.
Rather than learn the dynamics directly in the observation space, as in , we instead learn a mapping from the high-dimensional observation space to a low-dimensional latent space together with a dynamics model in this latent space, where encodes a short history (we use in all experiments). This latent state should capture sufficient information about the true state such that driving sufficiently well achieves the goal of driving to (where is the latent state corresponding to ). Given the current latent state and a sequence of actions we can predict future latent states by repeatedly applying our learned dynamics model. The forward model is then trained to minimize the dynamics prediction error (also know as simulation error) over a horizon
2 Learning a Visual Representation
The objective of the visual model is to produce a low-dimensional feature vector which serves as a suitable latent state in which to learn the dynamics. For the types of tasks and environments that we are interested in, spatial information about object locations is a critical piece of information. Pose estimation has played a critical role in classical manipulation pipelines, and was also used in dynamics learning approaches such as . In general producing pose information from high-dimensional observations (such as RGBD images) requires a dedicated perception system. Although pose can be a powerful state representation when dealing with a single known object, as noted in it has several drawbacks that limit its usefulness in more general manipulation scenarios. In particular it doesn’t readily (i) extend to the case of deformable objects, (ii) generalize to novel objects or (iii) extend to category-level tasks.
Our strategy is to leverage visual-correspondence pre-training to track points on the object of interest. The locations of these tracked points can then serve as the latent state on which we learn the dynamics. Similar to the approach taken in , we use visual-correspondence learning, which is trained in a completely self-supervised fashion, to train a visual model which that can be used to find correspondences across RGB images. We then propose several approaches to produce a low-dimensional latent using the pre-trained dense-correspondence model.
Descriptor Set (DS): In our simplest variant the latent state is simply made up of keypoint locations for a set of descriptor keypoints randomly sampled from the object. Specifically, we sample (we use in all experiments) descriptors corresponding to pixels from a masked reference descriptor image in our training set. Thus
Spatial Descriptor Set (SDS): Rather than randomly sampling descriptors, as in (DS), this method attempts to choose descriptors having specific properties. In particular we would like the descriptors to be (i) reliable, and (ii) spatially separated. By reliable we mean that they can be localized with high accuracy and don’t become occluded during the typical operating conditions, see Figure 2 for an example. Spatially separated means that the chosen descriptors aren’t all clustered around the same physical location on the object(s) of interest, but rather are sufficiently spread out (either in 3D space or pixel space) to provide meaningful information about both object position and orientation. Our dense descriptor model can provide a confidence score associated with descriptors and their associated correspondences.see Appendix 6.3 for more details Figure 2 shows a clear example of high confidence for a valid match and low-confidence when no valid correspondence exists due to an occlusion. We use this feature of our visual model to compute a confidence score for each each descriptor according to what fraction of images in the training set contain a high probability correspondence for . The intuition is that descriptors corresponding to points on the object that are easy to localize and remain unoccluded will have a high confidence score. As in the DS method we initially select a large number of descriptors () corresponding to points on the object. We then select the descriptors with the highest average confidence on the training set and which additionally satisfy a threshold on minimum separation distance (typically pixels in a image). In the experiments we use or .
The learnable weights are trained jointly with the parameters of the dynamics model, and are fixed at test time. The fact that is gotten from by taking a weighted linear combination preserves the interpretation of as tracking keypoints on the object, while allowing some flexibility to ignore unreliable keypoints.
Weighted Spatial Descriptor Set (WSDS): This method is simply the combination of (SDS) and (WDS).
3 Learning the Dynamics
We adopt a standard dynamics learning framework where we aim to predict the evolution of the latent state . We train our dynamics model to minimize the prediction error in Equation (1) where . Our proposed methods differ in the structure of the mapping and the set of trainable parameters .see Appendix 7.2 for details For all of our methods we keep the weights of the visual-correspondence network, , fixed.
4 Online Planning for Closed-Loop Control
Once we have learned a dynamics model , we use online planning with MPC to select an action. Given a goal latent state , we want to find an action sequence that maximizes the reward . Our dynamics learning approach is agnostic to the type of optimizer used to solve the MPC problem and the focus of our work is on the visual and dynamics learning, rather than the specifics of the MPC. Many model-based RL approaches use a random-sampling based planner (e.g. cross-entropy method) to solve the underlying MPC problem. We experimented with random shooting, gradient based shooting, cross-entropy and model-predictive path integral (MPPI) planners and found that MPPI worked best in our scenarios.See Appendix 8 for more details.
Results
We perform experiments aimed at answering the following questions: (1) Is it possible to successfully use self-supervised descriptors as the latent state for a model-based RL system? (2) What is the effect of various design decisions in our algorithm? (3) How does visual-correspondence learning compare to several benchmark methods in terms of enabling effective model-based RL policies? (4) Can we apply the method on real hardware?
Our main contribution is on the formulation of the visual model and dynamics learning problem rather than the specifics of the MPC. However, to accurately compare our approach to various baselines, we need to perform experiments in which the dynamics model is used in closed-loop to solve a manipulation task. Our formulation of dynamics learning is very general, and in principle can handle a wide variety of manipulation scenarios. However, even with an accurate model (whether it comes from first principles or is learned), using this model to perform closed-loop feedback control remains a challenging problem. Thus, for our closed-loop experiments, we limit ourselves to pushing tasks that can adequately be solved by the planners outlined in Section 3.4.
Tasks: Extended experimental details are provided in the Appendix but we provide a brief overview of the tasks here. We perform experiments with four simulated tasks (depicted in Figure 3) and one hardware task. All tasks involve pushing an object to a desired goal state. The first three simulation tasks, denoted as top-down, angled, occlusions involve pushing a single object. top-down has cameras in a top-down orientation while they are angled at 45 degrees in angled. Task occlusions keeps the angled camera positions but the box is resting on a different face, resulting in significant self-occlusions. Task mugs involves pushing many different objects from a category, in this case mugs with different size, shape and textures. The hardware task is essentially identical to the angled sim task.
Figures 1, 2 show the performance of our dense visual-correspondence model. In particular Figure 1 shows the localization performance on real data along a trajectory while Figure 2 shows an example of the confidence scores used as in the SDS method. Figure 3 shows the descriptors used in the SDS method for each of our simulation tasks. In particular Figure 3 (d)-(e) shows the ability of the descriptors to accurately find correspondences across different object instances within a category, despite differences in color and shape.
2 Ablations on visual-correspondence for dynamics learning
Ablation studies show that the choice of what to track can have a substantial impact on model performance, especially in the case of partial occlusions. Quantitative results are detailed in Table 1. On tasks top-down and angled camera all methods perform reasonably well, almost matching the performance of the GT 3D baseline that uses ground truth state information. Intuitively this is because there are minimal occlusions in these settings, and so almost all keypoints can be tracked reliably using dense-visual-correspondence. In contrast the occlusions introduces the potential for significant occlusions. Given the camera angle as shown in Figure 3 (c) and the fact that the object rotates through the full degrees of yaw during the task, only the top face of the box remains unoccluded while the 4 side faces are alternately occluded and visible. This task exposes significant differences in performance among our various ablations. In particular SDS, WDS and WSDS perform significantly better than DS. We believe that this is due to the fact that some of the descriptors that are tracked in DS correspond to locations on the object that become occluded during an episode. The dense visual-correspondence model is not able to track keypoints through occlusions, and when trying to localize an occluded point our dense-correspondence model maps it to the closest point in descriptor space, which is not the location of the true correpondence. Hence the keypoint locations that makeup the latent state for DS suffer reduced accuracy, leading to a less accurate dynamics model and ultimately lower performance when used for closed-loop MPC.
On the category-level task mugs the camera is in a top-down position and thus occlusions are no longer an issue. However because the task involves different objects from a category there is shape variation among the different objects. Thus the methods that use keypoints, namely DS and WDS, perform better than the sparser variants SDS, WSDS that use only keypoints. We believe that this is due to the fact that having a larger number of keypoints better captures the shape variation across object instances and allows for a more accurate dynamics model.
3 Comparison of visual-correspondence pretraining with baselines
In our comparisons against baselines, our model outperforms alternatives on all experimental tasks. The largest differences are apparent in tasks occlusions and mugs which involve partial occlusions and category-level generalization, respectively. The transporter baseline uses the keypoint locations from as the latent state , while the autoencoder baseline jointly trains an autoencoder with a forward dynamics model. For a detailed discussion of the baselines see Appendix 9.3. Quantitative results are detailed in Table 2.
On task top-down camera, the transporter model was able to achieve performance that was only slightly worse than WDS and SDS, while the autoencoder performed significantly worse. The top-down setting is ideally suited for the feature transport approach of the transporter model.
On tasks angled camera and occlusions, which have angled camera positions as opposed to the top-down task, the performance of our methods remained consistent while transporter suffered. This is potentially due to the fact that the feature transport mechanism of transporter is not well-suited to off-axis camera positions in 3D worlds. The performance of the autoencoder baseline remainded consistent, but worse than our approach, across tasks without category-level generalization.
Task mugs contains a variety of different mug shapes with varied visual appearances and hence tests category-level generalization of both the perception and dynamics models. As discussed in Section 4.1, our dense-correspondence model is able to find correspondences across these variations in appearance using only self-supervision, which allows us to learn a dynamics model that is effective for completing the task. This task is significantly harder than the other tasks not only because of the presence of novel objects, but because the goal states involve much larger rotations, as detailed in the first row of Table 1. Both the transporter and autoencoder baselines perform poorly in this task. We hypothesize that this is due to the fact that there is much more variance in the visual appearance of the objects as compared to the other tasks and thus the latent state produced by these baselines is not amenable to dynamics learning.
4 Hardware
Experimental Setup: We used a Kuka IIWA LBR robot with a custom cylindrical pusher attached to the end-effector to perform our hardware experiments, see Figure 1. RGBD sensing was provided by two RealSense D415 cameras rigidly mounted offboard the robot and calibrated to the robot’s coordinate frame. To enable effective correspondence learning between views, it is ideal to have views with some overlap such that correspondences exist, but still maintain different-enough views. At test time only a single camera is used to localize the dense-descriptor keypoints. The robot is controlled by commanding end-effector velocity in the plane at 5Hz.
Hardware Results: For visual learning we collected a small dataset of the object in different positions to provide a diverse set of views for training the dense-correspondence model. For dynamics learning, we collected a dataset of the robot randomly pushing the object around. This amounted to approximately 10 minutes of interaction time and was used to train the dynamics model. All hardware experiments used the SDS method. To enable our MPC controller to accomplish long-horizon tasks, we supplied the controller with a reference trajectory for the object keypoints that came from a single demonstration, see Figure 1 (a),(d). The MPC controller then tracked this reference trajectory using a second MPC horizon, which corresponds to since we are commanding actions at Hz. We collected 4 different reference trajectoriessee Appendix 10 for details on the reference trajectories and ran multiple rollouts for each trajectory, varying the initial condition of the object pose during each run to test the region of attraction of our controller. In all cases our controller showed the ability to stabilize the system to the reference trajectory in spite of perturbations to the initial condition. Quantitative results are detailed in Figure 4. In particular, we see that the ability of the controller to stabilize the trajectory in the face of disturbances in the initial condition depends on the trajectory. For trajectory (1) the controller is able to stabilize disturbances of up to 60 degrees, while trajectory (4) has a much smaller region of attraction. As can be seen in Appendix 10, trajectory (1) is a relatively simple trajectory with minimal orientation change between start and goal, while trajectory (4) involves a challenging 180-degree orientation change and requires the robot to operate at the edge of its kinematic workspace, reducing its control authority. Overall, our system exhibits impressive feedback and is able to track a trajectory in the keypoint latent space, enabling one-shot imitation learning. We encourage the reader to watch the videos on our project page to see the system in action.
Conclusion
We presented a method for using self-supervised visual-correspondence learning as input to a predictive dynamics model. Our approach produces interpretable latent states that outperform competing baselines on a variety of simulated manipulation tasks. Additionally, we demonstrated how the category-level generalization of our visual-correspondence model enables learning of a category-level dynamics model, resulting in large performance gains over baselines. Finally, we demonstrated our approach on a real hardware system, and showed its ability to stabilize complex long-horizon plans by tracking the latent state trajectory from a single demonstration.
This work was supported by Amazon.com Services LLC (Award No. CC MISC 00272683 2020 TR) and an Amazon Research Award. The views expressed are not endorsed by our funding sponsors.
References
Appendix
Dense Correspondence
Here we give a brief overview of the dense correspondence model formulation with spatial distribution losses . We briefly explain the loss functions and how the descriptors are localized.
We use an architecture that produces a full resolution descriptor image. Namely it maps
We use a FCN (fully-convolutional network architecture) with a ResNet-50 or ResNet-101 with the number of classes set to the descriptor dimension. Note that the FCN used in this work does not use striding and upsampling as in the architecture originally used in .
2 Loss Function
For all shown experiments we use the spatially-distributed loss formulation with a combination of heatmap and 3D spatial expectation losses, as described in , Chapter 4.
Let be the pixel space location of a ground truth match. Then we can define the ground-truth heatmap as
represents a pixel location. A predicted heatmap can be obtained from the descriptor image together with a reference descriptor . Then the predicted heatmap is gotten by
The heatmap can also be normalized to sum to one, in which case it represents a probability distribution over the image.
The heatmap loss is simply the MSE between and with mean reduction.
2.2 Spatial Expectation Loss
Given a descriptor together with a descriptor image we can compute the 2D spatial expectation as
If we also have a depth image then we can define the spatial expectation over the depth channel as
The spatial expectation loss is simply the L1 loss between the ground truth and estimated correspondence using
We can also use our 3D spatial expectation to compute a 3D spatial expectation loss. In particular given a depth image let the depth value corresponding to pixel be denoted by . The spatial expectation loss is simply
being careful to only take the expectation over pixels with valid depth values .
2.3 Total Loss
The total loss is a combination of the heatmap loss and the spatial loss
3 Correspondence Function
Given a reference descriptor the correspondence function in Section 3.2 is computed using the spatial expectations in Equations (10), (11). These spatial expectations localize the descriptor in either pixel space or 3D space. If in 3D we additionally use the known camera extrinsics to express the localized point in world frame.
If we define as the estimated pixel location of the correspondence for descriptor , then our visual model can additionally provide a confidence scores that is a valid correspondence for this descriptor. The confidence is defined as the value of the unnormalized heatmap (see Equation (7)) at the pixel location . Specifically the confidence is given by
Note that this value always lies in the range $$.
Training Details
This section provides details on the simulation and hardware experiments.
2 Training Details
All methods used the same architecture for the dynamics model , an MLP with two hidden layers of units. All of our variants, along with the transporter baselines, use visual pre-training. The visual models are trained for 100 epochs and the model with the best test error is used. The dense-correspondence model uses both camera views for visual pretraining, while the transporter model uses only the images from the camera used at test time. For the dynamics learning all methods are trained for epochs using an Adam optimizer with a learning rate of . For each method the model with the best test error was used for evaluation. Table 3 details the set of learnable parameters for each method.
Online Model-Predictive Control
Following we use the model-predictive path integral (MPPI) approach derived in . Here we provide a brief overview but refer the reader to for more details. MPPI is a gradient-free optimizer that considers coordination between timesteps when sampling action trajectories. The algorithm proceeds by sampling trajectories, rolling them out using the learned model, computing the reward/cost for each trajectory, and then re-weighting the trajectories in order to sample a new set of trajectories. Let be look-ahead horizon of the MPC, then a single trajectory consists of state-action pairs . Let be the reward of the -th trajectory. Define
A filtering technique is then used to sample new trajectories from the previously computed mean . Specifically
where the noise is sampled via
This procedure is repeated for iterations at which point the best action sequence is selected. All of our experiments we used . The cost/reward in the MPC objective varied slightly between the hardware and simulation experiments, more details are provided below.
Simulation Experiments
To evaluate our method we consider four manipulation tasks in simulation. We use the Drake simulation environment which provides both the underlying physics simulation and rendering of RGBD images at VGA resolution . Figure 3 shows image from the four simulation tasks that we consider.
top down camera: This environment, depicted in Figure 3 (a), consists of the sugar-box object from the YCB dataset laying flat on a table. The robot is represented as cylindrical pusher (shown in green) and the action is the velocity of the pusher in the plane. The environment timestep is , so the agent must command actions at 10Hz. Two cameras are placed directly above the table, facing downwards. The camera positions are offset by degrees about their z-axis. Our methods use both camera feeds for training the visual correspondence model, but only one camera feed at test time. An image from this camera is shown in Figure 3 (a). All other methods use only a single camera feed at both train and test time.
angled camera: This environment is identical to top down camera but has different camera positions. Instead of being top down the two cameras are located on adjacent sides of the table and angled at 45 degrees, see figure 3 (b). The setting of angled cameras is more similar to our hardware experimental setup and is useful for comparing approaches that use pixel space vs. 3D representations.
occlusions: This environment uses the same setup of task angled camera the only difference being that the object is now laying on its side, see Figure 3 (c). This, together with the angled camera position, means that occlusions become a significant factor. In particular as the box rotates through the full degrees in yaw, the sides of the box become alternately occluded or visible. The top face of the box is the only one that remains unoccluded for all poses of the object.
mugs (category): This environment has the same top-down camera placements and cylindrical pusher as task angled camera. Instead of a single object however, we use a collection of 10 different mug models and vary the color and texture on each episode. This environment tests category-level vision and dynamics generalization. Two mug instances are shown in Figures 3 (d) and (e).
For each environment we collect a static dataset that is then used to learn the visual dynamics model. All methods have access to exactly the same dataset and the visual pretraining for our method and the transporter baseline is done using this same dataset. For each task the dataset is generated by collecting 500 trajectories of length 40 using a scripted random policy. The simulator timestep is seconds so a trajectory of length equates to seconds.
2 Evaluating closed-Loop MPC performance
For each environment we evaluate the different methods by planning to a desired goal-state image and computing the pose error (both translation and rotation) using the ground truth simulator state. Goal states are generated sampling a random control input and applying it to the environment for time steps. We further require that goal states are sufficiently far from initial states (in both translation and rotation). This generates a diverse set of initial and goal state pairs for evaluation. The simulator state is then reset to the initial state and we use closed-loop MPC to control the system to the goal state. The MPC cost function is simply the L2 distance between the latent state and the goal state. Ground truth state information is used to compute the error between the final and goal poses for the object.
3 Baselines
To demonstrate the benefits of our approach over prior methods we compare against several baselines.
Ground Truth 3D points (GT_3D): This baseline is used for tasks top down camera, angled camera, occlusions since those environment use just a single object. contains ground truth world-frame 3D locations of 4 points on the object. Knowing the location of 4 points is equivalent to knowing the object pose for a rigid object. We believe that this is a strong baseline that provides an upper bound on what is achievable with our descriptor-based methods that attempt to track points on the object.
Transporter: We use the Transporter autoencoder formulation from to pre-train a visual model. Following the original paper we use keypoints and freeze the visual model while training the dynamics model. We investigate two variants using the transporter approach. In transporter 2D are the pixel-space locations of the keypoints. In transporter 3D are the 3D world frame locations of the keypoints, computed from the pixel space by using the depth image together with the camera intrinsics and extrinsics.
Autoencoder: This method jointly learns the visual model and the dynamics model. Specifically we jointly train a convolutional autoencoder together with a forward dynamics model. The loss is a combination of the dynamics loss, Equation (1), and an image reconstruction loss, which penalizes the L2 distance between the reconstructed and actual images. Note that this is exactly the autoencoder baseline from . Following images are downsampled to before being passed into the network. The encoder architecture contains 6 2D convolutions with kernel sizes and filter sizes $\bm{z}_{\text{object}}\bm{z}1664646464\times 64$.
Hardware Experiments
We used a Kuka IIWA LBR robot with a custom cylindrical pusher attached to the end-effector to perform our hardware experiments, see Figure 5. RGBD sensing was provided by two RealSense D415 cameras rigidly mounted offboard the robot and calibrated to the robot’s coordinate frame. To enable effective correspondence learning between views, it is ideal to have views with some overlap such that correspondences exist, but still maintain different-enough viewpoints from each camera. At test time only a single camera is used to localize the dense-descriptor keypoints. The robot is controlled by commanding end-effector velocity in the plane at 5Hz. A high-rate Jacobian space controller consumes these 5Hz end-effector velocity commands and closes the loop to command the robot’s joint positions at 200Hz.
2 One-Shot Imitation Learning
Although our learned dynamics model together with online MPC is able to plan over short to medium horizons, we can track much longer horizon plans by providing a single demonstration and using a trajectory tracking cost in our MPC formulation. This demonstration can in principle come from any source, in our case we used a human teleoperating the robot. We capture observations throughout the trajectory at resulting in a trajectory of observations . Using our visual model we convert these observations into keypoints and latent state trajectories. These trajectories are used to guide the MPC. In particular the cost/reward function in the MPC is
This trajectory cost allows us to accurately track long horizon plans (where the demonstrations are as long as 15 seconds) using an MPC horizon of seconds. The four demonstrations trajectories used in the hardware experiments are illustrated in Figure 6.
3 Results
In this section we provide more details on the hardware experiments from Section 4.4. Figure 7 expands on Figure 4 showing the region of attraction of our MPC controller when attempting to stabilize the four different trajectories shown in Figure 6. We define a trajectory as a success if the final object position is within cm and degrees of the goal position. Given that during trials we explicitly chose initial conditions to test the region of attraction of the MPC controller, success rates are not particularly meaningful, as the success rate depends on the initial condition. Table 4 shows the average translational and angular errors among successful trajectories.