Self-Supervised Visual Planning with Temporal Skip Connections
Frederik Ebert, Chelsea Finn, Alex X. Lee, Sergey Levine
Introduction
A key bottleneck in enabling robots to autonomously learn a wide range of skills is the need for human involvement, for example, through hand-specifying reward functions, providing demonstrations, or resetting and arranging the environment for episodic learning. Learning action-conditioned predictive models of the environment is a promising approach for self-supervised robot learning as it allows the robot to learn entirely on its own from autonomously collected data. In order for a robot to be able to predict what will happen in response to its actions, it needs a representation of the environment that is suitable for prediction. In complex, open-world environments, it is difficult to construct a concise and sufficient representation for prediction. At a high level, the number and types of objects in the scene might change substantially trial to trial, making it impossible to provide a single fixed-size representation. At a low level, the type of information that is important about each object might change from object to object: the motion of a rigid box might depend only on its shape and friction coefficient, while the motion of a deformable stuffed animal might depend on a variety of other factors which are hard to determine by hand. Instead of engineering representations for different object types and object scenes, we can instead directly predict the robot’s sensory observations. The rationale behind this is that, if the robot is able to predict future observations, it has acquired a sufficient understanding of the environment that it can leverage for planning actions.
A major challenge in directly predicting visual observations for robotic control lies in deducing the spatial arrangement of objects from ambiguous observations. For example, when one object passes in front of another, the robot must remember that the occluded object still persists. In human visual perception, this is referred to as object permanence, and is known to take several months to emerge in infants . When the robot is commanded to manipulate an object which can become occluded during a manipulation, it must be able to accurately predict how that object will respond to an occlusion. In this work, we propose a visual predictive model for robotic control that can reason about spatial arrangements of objects in 3D, while using only monocular images and without providing any form of camera calibration. Previous learning methods have proposed to explicitly model object motion in image-space or to explicitly model 3D motion using point cloud measurements from depth cameras , and lacked the capability to maintain information about objects which are occluded during the predicted motion. We propose a simple model that does not require explicitly deducing full 3D structure in the scene, but does provide effective handling of occlusions by storing the appearance of occluded objects in memory. Furthermore, unlike prior prediction methods [4, optical_flow_pred] our model does not require any additional, external systems for providing point-to-point correspondences between frames of different time-steps.
The technical contribution of this work is three-fold. First, we present a video prediction method that can more accurately maintain object permanence through occlusions, by incorporating temporal skip connections. Second, we propose a planning objective for control through video prediction that leads to significantly improved long-term planning performance, including planning through occlusions, when compared to prior work . Finally, we propose a mechanism for planning with both discrete and continuous actions with video prediction models. Our evaluation demonstrations that these components can be combined to enable a learned video prediction model to perform a range of real-world pushing tasks. Our experiments include manipulation of previously unseen objects, handling multiple objects, pushing objects around obstructions, and moving the arm around and over other obstacle-objects, representing a significant advance in the range and complexity of skills that can be acquired through entirely self-supervised learning.
Related Work
Large-scale robotic data collection has been explored in a number of recent works. Oberlin and Tellex proposed to collect object scans using multiple robots. Several prior works have focused on autonomous data collection for individual skills, such as grasping or obstacle avoidance . In contrast to these methods, our approach learns predictive models that can be used to perform a variety of manipulations, and does not require a success measure or reward function during data collection. Several prior approaches have also sought to learn inverse or forward models from raw sensory data without any supervision . While these methods demonstrated effective generalization to new objects, they were limited in the complexity of tasks and time-scale at which these tasks could be performed. The method proposed by Agrawal et al. was able to plan single pokes, and then greedily execute multiple pokes in sequence. The method of Finn and Levine performed long-horizon planning, but was only effective for short motions. In our comparisons to this method, we demonstrate a substantial improvement in the length and complexity of manipulations that can be performed with our models.
Sensory Prediction Models.
Action-conditioned video prediction has been explored in the context of synthetic video game images and robotic manipulation , and video prediction without actions has been studied for unstructured videos and driving .
Several works have sought to use more complex distributions for future images, for example by using autoregressive models . While this often produces sharp predictions, the resulting models are extremely demanding computationally, and have not been applied to real-world robotic control. In this work, we extend video prediction methods that are based on predicting a transformation from the previous image . Prior work has also sought to predict motion directly in 3D, using 3D point clouds obtained from a depth camera , requiring point-to-point correspondences over time, which makes it hard to apply to previously unseen objects. Our predictive model is effective for a wide range of real-world object manipulations and does not require 3D depth sensing or point-to-point correspondences between frames. Prior work has also proposed to plan through learned models via differentiation, though not with visual inputs . We instead use a stochastic, sampling-based planning method , which we extend to handle a mixture of continuous and discrete actions.
Preliminaries
In this section, we define our image-based robotic control problem, present a formulation of visual model predictive control (visual MPC) over pixel motion, and summarize prior video prediction models based on image transformation.
2 Video Prediction via Pixel Transformations
Skip Connection Neural Advection Model
To enable effective tracking of objects through occlusions, we propose an extension to the dynamic neural advection (DNA) model that incorporates temporal skip connections. This model uses a similar multilayer convolutional LSTM structure; however, the transformations are now conditioned on a context of previous images. In the most generic version, this involves conditioning the transformation at time on all of the previous images , though in practice we found that a greatly simplified version of this model performed just as well in practice. We will therefore first present the generic model, and then describe the practical simplifications. We refer to this model as the skip connection neural advection model (SNA), since it handles occlusions by copying pixels from prior images in the history such that when a pixel is occluded (e.g., by the robot arm or by another object) it can still reappear later in the sequence. When predicting the next image , the generic SNA model transforms each image in the history according to a different transformation and with different masks to produce (see Figure 10 in the appendix):
In the case where , negative values of simply reuse the first image in the sequence. This generic formulation can be computationally expensive, since the number of masks and transformations scales with . A more tractable model, which we found works comparably well in practice in our robotic manipulation setting, assumes that occluded objects are typically static throughout the prediction horizon. This assumption allows us to dispense the intermediate transformations and only provide a skip connection from the very first image in the sequence, which is also the only real image, since all of the subsequent images are predicted by the model:
Visual MPC with Pixel Distance Costs
The choice of objective function for visual MPC has a large impact on the performance of the method. Intuitively, as the model’s predictions get more uncertain further into the future, the planner relies more and more on the cost function to provide a reasonable estimate of the overall distance to the goal. Prior work used the probability of the chosen pixel(s) reaching their goal position(s) after time steps, as discussed in Section 3.1. When the trajectories needed to reach the goal are long, the objective provides relatively little information about the progress towards the goal, since most of the predicted probabilities will have a value close to zero at the goal-pixel. In this work, we propose a cost function that still makes use of the uncertainty estimates about the pixel position, but provides a substantially smoother planning objective, resulting in improved performance for more complex, longer-horizon tasks. A straightforward choice of smooth cost function in deterministic settings is the Euclidean distance between the current and desired location of the desired pixel(s), given by . Since our video prediction model produces a distribution over the pixel location at each time step, we can use the expected value of this distance as a cost function, summed over the entire horizon :
where the expectation can be computed by summing over all of the positions in each predicted image, and the cost corresponds to a (element-wise) Hadamard product of the pixel location probability map and the distances between each pixel position and the goal. This cost function encourages the movement of the designated objects in the right direction for each step of the execution, regardless of whether the position can be reached within time steps. For multi-objective tasks with multiple designated pixels the costs are summed to together weighting them equally. Although the use of well-shaped cost functions in MPC has been explored extensively in prior work , the combination of cost shaping and visual MPC has not been studied extensively.
Sampling-Based MPC with Continuous and Discrete Actions
The choice of action representation for visual MPC has a significant effect on the performance of the resulting controller. The action representation must allow the planner sufficient freedom to maneuver the arm to perform a wide variety of tasks, while constraining sufficiently so as to create a tractable search space. We found that a particularly well-suited action representation for tabletop manipulation can be constructed by combining continuous and discrete actions, in contrast to prior work that used only continuous end-effector motion vectors as actions . The actions consist of the horizontal motion of the end-effector, as well as a discrete “lift” action that allows the robot to command the end-effector to lift off the table vertically. The discrete lifting action can take on values (4 in our implementation) that specify for how many time steps the robot should lift the end-effector off the table. Unlike with continuous vertical motion commands, these discrete commands result in the end-effector staying off the table for multiple time steps even during random data collection. In order to incorporate this hybrid action space into the stochastic CEM-based optimization in visual MPC, we sample real-valued parameters for the discrete action, and then round them to the nearest valid integer to obtain discrete actions. As usual with CEM, we iteratively refit a multivariate Gaussian distribution to the best performing samples (which in our case are chosen to be in the 90 percentile of samples), treating the entire action vector as if it was continuous. Figure 6 shows an example of a situation where the discrete vertical motion component of the action is used by our model to lift the gripper over the object in order to push it from the opposite side.
Experiments
Our experimental evaluation compares the proposed occlusion-aware SNA video prediction model, as well as the improved cost function for planning, with a previously proposed model based on dynamic neural advection (DNA) . We use a Sawyer robot, shown in Figure 1, to push a variety of objects in a tabletop setting. In the appendix, we include a details on hyperparameters, analysis of sample complexity, and a discussion of robustness and limitations. We evaluate long pushes and multi-objective tasks where one object must be pushed without disturbing another. The supplementary video and links to the code and data are available at https://sites.google.com/view/sna-visual-mpc
Figure 7 shows an example task for the pushing benchmark. We collected 20 trajectories with 3 novel objects and 1 training object. Table 1 shows the results for the pushing benchmark. The column distance refers to the mean distance between the goal pixel and the designated pixel at the final time-step. The column improvement indicates how much the designated pixel of the objects could be moved closer to their goal (or further away for negative values) compared to the starting location. The true locations of the designated pixels after pushing were annotated by a human.
The results in Table 1 show that our proposed planning cost in Equation (missing) 3 substantially outperforms the planning cost used in prior work . The performance of the SNA model in these experiments is comparable to the prior DNA model when both use the new planning cost, since this task does not involve any occlusions. Although the DNA model has a slightly better mean distance compared to SNA, it is well within the standard deviation, suggesting that the difference is not significant.
To examine how well each approach can handle occlusions, we devised a second task that requires the robot to push one object, while keeping another object stationary. When the stationary object is in the way, the robot must move the goal object around it, as shown in Figure 8 on the left. While doing this, the gripper may occlude the stationary object, and the task can only be performed successfully if the model can make accurate predictions through this occlusion. These tasks are specified by picking one pixel on the target object, and one on the obstacle. The obstacle is commanded to remain stationary, while the target object destination location is chosen on the other side of the obstacle.
We used four different object arrangements, with two training objects and two objects that were unseen during training. We found that, in most of the cases, the SNA model was able to find a valid trajectory, while the prior DNA model was mostly unable to find a solution. Figure 8 shows an example of the SNA model successfully predicting the position of the obstacle through an occlusion and finding a trajectory that avoids the obstacle. These findings are reflected by our quantitative results shown in Table 2, indicating the importance of temporal skip connections.
Discussion and Future Work
We showed that visual predictive models trained entirely with videos from random pushing motions can be leveraged to build a model-predictive control scheme that is able to solve a wide range multi-objective pushing tasks in spite of occlusions. We also demonstrated that we can combine both discrete and continuous actions in an action-conditioned video prediction framework to perform more complex behaviors, such as lifting the gripper to move over objects.
Although our method achieves significant improvement over prior work, it does have a number of limitations. The behaviors in our experiments are relatively short. In principle, the visual MPC approach can allow the robot to repeatedly retry the task until it succeeds, but the ability to retry is limited by the model’s ability to track the target pixel: the tracking deteriorates over time, and although our model achieves substantially better tracking through occlusions than prior work, repeated occlusions still cause it to lose track. Improving the quality of visual tracking of the designated pixels may allow the system to retry the task until it succeeds. More complex behaviors, such as picking and placing (e.g., to arrange a table setting), may also be difficult to learn with only randomly collected data. We expect that more goal-directed data collection would substantially improve the model’s ability to perform complex tasks. Furthermore, better predictive models that incorporate hierarchical structure or reason at variable time scales would further improve the capabilities of visual MPC to carry out temporally extended tasks. Fortunately, as video prediction methods continue to improve, we expect methods such as ours to further improve in their capability.
Acknowledgments
We would like to thank Prof. Dr. Patrick van der Smagt from the Technical University of Munich for insightful technical discussions and helping to organize Frederik Ebert’s visit at UC Berkeley. We would also like to thank Roberto Calandra for assistance the robotic experiments and 3D printing. Furthermore we would like to thank Ashvin Nair and Pulkit Agrawal for insightful discussions. This work was supported by a fellowship within the FITweltweit program of the German Academic Exchange Service (DAAD), a stipend from the German Industrial Foundation (Stiftung Industrieforschung), support from Siemens, the National Science Foundation through IIS-1614653 and IIS-1651843, and an equipment donation from NVIDIA. Chelsea Finn and Alex X. Lee were also partially supported by National Science Foundation Graduate Research Fellowships.
References
Appendix
The video prediction network was trained on sequences of 15 steps taken from 44,000 trajectories of length 30 by randomly shifting the 15-step long window thus providing a form of data augmentation. The model was trained on 66,000 iterations with a minibatch-size of 32, using a learing rate of 0.001 with the Adam optimizer [ADAM]. For visual MPC we use CEM-iterations. At every CEM iteration action sequences are sampled and the best samples are selected, fitting a Gaussian to these examples. Then new actions are sampled according to the fitted distribution and the procedure is repeated for iterations.
Prototyping in Simulation
The visual MPC algorithm has been prototyped and tested on a simple block pushing task simulated in the MuJoCo physics engine. The code has been made available in the same repository. Using the simulator, training data could be collected orders of magnitude faster. Furthermore a benchmark was set up which does not require any manual reset or manual labeling of the objects’ positions. The main downside of simulation however is that the complexity of both real-world dynamics and real-world visual scenes cannot be matched. (More realistic simulators exist, but require very large computational resources and large amounts of hand-engineering for setup).
Dependence of Model Performance on Amount of Training Data
We tested the effect of using different amounts of training data and compare the results visually, see Figure 9. The complete size of the data set is 44,000 trajectories. For this evaluation all models were trained for 76k iterations. The quality of the predictions for the probability distribution of the designated pixel vary widely with different amounts of training data and the quality of these predictions is crucial for the performance of visual MPC. The fact that performance still improves with very large amounts of training data suggests that the model has sufficient expressive power to leverage this amount of experience. In the future we will investigate the effect of using even more data and evaluate the effect of a larger number of different training objects.
Failure Cases and Limitations
Visual MPC fails when objects have vastly different appearance than those used in training. For example when objects are much bigger than any of the training objects these methods tend to perform poorly. Also we observed failure for one object with a very bright color which did not occur in the training set. However the model usually generalizes well to novel objects of similar size and appearance as the training objects. Apart from that we observed two main failure modes: The first one occurs when the probability mass of the designated pixel, drifts away from the object of interest (as discussed in the Section 8) since there is no feedback correcting the position of the designated pixel. To alleviate this problem a tracker can be integrated into the system enabling visual MPC to keep retrying when object behave differently than expected.
The second failure mode occurs when visual MPC does not find an action sequence which moves the designated pixel closer to the goal. This can happen, when the goalpoint for pushing the object is far away and it is very unlikely to sample a trajectory moving closer to the goal. In some cases the problem can be avoided simply by increasing the number of samples used when applying CEM. However this slows down the planning process. In order to enable more temporally extended action sequences a potential solution could be to increase the efficiency of the sampling process e.g. by introducing macro-actions which consist of a sequence of actions (learning macro-actions is out of the scope of this paper).
Robustness of the Model
We did not find any negative influence of clutter in the workspace, as long as the object to be moved has a free way path to its goal. When an obstacle is on the path it can be marked by a second designated pixel to avoid collision using an additional cost function. In this case the planner usually manages to find a way around the obstacle. However visual MPC with the current type of model is not yet capable of reasoning by itself that it has to push and object around obstacles when it is not marked explicitly. The reason is that the model has large uncertainty when multiple collisions occur (i.e. when the arm pushes object 1 and object 1 collides with object 2).
Our model is robust to small changes of the viewpoint (in the order of several centimetres in translation and several degrees in orientation). A likely reason is that during data collection the camera has been slightly displaced in position. Robustness could be improved by training a model from several viewpoints at the same time. As the prediction problem also becomes harder in this way more data might be required.
General skip connection neural advection model
We show a diagram of the general skip connection neural advection model in Figure 10
Alternative Model Architectures
Instead of copying pixels from the first image of the sequence , the pixels in the first image can also be transformed and then merged together with the transformed pixels from the previous time step as