IFOR: Iterative Flow Minimization for Robotic Object Rearrangement
Ankit Goyal, Arsalan Mousavian, Chris Paxton, Yu-Wei Chao, Brian Okorn, Jia Deng, Dieter Fox
Introduction
Object rearrangement is the capability of an embodied agent to physically re-configure the objects in a scene into a desired goal configuration . It is an essential skill in day-to-day activities like setting a dining table, putting away groceries, and organizing a desk. Endowing robots with this capability is crucial for deploying them to assist people with everyday tasks .
With varying task setups, the desired goal state can be provided in different forms, for instance, a compact state representation or natural language descriptions . In this work, we address the rearrangement task where the goal state is specified by an RGB-D image , as shown in Fig. LABEL:fig:teaser. This setup lends itself well to many scenarios where the goal state can be snapped once, either in the first place or from a one-time demonstration. For instance, a user can set the dining table once to their preference and take a photo, and a robot assistant can restore the table back to the desired state from any configuration.
Traditionally, object rearrangement problems have been studied in the robotics community, often in the context of Task and Motion Planning (TAMP) . Despite much recent progress , most TAMP approaches still rely on a strict set of assumptions on the perception front. First, the objects and scenes are often assumed to be known a priori, provided with high fidelity 3D models. This makes the approaches difficult to be deployed in unseen environments or environments without models. Second, given the models, the planning front often assumes accurate pose information at the input. This makes explicit object pose estimation a necessary part of the pipeline, and the full system susceptible to pose estimation error from real vision systems.
Recent efforts in robotics have attempted to relax these constraints by leveraging the power of deep learning. A recent approach called NeRP, proposed by Qureshi et al. , has allowed for rearranging objects unseen at the training time, by representing the observed objects with learned embeddings. It also removes the need of explicit object pose estimation for planning by leveraging recent progress on learning-based grasp planners and collision detectors . However, NeRP only allows moving objects with 2D in-plane translations on the table surface and allows no change in their orientation. This prevents its applications in realistic scenarios that require moving objects with more complex transformations such as those shown in Fig. LABEL:fig:teaser.
We propose a new approach to image-guided robotic object rearrangement with RGB-D input. It achieves, for the first time to the best our of knowledge, the ability to handle unknown objects with translation as well as planar rotations.
The key to our method is re-formulating object rearrangement as an iterative minimization of optical flow between the current observed image and the goal image. By using optical flow as an intermediate representation, we can capitalize on the cutting edge development in flow estimation models . Although these flow models were originally developed for consecutive video frames with small pixel displacements, we show that with proper training, these models can excel at estimating flow with large displacements from arbitrary transformation of objects. Using this estimated flow, together with the depth input and generic object segmentation models , we can obtain dense 3D correspondences for each object. This provides a general representation that allows us to solve for the desired transformation of objects with simple optimization. Furthermore, with such a general representation, our method trained entirely on synthetic data transfers well to the real world in a zero shot manner.
To summarize, we introduce IFOR, Iterative Flow Minimization for Unseen Object Rearrangement. IFOR is to our knowledge the first system capable of rearranging unseen objects, given an RGB-D image goal, that handles both translation and rotation. Our approach is trained solely on simulation data and transfers to the real world in a zero-shot manner. Finally, we perform a set of experiments showing our method allows rearrangement of novel objects in cluttered scenes with a real robot.
Related Work
Our work falls in the broad area of robotic manipulation. Conventional manipulation systems often take a modular approach, decomposing the full system into perception, planning, and actuation components. The perception module is charged with estimating the state of the environment, e.g., detecting and segmenting objects and estimating their 6D poses . With the perception output, the planning module then searches for a sequence of actions to accomplish the manipulation task. In robotics, this is conventionally formalized and studied in the problem of Task and Motion Planning (TAMP) . However, setting up a real-world TAMP system often requires substantial task-specific knowledge and accurate 3D models of the environment, significantly limiting the environments to which the system can generalize. To address this challenge, recent work has adopted deep learning-based approaches for robotic manipulation, for instance, on grasp planning , motion planning , and reasoning about spatial relations .
Our work is concerned with rearranging objects, an area that has a long history in robotics but has recently gained traction in the vision and learning communities thanks to the advances in simulation platforms. The works most relevant to ours are that of Labbé et al. and NeRP , which also address the rearrangement task with the goal state specified by an image. A key step to address this problem is to establish the correspondence of objects between the current and goal image and solve for their desired transformations. However, both and only consider 2D planar translation with no orientation change. This is arguably because their feature descriptors for objects are generic and not rotation sensitive. In contrast, we propose to use optical flow as the low-level feature descriptors, which can be naturally used to infer the full 6D transformations. In parallel to our work, recent efforts have also addressed rearrangement particularly learned from human demonstrations and also with different goal specifications such as language .
Optical Flow and Feature Correspondence.
Optical flow is a long-standing vision problem that addresses the estimation of pixel motions between two video frames. Alike other sub-fields in computer vision, the take-off of deep learning has replaced the traditional pipelines for optical flow with learning-based end-to-end architectures . In tackling object rearrangement, our work proposes to use the predicted optical flow between the current and goal image to establish the desired object transformations. However, rather than estimating flow from consecutive video frames, we demonstrate that state-of-the-art models can excel at predicting flow with large displacements from arbitrary goal images in object rearrangement.
In addition to optical flow, our work can potentially also leverage the body of work on scene flow, which directly predicts the motion in 3D rather than in the 2D image space. Recent work has studied various setups for scene flow including predicting from monocular frames , stereo images , RGB-D pairs , to 3D point clouds . Finally, our flow prediction task is closely related to the problem of establishing feature correspondences in 3D reconstruction (e.g., SfM) and visual localization (e.g., SLAM).
Embodied AI.
Our work fits well with recent trends in embodied AI. Initially, this work largely centered around the family of tasks pertained to navigation , but has gradually grown to encompass tasks pertained to physical manipulation . Moreover, a recent notable article by Batra et al. has recognized the rearrangement problem as a “canonical task” for evaluating embodied AI. Our work drives progress in this exact frontier. And unlike most prior embodied AI works, which run evaluation only in simulation, we also evaluate the performance of our method on a real-world robotic platform.
Method
IFOR takes as input the RGB-D images of the current and the goal scene, and iteratively generates a pick-and-place action for one object at a time. At each iteration, the RGB-D image of current and goal scene is passed through two components: (1) perception and (2) planning (Fig. 1).
The perception component is responsible for estimating the relative transformation of all objects between the current and goal scene. Given the estimated transforms, the planning component selects an object to be moved along with the required transformation, by taking into account collision and kinematic feasibility. Finally, after executing the planned pick-and-place action, the system will take a new observation of the scene and repeat the process.
The task of the perception component is to segment out objects in the current and goal images, establish their correspondences, and predict a 6-DoF transformation for each object from its current pose to its goal pose. We achieve this with a pipeline that combines optical flow estimation, unseen object segmentation, and a RANSAC-based transformation optimization.
The first step is to estimate the optical flow between the current and the goal image. This provides us with a pixel-level correspondence between them. In conventional settings, optical flow is generally estimated between temporally close images in a video. The displacement of the flow is often small as temporally close images do not differ much from one another. In fact, this small displacement prior is crucial for classical methods like Lucas-Kanade method . This assumption however does not hold in rearrangement, as the objects could be moved a large distance and rotated by large angles from the initial scenes. This makes classical approaches of optical flow estimation ill-suited for rearrangement.
Contrary to classical methods, deep learning based optical flow methods like Recurrent All Pairs Field Transforms (RAFT) learn to make predictions from data. But, they too were developed for and trained on estimating flow in videos and hence pre-trained RAFT model does not work well on rearrangement scenes. However, RAFT’s underlying architecture does not make any assumption about the flow displacement being small, as it compares every pixel in one image to every other pixel in the other image. Consequently, given suitable data, RAFT can be trained for object rearrangement with large translations and rotations.
RAFT.
Recurrent All Pairs Field Transforms (RAFT) estimates optical flow by constructing a 4D correlation volume of which each pixel in one image is compared to every pixel in the other image . It then updates the flow estimates using a recurrent unit, starting with zero optical flow at all locations. In each iteration, the recurrent unit does a lookup around the current estimate of flow to decide how to update the flow estimates.
During training, RAFT applies supervision over all these intermediate flow estimates made by the recurrent unit. Suppose, are the intermediate flow estimates and is the ground-truth flow, then the loss () is defined as the weighted sum of distance between estimated and ground-truth flow. Specifically,
where is a discount factor of value 0.8. For more details, please refer to the work by Teed an Deng .
Synthetic Data.
To train RAFT for object rearrangement, we create a visually realistic dataset of synthetic scenes. In these scenes, objects are placed on supports like tables or beds. We sample the supports from ShapeNet and the objects from the Google Scanned Dataset . We use the NViSII renderer to render scenes with realistic lighting via ray tracing, and diverse view points by randomizing camera poses. We also randomize the lighting of the scenes, the texture of the supports and the background image. In the end, we created a dataset of 54K training samples and 1000 test samples. Some samples of our synthetic data are shown in Fig. 2.
Unseen Object Segmentation.
Optical flow alone is not sufficient to estimate the transformation of the object as they lack the grouping or “objectness” information. Moreover, unlike most prior segmentation methods that only segment object classes they were trained on, we need a zero-shot segmentation method to deal with unseen objects at test time. We use pre-trained UCN , which is trained to segment unknown objects from RGB-D image of the scene. It learns per-pixel embedding such that pixels belonging to the same object instance have similar embeddings but different from other object instances in the scene. At the test time, objects are segmented using mean-shift clustering.
Relative Object Transform from Flow.
The next step is to estimate the relative transformation of the objects between the two frames. The relative transformation of each object is computed by first unprojecting each pixel to 3D from the depth map of the scene and camera intrinsics. The 3D correspondence between current image and goal image is computed from predicted flow and unprojected points.
In practice, we observed that the matched correspondences from flow can contain many outliers. This is especially prominent when the object undergoes extreme transformations (e.g., large rotations), resulting in only a small portion of the object surface that is visible in both images. To handle the outliers reliably, we adopt RANSAC in solving the relative pose. In Table 1, we show that RANSAC is effective in removing outliers and estimating accurate transformations. In Figure 3, we show qualitative examples of the transformations estimated by RANSAC using our trained RAFT model.
2 Planning and Execution
Given the list of desired object transformations, the planner produces a pick-and-place action to execute, while taking into account various kinematic and geometric constraints. Our planning algorithm iterates over the list of desired object transformation from perception module, and finds out which objects can be moved directly using the predicted transform. The planner classifies each relative transform as feasible if the object at the predicted transform is not colliding with any other objects in the scene. We used pre-trained SceneCollisionNet to check the collision of the object at the predicted transformation. The feasible objects are then ranked based on score , where r is the relative rotation transformation in radians, t is the relative translation in cm and in our experiments. The planning algorithm greedily picks objects with larger relative transformations.
If the policy is not able to find a feasible movement for any object, it will try to move one of the objects to a random collision-free location. Provided that the estimated transformation of the objects are correct, the proposed planning and execution policy is guaranteed to converge and successfully rearrange the objects, i.e., in the worse case, it will move all but one object to collision-free locations, followed by moving each of them to the goal location one by one.
The system terminates when the estimated change in rotation and translation for all objects is smaller than than a fixed threshold. Specifically, we use a threshold of for rotation and for translation. Our experiments show that this heuristic is quite effective in handling challenging object rearrangement scenarios in the real world (discussed in Sec. 4.1). The pseudo-code for the planning algorithm is outlined in Algo. 1.
Experiments
We present results in two settings. First, we perform an integrated system comparison in the real world, with a physical robot picking and placing objects using either IFOR or the previous state of the art . Second, we perform large scale ablation studies in simulation to evaluate the effect of each design choice in IFOR.
We used a Franka Panda robot to conduct the real world experiments. The world is perceived from an external RealSense L515 camera, and a wrist-mounted RealSense D415 camera. The external camera is used for planning with IFOR, collision avoidance, and controlling the robot. The wrist mounted camera is only used for grasping to improve the robustness of the system.
IFOR generates a list of objects and transforms representing the placement location. These poses are passed to the pick-and-place system which takes the ordered list of actions from the planner. It selects the first object that can be grasped and places it at the desired location. Grasps for each object are computed by Contact-GraspNet and the robot motions are generated using a model predictive control pipeline and SceneCollisionNet. Refer to for details of the pick and place system. None of the components are trained on any real objects nor the objects we tested on.
We evaluated on 6 scenes, where each scene has between 2 to 5 objects in the initial configuration and a distinct goal configuration. Although it is not possible to replicate the exact initial conditions in the real world, we tried to duplicate the initial configuration visually, and used the same RGB-D goal image for testing both methods. In order to quantitatively evaluate the performance of the methods in the real world, we conducted a user study with 10 users, where we asked users to select the method that performed better: IFOR or NeRP . We also asked users to rate IFOR and NeRP on a scale of 1-4, where 1 is “very bad“ and 4 is “very good“.
All the components of the pick-and-place system were the same for both IFOR and NeRP, except the estimation of the objects’ final pose. Fig. 4 shows a qualitative comparison of IFOR and NeRP given similar target images. The video of the experiments are included in the Supplementary Materials. Since NeRP does not handle changing the orientation, we asked users to rank the two methods only based on translation as well as considering both rotation and translation. Fig. 5 shows that users significantly prefer IFOR over NeRP both on the translation only setting and full pose setting. NeRP failures were often due to an incorrect correspondence between the current image and target image. In addition, the learned placement generator does not generalize well to arbitrary rearrangements. IFOR, on the other hand, can find corresponding objects and their relative pose transform reliably from the predicted optical flow.
2 Ablation Studies
We conducted our ablation studies on a synthetic dataset which consisted of 200 tabletop scenes with randomized lighting, texture and background images. Each scene contained between 1 and 9 objects. Given a random target scene, the current scene was generated by applying a random planar rotation and translation to each object. We first define the evaluation metrics below and then provide the analysis on different ablation studies.
We report the median translation and rotation error averaged over all the objects. We also report the percentage of objects that are within a threshold of rotation and position error for different thresholds. In addition, for Table 2, we divide the scenes by three task difficulty levels: “easy“ for position error less than 2 cm and rotation error less than ; “medium“ for position error less than 5 cm and rotation error less than ; and “hard“ for position error less than 10 cm and rotation error less than .
Learning-based versus RANSAC for relative transform prediction.
We explored learning based solution to predict transformation from flow instead of optimizing it with RANSAC. For estimating location, we predict the heat-map of the centroid of the object in the goal scene given the flow and depth image of the scene. For estimating rotation, we use a cropped image of the flow around the object and regress the change in rotation. Table 1 compares the accuracy of the learning-based baseline with our RANSAC-based optimization. Our RANSAC-based optimization with re-trained RAFT (rearrangement RAFT) achieves significantly lower errors.
Pre-trained RAFT versus Rearrangement RAFT.
We compared the accuracy of single step transformation prediction between pre-trained RAFT which is trained on video sequences vs rearrangement RAFT where it is trained on our synthetic data with large translation and orientation changes. Table 1 shows that training RAFT on rearrangment scenes is cruicial for achieving high accuracy for predicting relative transforms of the objects.
Teleporting objects versus planning one action at the time.
In order to provide an approximate upper-bound performance of IFOR, we implemented a baseline where the policy computes transformations for all the objects and move all of them at once and observe the scene again. Fig. 6 shows the execution of teleport policy in simulation. This policy is clearly not practical in any real robotic setting but removes planning constraints and achieves significantly better results as shown in first row of Table 2. In addition, we can see the the heuristic planning does not lose the performance compared to the teleport policy which shows its effectiveness.
Ground-truth Collision Checker versus Learned Collision Checker [7].
An important component of IFOR is the collision checker which indicates whether a predicted rotation and translation should be accepted by the policy or not. Table 2 shows that the placement accuracy drops due to imperfections of the learned collision checker. IFOR benefits from improvements in collision checking for arbitrary scenes.
Planar Rotation versus SO(3) Rotation.
Given the correspondences in 3D, we optimize for relative transform using RANSAC. However, we can use the knowledge that objects rotations are planar and only keep the rotation component around z-axis. Table 2 shows that adding planar rotation assumption improves the accuracy significantly.
Object-wise Performance Analysis.
In order to get more insight into the performance of the system, we divided the objects of into 7 classes that are shown in Fig. 7. The general trend among the categories is that the accuracy improves with the increase in the number of steps which stems from having more overlapping views of the object. The category that IFOR struggles with most is bottle. Bottles are challenging because the method needs to match not only the geometry of object but also needs to match the texture details on the surface of the bottles, which can be quite challenging for smaller objects where there is not enough resolution on those objects. We can see a similar issue with the dish category, which includes pots, pans, and mugs. Unlike with bottles, these objects tend to have little to no texture and are mostly symmetric.
Discussion
We proposed IFOR, an iterative optical flow based system for object rearrangement. To the best of our knowledge, IFOR is the first method that can solve the image-based object rearrangement task, with both translation and rotation, for unseen objects. We conducted experiments with a robot in the real world to show its effectiveness and generalization to real perception data. Our results show a significant performance boost over previous work. In addition, we conducted a comprehensive analysis on simulation data to evaluate the effect of each component of our approach.
There are multiple promising future directions for this work. The first is to extend IFOR to poses. Even though there is no inherent assumption about planar rotation in our method, we only evaluated the object rearrangement problem with planar rotations because the models are only trained on data with planar rotations.
Another interesting direction would be to monitor the scene during robot execution to account for any in-gripper object motions and update the placement pose accordingly.