O2O-Afford: Annotation-Free Large-Scale Object-Object Affordance Learning

Kaichun Mo, Yuzhe Qin, Fanbo Xiang, Hao Su, Leonidas Guibas

Introduction

We humans accomplish everyday tasks by interacting with a wide range of objects. Besides mastering the skills of manipulating objects using our fingers (e.g. grasping ), we must also understand the rich space of object-object interactions. For instance, to put a book inside a bookshelf, not only do we need to pick up the book by our hand, to place it, we still have to figure out possible good slots — the distance between the two shelf boards must be larger than the book height, and the slot should afford the book at a suitable pose. For us humans, we can instantly form such understanding of object-object interactions after a glance of the scene. Theories in cognitive science conjecture that this is because human beings have learned certain priors for object-object interactions based on the shape and functionality of objects. Can intelligent robot agents acquire similar priors and skills?

While there is a plethora of literature studying agent-object interaction, very few works have studied the important task of object-object interaction. In an earlier work, Sun et al. proposed to use Bayesian network to model human-object-object interaction and performed experiments on a small-scale (six objects in total) labeled training data with relative motions of humans and the two objects. Another relevant work by Zhu et al. studied the problem of tool manipulation, which is an important special case of object-object interaction, and proposed a learning framework given RGB-D scanned human and object demonstration sequences as supervision signals. Both works modeled object-object interaction in a small scale and trained the models with human annotations or demonstrations. In contrary, we propose a large-scale solution by learning from simulated interactions without any need for human annotations or demonstrations.

In this paper, we consider the problem of learning object-object interaction priors (abbreviated as the O2O priors for brevity). Particularly, in our setup, we consider an acting object that is directly manipulated by robot actuators, and a 3D scene that will be interacted upon. We are interested in encoding and predicting the set of feasible geometric relationships when the acting object is afforded by objects in the scene along accomplishing a certain specified task. One important usage of our O2O priors is to reduce the search space for downstream planning tasks. For example, to place a big object inside a cabinet having a few drawers with various sizes, the O2O priors may rule out those small drawers without any interaction trials and identify the ones big enough for the motion planners. Then, for the identified drawers, the O2O priors may further propose where the acting object could be placed by considering other factors, such as collision avoidance with existing objects in the drawers.

Next, we introduce how we encode the O2O priors. We formulate a per-point affordance labeling framework (Fig. 1) that unifies representations for various kinds of object-object interaction tasks. Given as input the acting object in different geometry, orientations and sizes, along with a partial 3D scan of an existing scene, we produce a point-wise affordance heatmap on the scene that measures the likelihood for each point on the scene point cloud of successfully accomplishing the task.

The affordance labels for supervision are generated by simulating the object-scene interaction process. Using the SAPIEN physical simulator and large-scale 3D shape datasets , we build up a large-scale learning-from-interaction benchmark that covers a rich space of object-object interaction scenarios. Fig. 1 illustrates the four diverse tasks we use in this work, in which different visual, geometric and dynamic attributes are essential to be learned for accurately modeling the task semantics. For example, to place a jar on a messy table, one need to find a flat tabletop area with enough space considering the volume of the jar; to fit a mug inside a drawer, the size and height of the drawer has to be big enough to contain the mug; etc.

As a core technical contribution, we propose an object-kernel point convolution network to reason about detailed geometric constraints between the acting object and the scene. We perform large-scale training and evaluation over 1,785 shapes from 18 object categories. Experiments prove that our proposed method learns effective features reasoning about the contact geometric details and task semantics, and show promising results not only on novel shapes from the training categories, but also over unseen object categories and real-world data.

we revisit the important problem of object-object interaction and propose a large-scale annotation-free learning-from-interaction solution;

we propose a per-point affordance labeling framework, with an object-kernel point convolution network, to deal with various object-object interaction tasks;

we build up four benchmarking environments with unified specifications using SAPIEN and ShapeNet that covers various kinds of object-object interaction tasks;

we show that the learned visual priors provide meaningful semantics and generalize well to novel shapes, unseen object categories, and real-world data.

Related Works

Learning from Interaction. Annotating training data has always been a heavy burden for supervised learning tasks. In robotics community, there has been growing interest in scaling up data collection via self-supervised interaction. This approach has been widely used to learn robot manipulation skills , facilitate object representation learning , and improve the result of perception tasks, e.g. segmentation , pose estimation . However, collecting interaction experience by a real robot is slow and even unsafe. A surrogate solution is physical simulation. Recent works explore the possibility of using data collected from simulation to train perception model . By leveraging the interactive nature of physical simulation, researchers can also train network to reason object dynamics, e.g. mass , force , stability . Benefits from the large-scale ShapeNet models, we can collect various type of interaction data from simulators.

Agent-Object Interaction Affordance. Agent-object interaction affordance describes how agents may interact with objects. The most common kind is grasping affordance. Recent works formulate grasping as visual affordance detection problems that anticipate the success of a given grasp. An affordance detector predicts the graspable area from image or point cloud for robot grippers. Other works extend contact affordance from simple robot gripper to more complex human object interaction. However, most of the works require additional annotation as training data . Recently, it was shown that visual affordances can also be reasoned from human demonstration videos in a weakly-supervised manner . Recent works proposed automated methods for large-scale agent-object visual affordance learning.

Object-Object Interaction Affordance. Very few works have explored learning object-object interaction affordance. Sun et al. built an object-object relationship model and associated it with a human action mode. It shows that the learned affordance is beneficial for downstream robotic manipulation tasks. Another line of works on object-object affordance focus on one particular object relationship: tool manipulation , object placement , pouring , and cooking . However, all these works are performed on a limit number of objects. Most of these works also require human annotations or demonstrations. In our work, we conduct large-scale annotation-free affordance learning that covers various kinds of object-object interaction with diverse shapes and categories.

Problem Formulation

Task Definition and Data Generation

As summarized in Fig. 1, we consider four object-object interaction tasks: placement, fitting, pushing, stacking. While placement and stacking are commonly used in manipulation benchmark , the fitting and pushing tasks are also interesting for bin packing and tool manipulation applications . One may create more task environments depending on downstream applications. Although having different task semantics and requiring learning of distinctive geometric, semantic, or dynamic attributes, we are able to unify the task specifications to share the same framework.

Task Initialization and Inputs. Each task starts with creating a static scene, including one or many randomly selected ShapeNet models and a possible ground floor. Some objects in the scene may have articulated parts with certain starting part poses, depending on different tasks. The scene objects may be fixed to be always static (e.g. when we assume a cabinet is very heavy), or dynamic but of zero velocity at the beginning of the simulation (e.g. for an object on the ground to be pushed). In all cases, a camera takes a single snapshot of the scene to obtain a scene partial 3D scan SS as the input to the problem. Fig. 2 (a) illustrate example initialization scenarios in the four task environments. To interact with the scene, we then randomly fetch an acting object and initialize it with random orientation and size. We provide a complete 3D point cloud OO, in the same camera coordinate frame as the scene point cloud, as another input to the problem.

Simulated Interaction Trajectory. For each interaction trial, we execute a hard-coded short-term trajectory τT(pi)\tau_{\mathfrak{T}}({p_{i}}) to simulate the interaction between the acting object OO and the scene objects SS at positition pip_{i}. Every motion trajectory is very short, e.g. taking place within <0.1<0.1 unit length, so that it can be preceded by long-term task trajectories, to make the learned visual priors possibly useful for many downstream tasks. The white arrows in Fig. 2 (b) show the trajectory moving directions – the forward direction for pushing and the gravity direction for the other three environments. The executed hard-coded trajectories are always along straight lines, though the final object state trajectories may be of free forms due to the object collisions. Fig. 2 (c) present some example ending object states.

Applicable and Possible Regions. For every simulated interaction trial, we randomly pick pp over the regions where the task is applicable and possible to succeed. Some scene points may not be applicable for a specific task. For example, for the stacking task, only points on the ground are applicable since we have to put the acting object on the floor for the scene object to stack over. Among applicable points, we only try the positions that are possible to be successfully interacted and directly mark the impossible points as failed interactions. For example, impossible points include the positions whose normal directions are not nearly facing upwards for the placement and fitting tasks. Fig. 2 (d) illustrate example applicable and possible masks over the input scene geometry.

Metrics and Outcomes. For each interaction trial, the outcome could be either successful or failed, measured by task-specific metrics. The metric measures if the intended task semantics has been accomplished, by detecting state changes of the acting and scene objects during the interaction and at the end. For example, in Fig. 2 (c), the fitting example shows that the drawer will be driven to close to check if the acting object can be fitted inside the drawer. See Sec. 4.2 for detailed definitions.

2 Four Task Environments

Placement. Each scene is initialized with a static root object (e.g. a table) with 0∼\sim15 movable small item objects randomly placed on the root object to simulate a messy tabletop. The root object may have articulated parts, which are initialized to be closed or of a random starting pose with equal probabilities. The acting object is another small item object to be placed. All points are applicable on the scene, but only the positions with normal directions that are close enough to the world up-direction are possible. For the acting object center at start, we have an up-directional offset rz=sz/2+0.01r_{z}=s_{z}/2+0.01, where szs_{z} is the up-directional object size, so that the acting object is 0.010.01 unit length away from contacting the intended interaction position pp. The motion trajectory is along the gravity direction. The metric for success is: 1) the acting object has no collision at start; 2) the acting object finally stays on the countertop; and 3) the acting object drops off stably with no big orientation change.

Fitting. The scene contains only one static root object with articulated parts (e.g., doors, drawers). At least one articulated part is randomly opened, while the other parts may be closed or randomly opened with equal probabilities. The acting object is a small item object to be fitted inside the drawer or shelf board. All points except the countertop points are applicable, since placing the item on countertop is concerned by the placement task. The positions with normal directions close enough to the world up-direction are possible. The starting acting object center and the motion trajectory are the same as in the placement task. Besides the three criteria for placement, there is one additional checking point for the metric: the door or drawer can be closed containing the acting object without being blocked.

Pushing. The scene object is dynamically placed on an invisible ground, initialized with zero velocity, and guaranteed to be stable by itself. There is no part articulation allowed in this task. The acting object is a small item object to push the scene object. All points on the scene are applicable and possible. The starting acting object center has offsets rx=−sx/2−0.1,rz=sz/2+0.02r_{x}=-s_{x}/2-0.1,r_{z}=s_{z}/2+0.02 where sxs_{x} and szs_{z} are respectively the forward-directional and up-directional object sizes. The acting object moves along the forward-direction to push the scene object. The metric for success is: 1) the acting object has no collision at start; 2) there is a big enough motion of the scene object; 3) the actual moving direction is within 30∘30^{\circ} aligned with the forward-direction; and 4) the scene object does not topple.

Stacking. We first place a dynamic acting object stably with zero velocity on a visible ground. Then, we pick a second item as the scene object to stack over the acting object. We simulate the stacking interaction that drops the scene object on top of the acting object. For the cases that the two objects finally touch each other and stay stably on the ground, we consider a stacking task based on the final two object states. The scene object is initialized at the final stably stacking pose. All points on the ground are applicable and possible. The acting object center has an up-directional offset rz=sz/2r_{z}=s_{z}/2, where szs_{z} is the up-directional object size. The metric for success is that the acting object has no big pose (center+orientation) change before and after the stacking interaction.

Method for Affordance Prediction

We propose a unified 3D point-based method, with an object-kernel point convolution network, to tackle the various O2O-Afford tasks. Though very simple, this method is quite effective and efficient in reasoning about the detailed geometric contacts and constraints.

Fig. 3 illustrates the proposed pipeline. We describe each network module in details below.

Object-kernel Point Convolution. Our O2O-Afford tasks require reasoning about the contact geometric constraints between the two input point clouds. Thus, we design an object-kernel point convolution module that uses the acting object as an explicit object kernel to slide over a subsampled scene seed points and performs point convolution operation to aggregate per-point features between the acting object and the scene inputs. This design shares a similar spirit to the recently proposed Transporter networks , but we carefully curate it for the 3D point cloud convolution setting. One may think of a naive alternative of simply concatenating two point clouds together at every seed point and training a classifier. However, this is computationally too expensive due to several forwarding passes over the two input points clouds with the acting object positioned at different seed locations.

2 Training and Loss

The whole pipeline is trained in an end-to-end fashion, supervised by the simulated interaction trials with successful or failed outcomes. The scene and acting objects are selected randomly from the training data of the training object categories. We equally sample data from different object categories to address the data imbalance issue. We empirically find that having enough positive data samples (at least 20,000 for task) are essential for a successful training. We train individual networks for different tasks and use the standard binary cross entropy loss. We use n=10000n=10000, m=1000m=1000, k=1000k=1000, t=3t=3, and f1=f2=f3=fg=128f_{1}=f_{2}=f_{3}=f_{g}=128 in the experiments. See supplementary for more training details.

Experiments

We use the SAPIEN physical simulator , equipped with ShapeNet and PartNet models, to do the experiments. We evaluate our proposed pipeline and provide quantitative comparisons to three baselines. Experiments show that we successfully learned visual priors of object-object interaction affordance for various O2O-Afford tasks, and the learned representations generalize well to novel shapes, unseen object categories, and real-world data.

Our experiments use 1,785 ShapeNet models in total, covering 18 commonly seen indoor object categories. We randomly split the different object categories into 12 training ones and 6 test ones. Furthermore, the shapes in the training categories are separated into training and test shapes. In total, there are 867 training shapes from the training categories, 281 test shapes from the training categories, and 637 shapes from the test categories. During training, all the networks are trained on the same split of training shapes from the training categories. We then evaluate and compare the methods by evaluating on the test shapes from the training categories, to test the performance on novel shapes from known categories, and the shapes in the test categories, to measure how well the learned visual representations generalize to totally unseen object categories. See supplementary for more details.

2 Baselines and Metric

We compare to three baseline methods B-PosNor, B-Bbox, and B-3Branch. B-PosNor replaces the per-point scene feature with 3-dim position and 3-dim ground-truth normal, while B-Bbox uses a 6-dim axis-aligned bounding box extents to replace the acting object geometry input. We compare to these two baselines to validate that the extracted scene features contain more information than simple normal directions and that object geometry matters. B-3Branch implements a naive baseline that employs two PointNet++ branches to process the acting object and scene point clouds as well as an additional branch taking as input the seed point position. Comparison to this baseline can help illustrate the necessity of correlating the two input point clouds. We use a success threshold 0.5 and employ two commonly used metrics: F-score and Average-Precision (AP).

3 Results and Analysis

Table 1 shows the quantitative evaluations and comparisons. It is clear to see that our method performs better than the three baselines in most entries. We visualize our network predictions in Fig. 4, where we observe meaningful per-point affordance labeling heatmaps on both test shapes from the training categories and shapes in unseen test categories. For placement, the network learns to not only find the flat surface, but also avoid collisions from the existing objects. We observe similar patterns for fitting, with one additional learned constraint to find two shelves with enough height. For pushing, one should push an object in the middle to cause big enough motions and at the bottom to avoid toppling. For stacking, our network successfully learns where to place the acting object on the ground for the scene object to stack over. See supplementary for more results.

We also perform an ablation study in Table 7 to prove the effectiveness of the proposed object-kernel point convolution. We further illustrate in Fig. 7 that our network predictions are sensitive to the acting object size and orientation changes. In the top-row fitting example, increasing the size of the mug reduces the chance of putting it inside the drawer, as the drawer cannot be further closed containing a big mug. For the bottom-row placement example, we observe some detailed affordance heatmap changes while we rotate the cuboid-shaped acting object.

We directly try to apply our network trained on synthetic data to real-world 3D scans. Fig. 5 and Fig. F.10 in the supplementary shows some qualitative results. We observe that, though trained on synthetic data only, our network transfers to real-world collected data to reasonable degrees.

Conclusion

We revisited the important but underexplored problem of visual affordance learning for object-object interaction. Using state-of-the-art physical simulation and the available large-scale 3D shape datasets, we proposed a learning-from-interaction framework that automates object-object interaction affordance learning without the need of having any human annotations or demonstrations. Experiments show that we successfully learned visual affordance priors that generalize well to novel shapes, unseen object categories, and real-world data.

Limitations and Future Works. First, our method assumes uniform density for all the objects. Future works may annotate such physical attributes for more accurate results. Second, we train separate networks for different tasks. Future study could think of a way for joint training, as many features may be shared across tasks. Third, there are many more kinds of object-object interaction that we have not included in this paper. People may extend the framework to cope with more tasks.

Acknowledgements

This research was supported by NSF grant IIS-1763268, NSF grant RI-1763268, a grant from the Toyota Research Institute University 2.0 programToyota Research Institute (”TRI”) provided funds to assist the authors with their research but this article solely reflects the opinions and conclusions of its authors and not TRI or any other Toyota entity., a Vannevar Bush faculty fellowship, and gift money from Qualcomm. This work was also supported by AWS Machine Learning Awards Program.

References

Appendix A More Data Details and Visualization

In Table A.2, we summarize our data statistics. In Fig. A.8, we visualize our simulation assets from ShapeNet and PartNet that we use in this work.

There are two kinds of object categories: big heavy objects Cheavy\mathbf{\mathcal{C}_{heavy}} and small item objects Citem\mathbf{\mathcal{C}_{item}}. In our experiments, Cheavy\mathbf{\mathcal{C}_{heavy}} include cabinets, microwaves, tables, refrigerators, safes, and washing machines, while Citem\mathbf{\mathcal{C}_{item}} contains baskets, bottles, bowls, boxes, cans, pots, mugs, trash cans, buckets, dispensers, jars, and kettles. For the placement and fitting tasks, we use the object categories in Cheavy\mathbf{\mathcal{C}_{heavy}} to serve as the main scene object. And, we use the Citem\mathbf{\mathcal{C}_{item}} categories as the scene objects for the other two task environments, as well as employ them as the acting objects for all the four tasks. Some objects may contain articulated parts. For the objects in Cheavy\mathbf{\mathcal{C}_{heavy}}, we sample a random starting part pose that is either fully closed or randomly opened to random degree with equal probabilities. For the acting objects, we fix the part articulation at the rest state during the entire simulated interaction.

Appendix B More Details on Settings

For the physical simulation in SAPIEN , we use the default setting of frame rate 500 frame-per-second, solver iterations 20, standard gravity 9.81, static friction coefficient 4.0, dynamic friction coefficient 4.0, and restitution coefficient 0.01. The perspective camera is located at a random position determined by a random azimuth [0∘,360∘)[0^{\circ},360^{\circ}) and a random altitude [30∘,60∘][30^{\circ},60^{\circ}], facing towards the center of the scene point cloud with 5 unit length distanced away. It has field-of-view 35∘35^{\circ}, near plane 0.1, far plane 100, and resolution 448×\times448. We use the Three-point lighting with additional ambient lighting. The scene point cloud SS samples n=10,000n=10,000 points using Furthest Point Sampling (FPS) from the back-projected depth scan of the camera.

Appendix C More Details on Networks

For the feature extraction backbones Encscene\mathbf{Enc_{scene}} and Encobject\mathbf{Enc_{object}}, we use the segmentation-version PointNet++ with a hierarchical encoding stage, which gradually decreases point cloud resolution by several set abstraction layers, and a hierarchical decoding stage, which gradually expands back the resolution until reaching the original point cloud with feature propagation layers. There are skip links between the encoder and decoder. We use the single-scale grouping version of PointNet++. The two PointNet++ networks do not share weights.

More specifically, for Encscene\mathbf{Enc_{scene}}, we use four set abstraction layers with resolution 1024, 256, 64 and 16, with learnable Multilayer Perceptrons (MLPs) of sizes , , , and respectively. There are four corresponding feature propagation layers with MLP sizes , , , and . Finally, we use a linear layer that produces a point-wise 128-dim feature map for all scene points. We use ReLU activation functions and Batch-Norm layers.

For Encobject\mathbf{Enc_{object}}, we use three set abstraction layers with resolution 512, 128, and 1 (which means that we extract the global feature of the acting object OO), with learnable MLPs of sizes , , and respectively. There are three corresponding feature propagation layers with MLP sizes , , and . Finally, we use a linear layer that produces a 128-dim feature for every point of the acting object, and use another linear layer to obtain a 128-dim global acting object feature. We use ReLU activation functions and Batch-Norm layers.

We use a PointNet to implement the proposed object-kernel point convolution module Convobject\mathbf{Conv_{object}}. The network has three linear layers that transforms each point feature. We use ReLU activation functions and Batch-Norm layers. Finally, we apply a point-wise max-pooling operation to pool over the mm acting object points to obtain the aggregated 128-dim feature for every sampled scene seed point pip_{i}.

The final affordance prediction module Deccritic\mathbf{Dec_{critic}} is implemented with a simple MLP with two layers , with the hidden layer activated by Leaky ReLU and final layer without any activation function. We do not use Batch-Norm for either of the two layers.

Appendix D More Details on Training

We collect hundreds of thousands of interaction trials in the simulated task environments for each task. The scene and acting objects are selected randomly from the training data of the training object categories. We equally sample data from different object categories, to address the data imbalance issue. For a generated pair of scene SS and acting object OO, we then randomly pick an interacting position pip_{i} from the task-specific possible region to perform a simulated interaction. The task environment provides us the final interaction outcome, either successful or failed, using the task-specific metrics described in Sec. 4.2.

Using random data sampling gives us different positive data rate for different tasks, ranging from the lowest one 7.4% for pushing and the highest 41.2% for placement. We empirically find that having enough positive data samples are essential for a successful training. Thus, for each task, we make sure to sample at least 20,000 successful interaction trials. It is also important to equally sample positive and negative data points in every batch of training.

We use batch size 32 and learning rate 0.001 (decayed by 0.9 every 5000 steps). Each training takes roughly 1-2 days until convergence on a single NVIDIA Titan-XP GPU. The GPU memory cost is about 11 GB. At the test time, the speed is fast (on average 62.5 milliseconds per data) since it only requires a single feed-forwarding inference throughout the network. Testing over a batch of 64 needs 6 GB GPU memory.

Appendix E More Results

Fig. E.9 presents more results of our affordance heatmap predictions, to augment Fig. 4 in the main paper.

Appendix F More Results on Real-world Data

Fig. F.10 presents more results testing if our learned model can generalize to real-world data, to augment Fig. 5 in the main paper. We use the Replica dataset , the RBO dataset , and Google Scanned Objects as the input 3D scans.

For qualitative results over the real-world scans shown in Fig. 5 in the main paper, we use an iPad Pro to collect the real-world scans by ourselves. We first mount the camera on a fixed plane and rotate the iPad to make the camera view 20∘20^{\circ} top down to the ground. We use the front structured light camera on iPad to capture a cluttered table top and a microwave oven.

Our method is designed to learn visual priors for object-object interaction affordance. Future works may further finetune our priors predictions by training on some real-world tasks to obtain more accurate posterior results.

Appendix G Failure Cases: Discussion and Visualization

Fig. G.11 summarizes common failure cases of our method. See the caption for detailed explanations and discussions.

We hope that future works may study improving the performance regarding these matters.