SAPIEN: A SimulAted Part-based Interactive ENvironment

Fanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia, Hao Zhu, Fangchen Liu, Minghua Liu, Hanxiao Jiang, Yifu Yuan, He Wang, Li Yi, Angel X. Chang, Leonidas J. Guibas, Hao Su

Table of Contents

Appendix A Details on PartNet-Mobility Annotation System.

Appendix B Experiment details on movable part segmentation and motion recognition tasks.

Appendix A: Annotation System

We developed a web interface (Figure 1) for mobility annotation. This tool is a question answering (QA) system, which proposes questions based on current stage of annotation. It exploits the hierarchical structures of PartNet to propose objects without relative mobility, and generates new questions based on past annotations. Using this tool, annotators will not miss any movable parts if they answer every question correctly, and they will not face any redundant questions by design. The output mobility annotations are guaranteed to satisfy tree properties suitable for simulation.

The annotation procedure has the following steps:

We start with a PartNet semantic tree, and traverse the tree nodes. Annotators are prompted with questions asking if current subtree has relative motion. If it does not, all parts in this tree will be fixed together; otherwise, the same question is asked again on the child nodes of this subtree.

When the PartNet semantic tree traversal is finished, annotators are asked to choose parts that are fixed together.

Next, annotators are asked to choose parts that are connected with a hinge (rotational) joint. They will then choose parent-child relation, and annotate axis position/motion limit with our 3D annotation tool.

Next, annotators are asked to choose parts that are connected with a slider (translational) joint. They will similarly choose motion parameters and decide if this axis also bears rotation (screw joint).

Finally, annotators will annotate each separate object in the scene as “fixed base”, “free””, or “out lier”.

The procedure is summarized in the following pseudo-code block.

Appendix B: Movable Part Segmentation and Motion Recognition

Table 1 shows the movable part segmentation results for all categories in PartNet-Mobility dataset.

Motion Recognition: experiment details

For this task, we normalize the [0,2π][0,2\pi] hinge joint range to .Forsliders,wenormalizebythemaximummotionrangeoverthedatasettomakethemotionrangepredictionwithin. For sliders, we normalize by the maximum motion range over the dataset to make the motion range prediction within.

In the following, letters without hat indicates their corresponding ground-truth labels.

In our experiment, we modify the input layer of a ResNet50 network to accept 5 channels, and output layer to output 13 numbers. In addition, we apply tanh activation to produce d^r,d^t\hat{\mathbf{d}}^{r},\hat{\mathbf{d}}^{t}, and sigmoid activation to produce x^door,x^drawer\hat{x}_{\text{door}},\hat{x}_{\text{drawer}}. The loss has 7 terms: Axis alignment loss, measured by cosine distance:

Pivot loss, measured by the distance from predicted pivot to ground truth joint axis:

Joint position loss, L2L_{2} loss between predicted position and ground truth position.

The final loss is a summation of all the losses above:

This objective is optimized on mini-batches using proper masking based on HH and SS values.

We repeat this experiment with PointNet++qi2017pointnet++ operating on 3D RGB-point cloud produced by the same images. For each image, we sample 10,000 points from the partial point cloud (create random copies if the total number of points is less than 10,000). Figure 2 shows the network structure for the motion recognition tasks.

Appendix C: Terminology

Articulation: An articulation is composed of a set of links connected together with transnational or rotational joints physx. The most common articulation is a robot.

Kinematic/Dynamic joint system: Both joint systems are an assembly of rigid bodies connected by pairwise constraints. Kinematic system does not respond to external forces while dynamic objects do.

Force/Joint/Velocity Controller: Controller which can control the force/position/velocity of one or multiple joints at once. Like real robot, controller may fail depending on whether the target is reachable.

Inertial Measurement Unit(IMU): A sensor which can measure the orientation, acceleration and angular velocity of the mounted link.

Trajectory Controller: A controller which receive trajectory command and execute to move through the trajectory points. Note that trajectory consist of a sequence of position, velocity and acceleration, while path is simply a set of points without a schedule for reaching each point hutchinson1996tutorial.

End-effector: End-effector is a manipulator that performs the task required of the robot, The most common end-effector is gripper.

Inverse Kinematics: Determine the joint position corresponding to a given end-effector position and orientation siciliano2010robotics.

Inverse Dynamics: Determining the joint torques which are needed to generate a given motion. Usualy, the input of inverse dynamics is the output of inverse kinematics or motion planning.

SAPIEN Renderer

GLSL is OpenGL’s shading language with describes how the GPU draws visuals.

Rasterization is the process of converting shapes to pixels. It is the pipeline used by most real-time graphics applications.

Ray tracing is a rendering technique by simulating light-rays, reflections, refractions, etc. It can achieve physically accurate images at the cost of rendering time. OptiX is Nvidia’s GPU based ray-tracing framework.

References