ACID: Action-Conditional Implicit Visual Dynamics for Deformable Object Manipulation

Bokui Shen, Zhenyu Jiang, Christopher Choy, Leonidas J. Guibas, Silvio Savarese, Anima Anandkumar, Yuke Zhu

I Metrics

In this section, we provide the formal definitions of the metrics that we use for evaluation. We define mIoU for geometry evaluation; vis and full mean-squared error for dynamics evaluation; FMR and acc for correspondence evaluation; Kendall’s τ\tau for ranking evaluation; F-score and Chamfer for planning execution evaluation. dcorrd_{corr} and success rate have been defined in the main paper.

mIoU: we follow Peng et al.[peng2020convolutional] for volumetric interaction over union calculation. Let Spred\mathcal{S}_{pred} and SGT\mathcal{S}_{GT} be the set of all points that are inside of the predicted and ground-truth shapes, respectively. The volumetric IoU is the volume of two shapes’ intersection divided by the volume of their union:

We randomly sample 100k points from the bounding boxes and determine if the points lie inside or outside Spred\mathcal{S}_{pred} and SGT\mathcal{S}_{GT}, respectively. The mIoU is the mean volumetric intersection over union across test set.

I-B Dynamics Metric

I-C Correspondence Metric

FMR (Feature-match Recall): we follow Deng et al.[deng2018ppfnet] for feature-match recall calculation, which measures the percentage of fragment pairs that is recovered with high confidence. Let ξgt\xi_{gt} be the ground truth mapping and ξpred\xi_{pred} be the predicted mapping. Mathematically, FMR is calculated as:

And RR is averaged across test set. τ1=0.1m\tau_{1}=0.1m inlier distance threshold and τ2=0.05\tau_{2}=0.05 or 5%5\% is the inlier recall threshold, following prior work [deng2018ppfnet, choy2019fully].

acc. (accuracy): is defined as the percentage of points that are correctly matched, where being correctly matched means that the predicted matched point is within 0.05m0.05m from the ground-truth matched point:

I-D Ranking Metric

Kendall’s τ\tau: we use Kendall’s τ\tau [kendall1938new] for ranking evaluation, which has been examined in a computer vision context before [dwibedi2019temporal]. Kendall’s τ\tau is a measure of the correspondence between two rankings, with values close to 1 indicating strong agreement, and values close to -1 indicating strong disagreement. In our scenarios, we have a predicted ranking of sampled action sequence rpredr_{pred} and a ground-truth ranking of the sampled action sequence rgtr_{gt}. For action sequence i,ji,j, the quadruplet of rank indices (rpred(i),rpred(j),rgt(i),rgt(j))(r_{pred}(i),r_{pred}(j),r_{gt}(i),r_{gt}(j)) is said to be concordant if rpred(i)<rpred(j)r_{pred}(i)<r_{pred}(j) and rgt(i)<rgt(j)r_{gt}(i)<r_{gt}(j) or rpred(i)>rpred(j)r_{pred}(i)>r_{pred}(j) and rgt(i)<rgt(j)r_{gt}(i)<r_{gt}(j). Otherwise it is said to be discordant. And Kendall’s τ\tau is defined as:

where PP is the number of concordant pairs, QQ the number of discordant pairs.

I-E Planning Execution Metric

For us, S1S_{1} represents the object shape specified in the target configuration, S2S_{2} represents the shape of the object after actions are executed.

F-Score: we follow Tatarchenko et al.[tatarchenko2019single] for F-score calculation. Recall is defined as the number of points in the ground-truth shape that lie within a certain distance to the source shape. Precision counts the percentage of points on the source shape that lie within a certain distance to the ground truth. The F-Score is then defined as the harmonic mean between precision and recall:

For us, the source shape is the final shape of the object after the action sequence is executed.

II Action-Conditional Visual Dynamics Visualizations

We show here more example visualizations of our model comparing to the baselines Pixel-Flow and Voxel-Flow. As shown in Fig. 1 and Fig. 2, our method consistently outperforms action-condition visual dynamics model baselines.

III Roll-outs Visualizations

We show here more example visualizations of ACID’s planning results. We visualize the roll-outs of the selected action sequences with lowest cost under different start and target configurations. As shown in Fig. 3, our method can estimate the dynamics of the object over multiple action commands and select reasonable action sequence.