ScrewNet: Category-Independent Articulation Model Estimation From Depth Images Using Screw Theory

Ajinkya Jain, Rudolf Lioutikov, Caleb Chuck, Scott Niekum

I Introduction

Human environments are populated with objects that contain functional parts, such as refrigerators, drawers, and staplers. These objects are known as articulated objects and consist of multiple rigid bodies connected via mechanical joints such as hinge joints or slider joints. Service robots will need to interact with these objects frequently. For manipulating such objects safely, a robot must reason about the articulation properties of the object. Safe manipulation policies for these interactions can be obtained directly either by using expert-defined control policies or by learning them through interactions with the objects . However, this approach may fail to provide good manipulation policies for all articulated objects that the robot might interact with, due to the vast diversity of articulated objects in human environments and the limited availability of interaction time. An alternative is to estimate the articulation models through observations, and then use a planning or model-based RL method to manipulate them effectively.

Existing methods for estimating articulation models of objects from visual data either use fiducial markers to track the relative movement between the object parts or require textured objects so that feature tracking techniques can be used to observe this motion . These requirements severely restrict the class of objects on which these methods can be used. Alternatively deep networks can extract relevant features from raw images automatically for model estimation . However, these methods assume prior knowledge of the articulation model category (revolute or prismatic) to estimate the category-specific model parameters, which may not be readily available for novel objects encountered by robots in human environments. Addressing this limitation, we propose a novel approach, ScrewNet, which uses screw theory to perform articulation model estimation directly from depth images without requiring prior knowledge of the articulation model category. ScrewNet unifies the representation of different articulation categories by leveraging the fact that the common articulation model categories (namely revolute, prismatic, and rigid) can be seen as specific instantiations of a general constrained relative motion between two objects about a fixed screw axis. This unified representation enables ScrewNet to estimate the object articulation models independent of the model category.

ScrewNet garners numerous benefits over existing approaches. First, it can estimate articulation models directly from raw depth images without requiring a priori knowledge of the articulation model category. Second, due to the screw theory priors, a single network suffices for estimating models for all common articulation model categories unlike prior methods . Third, ScrewNet can also estimate an additional articulation model category, the helical model (motion of a screw), without making any changes in the network architecture or the training procedure.

We conduct a series of experiments on two benchmarking datasets: a simulated articulated objects dataset , and the PartNet-Mobility dataset , and three real-world objects: a microwave, a drawer, and a toaster oven. We compare ScrewNet with a current state-of-the-art method and three ablated versions of ScrewNet and show that it outperforms all baselines with a significant margin.

II Related Work

Articulation model estimation from visual observations: proposed a probabilistic framework to learn the articulation relationships between different parts of an articulated object from the time-series observations of 6D poses of object parts . The framework was extended to estimate the articulation model for textured objects directly from raw RGB images by extracting SURF features from the images and tracking them robustly. Other approaches have explored modeling articulated objects that exhibit configuration-dependent changes in the articulation model, rather than having a single model throughout their motion . Recently, the problem of articulation model parameter estimation was posed as a regression task given a known articulation model category . The mixture density network-based approach predicts model parameters using a single depth image. However, in a realistic setting, an object’s articulation model category might not be available a priori to the robot.

Interactive perception (IP): IP approaches leverage the robot’s interaction with the objects for generating a rich perceptual signal for robust articulation model estimation . first studied IP to learn articulated motion models for planar objects , and later extended it to learn 3D kinematics of articulated objects . In more recent works, and have further extended the approach and used hierarchical recursive Bayesian filters to develop online algorithms from articulation model estimation from RGB images. However, current IP approaches still require textured objects for estimating the object articulation model from raw images, whereas, ScrewNet imposes no such requirement on the objects.

Articulated object pose estimation: For known articulated objects, the problem of articulation model parameter estimation can also be treated as an articulated object pose estimation problem. Different approaches leveraging object CAD model information and the knowledge of articulation model category have been proposed to estimate the 6D pose of the articulated object in the scene. These approaches can be combined with an object detection method, such as YOLOv4 , to develop a pipeline for estimating the articulation model parameters for objects from raw images. On the other hand, ScrewNet can directly estimate the articulation model for an object from depth images without requiring any prior knowledge about it.

Other approaches: and have proposed methods to learn articulation models as geometric constraints encountered in a manipulation task from non-expert human demonstrations. Leveraging natural language descriptions during demonstrations, Daniele et al. have proposed a multimodal learning framework that incorporates both vision and natural language information for articulation model estimation. However, these approaches use fiducial markers to track the movement of the object, unlike ScrewNet, that works on raw images.

III Background

Screw displacements: Chasles’ theorem states that “Any displacement of a body in space can be accomplished by means of a rotation of the body about a unique line in space accompanied by a translation of the body parallel to that line” . This line is called the screw axis of displacement, S\mathsf{S} . We use Plücker coordinates to represent this line. The Plücker coordinates of the line l=p+xll=\mathbf{p}+x\mathbf{l} are defined as (l,m)(\mathbf{l},\mathbf{m}), with moment vector m=p×l\mathbf{m}=\mathbf{p}\times\mathbf{l} . The constraints ∥l∥=1\left\lVert\mathbf{l}\right\rVert=1 and ⟨l,m⟩=0\langle\mathbf{l},\mathbf{m}\rangle=0 ensure that the degrees of freedom of the line in space are restricted to four. The rigid body displacement in SE(3)SE(3) is defined as σ=(l,m,θ,d)\sigma=(\mathbf{l},\mathbf{m},\theta,d). The linear displacement dd and the rotation θ\theta are connected through the pitch hh of the screw axis, d=hθd=h\theta. The distance between l1:=(l1,m1)l_{1}:=(\mathbf{l}_{1},\mathbf{m}_{1}) and l2:=(l2,m2)l_{2}:=(\mathbf{l}_{2},\mathbf{m}_{2}) is defined as:

where [t]×[\mathbf{t}]_{\times} denotes the skew-symmetric matrix corresponding to the translation vector t\mathbf{t}, and (Al,Am)(^{A}\mathbf{l},^{A}\mathbf{m}) and (Bl,Bm)(^{B}\mathbf{l},^{B}\mathbf{m}) represents the line ll in frames FA\mathcal{F}_{A} and FB\mathcal{F}_{B}, respectively .

IV Approach

Given a sequence of nn depth images I1:n\mathcal{I}_{1:n} of motion between two parts of an articulated object, we wish to estimate the articulation model M\mathcal{M} and its parameters ϕ\phi governing the motion between the two parts without knowing the articulation model category a priori. Additionally, we wish to estimate the configurations q1:nq_{1:n} that uniquely identify different relative spatial displacements between the two parts in the given sequence of images I1:n\mathcal{I}_{1:n} under model M\mathcal{M} with parameters ϕ\phi. We consider articulation models with at most one degree-of-freedom (DoF), i.e. M∈{Mrigid,Mrevolute,Mprismatic,Mhelical}\mathcal{M}\in\{\mathcal{M}_{\text{rigid}},\mathcal{M}_{\text{revolute}},\mathcal{M}_{\text{prismatic}},\mathcal{M}_{\text{helical}}\}. Model parameters ϕ\phi are defined as the parameters of the screw axis of motion, i.e. S=(l,m)\mathsf{S}=(\mathbf{l},\mathbf{m}), where both l\mathbf{l} and m\mathbf{m} are three-dimensional real vectors. Each configuration qiq_{i} corresponds to a tuple of two scalars, qi=(θi,di)q_{i}=(\theta_{i},d_{i}), defining a rotation around and a displacement along the screw axis S\mathsf{S}. We assume that the relative motion between the two object parts is governed only by a single articulation model.

IV-B ScrewNet

We propose ScrewNet, a novel approach that given a sequence of segmented depth images I1:n\mathcal{I}_{1:n} of the relative motion between two rigid objects estimates the articulation model M\mathcal{M} between the objects, its parameters ϕ\phi, and the corresponding configurations q1:nq_{1:n}. In contrast to state-of-the-art approaches, ScrewNet does not require a priori knowledge of the articulation model category for the objects to estimate their models. ScrewNet achieves category independent articulation model estimation by representing different articulation models through a unified representation based on the screw theory . ScrewNet represents the 1-DoF articulation relationships between rigid objects (rigid, revolute, prismatic, and helical) as a sequence of screw displacements along a common screw axis. A rigid model is now defined as a sequence of identity transformations, i.e., θ1:n=0∧d1:n=0\theta_{1:n}=0\wedge d_{1:n}=0, a revolute model as a sequence of pure rotations around a common axis, i.e., θ1:n≠0∧d1:n=0\theta_{1:n}\neq 0\wedge d_{1:n}=0, a prismatic model as a sequence of pure displacements along the same axis, i.e., θ1:n=0∧d1:n≠0\theta_{1:n}=0\wedge d_{1:n}\neq 0, and, a helical model as a sequence of correlated rotations and displacements along a shared axis, i.e., θ1:n≠0∧d1:n≠0\theta_{1:n}\neq 0\wedge d_{1:n}\neq 0).

Under this unified representation, all 1-DoF articulation models can be represented using the same number of parameters, i.e., 6 parameters for the common screw axis S\mathsf{S} and 2n=∣{(θi,di)∀i∈{1...n}}∣2n=|\{(\theta_{i},d_{i})\forall i\in\{1...n\}\}| parameters for configurations, which enables ScrewNet to perform category independent articulation model estimation. ScrewNet not only estimates the articulation model parameters without requiring the model category M\mathcal{M}, but is also capable of estimating the category itself. This ability can potentially reduce the number of control parameters required for manipulating the object . A unified representation also allows ScrewNet to use a single network to estimate the articulation motion models across categories, unlike prior approaches that required separate networks, one for each articulation model category . Having a single network grants ScrewNet two major benefits: first, it needs to train fewer total parameters, and second, it allows for a greater sharing of training data across articulation categories, resulting in a significant increase in the number of training examples that the network can use. Additionally, in theory, ScrewNet can also estimate an additional articulation model category, the helical model, which was not addressed in earlier work .

Architecture: ScrewNet sequentially connects a ResNet-18 CNN , an LSTM with one hidden layer, and a 3-Layer MLP. ResNet-18 extracts features from the depth images that are fed into the LSTM which encodes the sequential information from the extracted features into a latent representation. Using this representation, the MLP then predicts a sequence of screw displacements having a common screw axis. The network is trained end-to-end, with ReLU activations for the fully-connected layers. Fig. 2 shows the network architecture. The model category M\mathcal{M} is deduced from the predicted screw displacements using a decision-tree based on the displacements properties of each model class.

Loss function: Screw displacements are composed of two major components: the screw axis S\mathsf{S}, and the corresponding configurations qiq_{i} about it. Hence, we pose ScrewNet training as a multi-objective optimization problem with loss

where λi\lambda_{i} weights the respective component. LSori\mathcal{L}_{\mathsf{S}_{\text{ori}}} penalizes the screw axis orientation mismatch as the angular difference between the target and the predicted orientations. LSdist\mathcal{L}_{\mathsf{S}_{\text{dist}}} penalizes the spatial distance between the target and predicted screw axes as defined in Eqn. 1. LScons\mathcal{L}_{\mathsf{S}_{\text{cons}}} enforces the constraints ⟨l,m⟩=0\langle\mathbf{l},\mathbf{m}\rangle=0 and ∥l∥=1\left\lVert\mathbf{l}\right\rVert=1. Lq:=α1Lθ+α2Ld\mathcal{L}_{\text{q}}:=\alpha_{1}\mathcal{L}_{\theta}+\alpha_{2}\mathcal{L}_{d} penalizes errors in the configurations, where Lθ\mathcal{L}_{\theta} and Ld\mathcal{L}_{d} represent the rotational and translational error respectively:

with R(θ;l)R(\theta;\mathbf{l}) denoting the rotation matrix corresponding to a rotation of angle θ\theta about the axis l\mathbf{l}. We choose this particular form of the loss function for Lq\mathcal{L}_{\text{q}}, rather than a standard loss function such as an L2L2 loss, as it ensures that the network predictions are grounded in their physical meaning. By imposing a loss based on the orthonormal property of the 3D3D rotations, the proposed loss function ensures that the learned angle-axis pair (lpred,θpred)(\mathbf{l}_{\text{pred}},\theta_{\text{pred}}) corresponds to a rotation R(θtar;ltar)∈SO(3)R(\theta_{\text{tar}};\mathbf{l}_{\text{tar}})\in SO(3). Similarly, the loss function Ld\mathcal{L}_{d} calculates the difference between the two displacements along two different axes ltar\mathbf{l}_{\text{tar}} and lpred\mathbf{l}_{\text{pred}}, rather than calculating the difference between the two configurations, dtard_{\text{tar}} and dpredd_{\text{pred}}, which assumes that they represent displacements along the same axis. Hence, this choice of loss function ensures that the network predictions conform to the definition of a screw displacement. We empirically choose weights to be λ1=1,λ2=2,λ3=1,λ4=1,α1=1\lambda_{1}=1,\lambda_{2}=2,\lambda_{3}=1,\lambda_{4}=1,\alpha_{1}=1, and α2=1\alpha_{2}=1.

Training data generation: The training data consists of sequences of depth images of objects moving relative to each other and the corresponding screw displacements. The objects and depth images are rendered in Mujoco . We apply random frame skipping and pixel dropping to simulate noise encountered in real world sensor data. We use the cabinet, drawer, microwave, the toaster-oven object classes from the simulated articulated object dataset. The cabinet, microwave, and toaster object classes contain a revolute joint each, while the drawer class contains a prismatic joint. We consider both left-opening cabinets and right-opening cabinets. From the PartNet-Mobility dataset , we consider the dishwasher, oven, and microwave object classes for the revolute articulation model category, and the storage furniture object class consisting of either a single column of drawers or multiple columns of drawers, for the prismatic articulation model category.

V Experiments

We evaluated ScrewNet’s performance in estimating the articulation models for objects by conducting three sets of experiments on two benchmarking datasets: the simulated articulated objects dataset provided by Abbatematteo et al. and the recently proposed PartNet-Mobility dataset . The first set of experiments evaluated ScrewNet’s performance in estimating the articulation models for unseen object instances that belong to the object classes used for training the network. Next, we tested ScrewNet’s performance in estimating the model parameters for novel articulated objects that belong to the same articulation model category as seen during training. In the third set of experiments, we trained a single ScrewNet on object instances belonging to different object classes and articulation model categories and evaluated its performance in cross-category articulation model estimation. We compared ScrewNet with a state-of-the-art articulation model estimation method proposed by Abbatematteo et al. . Lastly, to evaluate how effectively ScrewNet transfers from simulation to real-world setting, we trained ScrewNet solely using simulated images and tested it to estimate articulation models for three real-world objects.

In all the experiments, we assumed that the input depth images are semantically segmented and contain non-zero pixels corresponding only to the two objects between which we wish to estimate the articulation model. Given this input, ScrewNet estimates the articulation model parameters for the pair of objects in an object-centric coordinate frame defined at the center of the bounding box of the object. Note while the approach proposed by Abbatematteo et al. can be used to estimate the articulation model parameters directly in the camera frame, for a fair comparison to our approach, we modified the baseline to predict the model parameters in the object-centric reference frame as well.

In the first set of experiments, we investigated whether our proposed approach can generalize to unseen object instances belonging to the object classes seen during the training. For this set of experiments, we trained a separate ScrewNet and a baseline network for each of the object classes and tested how ScrewNet fares in comparison to the baseline under similar experimental conditions. We generated 10,000 training examples for each object class in both datasets and performed evaluations on 1,000 withheld object instances. From Fig. 4, it is evident that ScrewNet outperformed the baseline in estimating the joint axis position and the observed joint configurations by a significant margin for the first dataset. However, for the joint axis orientation estimation, the baseline method reported lower errors than the ScrewNet. Similar trends in the performance of the two methods were observed on the PartNet-Mobility dataset (see Fig. 4). ScrewNet significantly outperformed the baseline method in estimating the joint axis displacement and observed joint configurations, while the baseline reported lower errors than ScrewNet in estimating the joint axis orientations. However, for both the datasets, the errors reported by ScrewNet in screw axis orientation estimation were reasonably low (<<$$), and the model parameters predicted by ScrewNet may be used directly for manipulating the object. These experiments demonstrate that under similar experimental conditions, ScrewNet can estimate the joint axis positions and joint configurations for objects better than the baseline method, while reporting reasonably low but higher errors in joint axis orientations.

V-B Same articulation model category

Next, we investigated if our proposed approach can generalize to unseen object classes belonging to the same articulation model category. We conducted this set of experiments only on the PartNet-Mobility dataset as the simulated articulated objects dataset does not contain enough variety of object classes belonging to the same articulation model category (only 3 for revolute and 1 for prismatic). For the revolute category, we trained ScrewNet and the baseline on the object instances generated from the oven and the microwave object classes and tested it on the objects from the dishwasher class. For the prismatic category, we trained them on the objects from the storage furniture class containing multiples columns of drawers and tested it on the storage furniture objects containing a single column of drawers. We trained a single instance of ScrewNet and the baseline for each articulation model category and used them to predict the articulation model parameters for the test object classes. We used the same training datasets as used in the previous set of experiments. Results are reported in Fig. 5. It is evident from Fig. 6 that ScrewNet was able to generalize to novel object classes belonging to the same articulation model category, while the baseline failed to do so. Both methods reported low errors in the joint axis orientation and the observed configurations. However, for the joint axis position, the baseline method reported mean errors of an order of magnitude higher than the ScrewNet for both the articulation model categories.

V-C Across articulation model category

Next, we studied whether ScrewNet can estimate articulation model parameters for unseen objects across the articulation model category. For these experiments, we trained a single ScrewNet on a mixed dataset consisting of object instances belonging to all object model classes. To test whether sharing training data across articulation categories can help in reducing the number of examples required for training, we used only half of the dataset available for each object class (50005000 examples each) while preparing the mixed dataset. We compared its performance with a baseline network that is trained specifically on the particular object class. We also conducted ablation studies to test the effectiveness of the various components of the proposed method.

Fig. 6 summarizes the results for the first dataset. Even though we used a single network to estimate the articulation model for objects belonging to different articulation model categories, ScrewNet performed at par or better than the baseline method for all the object model categories. ScrewNet outperformed the baseline while estimating the observed joint configurations for all object classes, even though the baseline was trained separately for each object class. For joint axis position estimation, ScrewNet reported significantly lower errors than the baseline for the cabinet and the drawer classes, and comparable errors for the microwave and the toaster classes. In estimating the joint axis orientations, both methods reported comparable errors for the cabinet, drawer, and the toaster classes. However, for the cabinet object class, ScrewNet reported a higher error than the baseline method, which may stem from the fact that the cabinet object class includes both left-opening and right-opening configurations that have a difference in their axis orientations. On the PartNet-Mobility dataset (see Fig. 6), the performances of the methods was similar, with ScrewNet outperforming the baseline method with a significant margin in estimating the joint axis positions and the observed joint configurations while reporting higher errors than the baseline in estimating the joint axis orientations. The results show ScrewNet leverages the unified representation and performs cross-category articulation model estimation with better on average performance than the current state-of-the-art method while using only half the training examples.

In comparison to its ablated versions, ScrewNet outperformed the L2-error and the two-images versions by a significant margin for both datasets and performed comparably to the NoLSTM version. For the first dataset, the NoLSTM version reported lower errors than ScrewNet in estimating the joint axis orientations, their positions, and the observed joint configurations for the microwave, cabinet, and toaster classes. However, the NoLSTM version failed to generalize across articulation model categories and reported higher errors than the ScrewNet for the drawer class, and sometimes even predicted NaNs. On the second dataset, ScrewNet reported much lower errors than the NoLSTM ablated version for all object model categories. These results demonstrate that for reliably estimating articulation model parameters across categories, both the sequential information available in the input and a loss function that grounds predictions in their physical meaning are crucial.

V-D Real world images

Finally, we evaluated how effectively ScrewNet transfers from simulation to a real-world setting. ScrewNet was trained solely on the combined simulated articulated object dataset. Afterwards, we used the model to infer the joint axis of a microwave, a drawer, and a toaster oven. Figure 7 qualitatively demonstrates ScrewNet’s performance for three different poses of the microwave. Despite only ever having seen simulated data, ScrewNet achieved a mean error of ∼\sim$inaxisorientationandin axis orientation and\sim 1.5\text{cm}$ in axis position on real-world sensor input. These results demonstrate that ScrewNet achieves reasonable estimates of the articulation model parameters for real-world objects when it is trained solely using simulated data. In order to obtain better performances a retraining on real world data would be required.

VI Conclusion

Articulated objects are common in human environments and service robots will be interacting with them frequently while assisting humans. For manipulating such objects, a robot will need to learn the articulation properties of such objects through raw sensory data such as RGB-D images. Current methods for estimating the articulation model of objects from visual observations either require textured objects or need to know the articulation model category a priori for estimating the articulation model parameters from the depth images. We propose ScrewNet that uses screw theory to unify the representation of different articulation models and performs category-independent articulation model estimation from depth images. We evaluate the performance of ScrewNet on two benchmarking datasets and compare it with a state-of-the-art method. Results show that ScrewNet can estimate articulation models and their parameters for objects across object classes and articulation model categories successfully with better on average performance than the baselines while using half the training data and without requiring to know the model category.For further details and results, refer: https://arxiv.org/abs/2008.10518

While ScrewNet successfully performs cross-category articulation model estimation, at present, it can only predict 1-DOF articulation models. For multi-DoF objects, an image segmentation step is required to mask out all non-relevant object parts. This procedure can be repeated iteratively on all object part pairs to estimate relative models between object parts. Future work will directly estimate multi-DoF objects by learning a segmentation network along with the ScrewNet. Further work work will predict the articulation model parameters directly in the robot’s camera frame rather than in an object-centric frame. Predictions in the camera frame can help the robot to learn articulation models in an active learning fashion.

VII Acknowledgement

This work has taken place in the Personal Autonomous Robotics Lab (PeARL) at The University of Texas at Austin. PeARL research is supported in part by the NSF (IIS-1724157, IIS-1638107, IIS-1749204, IIS-1925082), ONR (N00014-18-2243), AFOSR (FA9550-20-1-0077), and ARO (78372-CS). This research was also sponsored by the Army Research Office under Cooperative Agreement Number W911NF-19-2-0333. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the Army Research Office or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation herein. The authors would also like to thank Wonjoon Goo, Yuchen Cui, and Akanksha Saran for their insightful discussions and Tushti Shah for her help with graphics.

References

Appendix

Objects used in the experiments from each of the dataset are shown in the Figures 8 and 9. We sampled a new object geometry and a joint location for each training example in the simulated articulated object dataset, as proposed by . For the PartNet-Mobility dataset, we considered 1111 microwave (88 train, 33 test), 3636 dishwasher (2727 train, 99 test), 99 oven (66 train, 33 test), 2626 single column drawer (2020 train, 66 test), and 1414 multi-column drawer (1010 train, 44 test) object models. For both datasets, we sampled object positions and orientations uniformly in the view frustum of the camera up to a maximum depth dependent upon the object size.

VII-A2 Experiment 1: Same object class

Numerical error values for the first set of experiments for the simulated articulated objects dataset are presented in the Table I. It is evident from the Table I that the baseline succeeded in achieving nearly zero prediction error () in joint axis orientation estimation for all object classes. ScrewNet also performed well and reported low prediction errors (<<$)forthedrawer,microwave,andtoasterobjectclasses.Forthecabinetobjectclass,whileScrewNetreportedahighermeanerror() for the drawer, microwave, and toaster object classes. For the cabinet object class, while ScrewNet reported a higher mean error (\sim$$$), it is relatively small compared to the difference in axis orientations, , between the two possible configurations of the cabinet (left-opening or right-opening). For the other two model parameters, namely the joint axis position and the observed configurations, ScrewNet significantly outperformed the baseline method.

Numerical error values for the first set of experiments for the PartNet-Mobility dataset are reported in the Table II. Similar trends followed in the performance of the two approaches. The baseline achieved very high accuracy in predicting the joint axis orientation, whereas ScrewNet reported reasonably low but slightly higher errors (<<$$). For the joint axis position and the observed configurations, ScrewNet outperformed the baseline method on this dataset as well.

VII-A3 Experiment 2: Same articulation model category

Numerical results for the second set of experiments are reported in the Table III. It is evident from the Table III that the ScrewNet was able to generalize to novel object classes belonging to the same articulation model category, while the baseline method failed to do so. While both approaches reported comparable errors in estimating the joint axis orientations and the observed configurations, the baseline reported errors of an order of magnitude higher than ScrewNet in the joint axis position estimation.

VII-A4 Ablation studies

We consider three ablated versions of ScrewNet. First, to test the effectiveness of the proposed loss function, we consider an ablated version of ScrewNet which is trained using a raw L2-loss between the labels and the network predictions (named as L2-Error version while reporting results). As the second ablation study, we test whether using an LSTM layer in the network helps with the performance or not (named as NoLSTM version while reporting results). We replace the LSTM layer of the ScrewNet with a fully connected layer such that the two networks, ScrewNet and its ablated version, have a comparable number of parameters. Lastly, to check if a sequence of images is helpful in the model estimation or not, we consider an ablated version of ScrewNet that estimates the articulation model using just a pair of images (named as 2_imgs version while reporting results). Note that ScrewNet and all its ablated versions use a single network each. Numerical results for the simulated articulated objects dataset are presented in the Table IV, and for the PartNet-Mobility dataset are shown in the Table V.