LoopReg: Self-supervised Learning of Implicit Surface Correspondences, Pose and Shape for 3D Human Mesh Registration
Bharat Lal Bhatnagar, Cristian Sminchisescu, Christian Theobalt, Gerard Pons-Moll
Legend for Notations
We realise that the paper is a bit intensive in terms of notations. For improved readability, we present the definition of key symbols used in the main paper in Table 1.
Results: Correspondence prediction
Establishing correspondences across 3D shapes is a challenging problem in computer graphics and vision community. Though our work does not directly predict correspondences between two shapes we can still register the two shapes with a common template. This allows us to establish correspondences between the shapes. We compare the performance of our approach on the correspondence prediction task on FAUST . The FAUST test set contains 200 scans of undressed people in challenging poses and the scans themselves are noisy. The evaluation metric is based on the geodesic distance between the predicted correspondence and the GT correspondence. This metric heavily penalises the errors made due to self contacts on the body. In practice it can be seen that the overall distribution of errors is dominated by the contact errors making the evaluation less than ideal. Nonetheless we report the results as per the protocol in Table 2. For competing approaches we take the numbers from the corresponding papers. It can be clearly seen that our model trained primarily with self-supervision performs better than the competing approaches. Also note that none of these other approaches generalises to dressed humans where as in Fig. 3 we show that our method can predict correspondences for both dressed and undressed scans.
More qualitative results
We show additional qualitative results for undressed scan registration in Fig. 1 and for dressed scans in Fig. 2. It can be seen that our approach can produce high quality registrations for both undressed and dressed scans in complex poses. Our approach predicts continuous correspondences from the scan to the canonical human template. We visualise these correspondences in Fig. 3 and show that we can accurately predict correspondences for both undressed and dressed scans.
Limitations and Future Work
Our formulation allows us to jointly differentiate through the correspondences and the instance specific human model parameters. This allows us to create a self-supervised loop for registration. But in practice we find that in order for this loop to not get stuck in a local minima, it is important that correspondences are initialized well. So even though our formulation does not require labeled data, in practice we find that a supervised warm-start with a small amount of data is important for subsequent self-supervised training.
As shown in our results, our approach performs significantly better than other competing approaches both qualitatively and quantitatively. We still find that our registration is not as high quality as when they have access to precomputed 3D joints, facial landmarks and manual intervention (Note that our approach does not require this information). This is not necessarily a limitation as these additional cues can be integrated with our approach as well. We leave this as a potential future work.