Self-supervised Single-view 3D Reconstruction via Semantic Consistency
Xueting Li, Sifei Liu, Kihwan Kim, Shalini De Mello, Varun Jampani, Ming-Hsuan Yang, Jan Kautz
Introduction
Recovering both 3D shape and texture, and camera pose from 2D images is a highly ill-posed problem due to its inherent ambiguity. Existing methods resolve this task by utilizing various forms of supervision such as ground truth 3D shapes , 2D semantic keypoints , shading , category-level 3D templates or multiple views of each object instance . These types of supervision signals require tedious human effort, and hence make it challenging to generalize to many object categories that lack such annotations. On the other hand, learning to reconstruct by not using any 3D shapes, templates, or keypoint annotations, i.e., with only a collection of single-view images and silhouettes of object instances, remains challenging. This is because the reconstruction model learned without the aforementioned supervisory signals leads to erroneous 3D reconstructions. A typical failure case is caused by the “camera-shape ambiguity”, wherein, incorrectly predicted camera pose and shape result in a rendering and object boundary that closely match the input 2D image and its silhouette, as shown in Figure 2 (c) and (d).
Interestingly, humans, even infants who have never been taught about objects in a category, tend to mentally reconstruct objects in that category by perceiving them as a combination of several basic parts, e.g., a bird has two legs, two wings, and one head, etc., and use the parts to associate all the divergent instances of the category. By observing object parts, humans can also roughly infer the relative camera pose and 3D shape of any specific instance. In computer vision, a similar intuition is formulated by the deformable parts model, where objects are represented as a set of parts arranged in a deformable configuration .
Inspired by this intuition, we learn a single-view reconstruction model from a collection of images and silhouettes. We utilize the semantic parts in both the 2D and 3D space, along with their consistency to correctly estimate shape and camera pose. Specifically, we first leverage self-supervised co-part segmentation (SCOPS ) to decompose 2D images into a collection of semantic parts (Figure 1(b)). By exploiting the property of semantic part invariance, which states that the semantic part label of a point on the mesh surface does not change even when the mesh shape is deformed, we associate the semantic parts of different object instances with each other and build a category-level canonical semantic UV map (Figure 1(c)). The semantic part label of each point on the reconstructed mesh surface (Figure 1(d)) is then defined by this canonical semantic UV map. Finally, we resolve the aforementioned “camera-shape ambiguity” and learn the self-supervised reconstruction model by encouraging the consistency of semantic part labels in both the 2D and 3D space (Figure 1, orange arrow). Furthermore, we train our model by iteratively learning (a) instance-level reconstruction and (b) a category-level template mesh from scratch. Thus, our model also does not require a pre-defined 3D template mesh or any other shape prior. Our main contribution is a 3D reconstruction model that is able to:
Conduct single-view mesh reconstruction without any of the following forms of supervision: category-level 3D template prior, annotated keypoints, camera pose or multi-view images. In other words, the model can be generalized to other categories which do not have well-defined keypoints, e.g., penguin.
Leverage the semantic part invariance property of object instances of a category as a deformable parts model.
Learn a category-level 3D shape template from scratch via iterative learning.
Perform comparably to the state-of-the-art supervised methods trained with either pre-defined templates or annotated keypoints, while also improving the self-supervised semantic co-part segmentation model (SCOPS ).
Related Work
Various representations have been explored for 3D processing tasks, including point clouds , implicit surfaces , triangular meshes and voxel grids . Among these, while both voxels and point clouds are more friendly to deep learning architectures (e.g., VON , PointNet , etc), they suffer either from issues of memory inefficiency or are not amenable to differentiable rendering. Hence, in this work, we adopt triangular meshes for 3D reconstruction.
Single-view 3D reconstruction aims to reconstruct a 3D shape given a single input image. One line of works have explored this ill-posed task with varying degree of supervision. Several methods utilize image and ground truth 3D mesh pairs as supervision. This either requires significant manual annotation effort or is restricted to synthetic data . More recently, a few works avoid 3D supervision by taking advantage of differentiable renderers and the “analysis-by-synthesis” approach, with either multiple views, or known ground truth camera poses.
To further relax constraints on supervision, Kanazawa et al. explored 3D reconstruction from a collection of images of different instances. However, their method still requires annotated 2D keypoints to infer camera pose correctly. It is also the first work to propose a learnable category-level 3D template shape, which, however, needs to be initialized from a keypoint-dependent 3D convex hull. Similar problem settings have also been explored in other methods , but with object categories restricted to rigid or structured objects, such as cars or faces. Different from all these works, we target both rigid and non-rigid objects (e.g., birds, horses, penguins, motorbikes and cars shown in Figure 1 (e)-(g)) and propose a method that jointly estimates a 3D mesh, texture, and camera pose from a single-view image, using only a collection of images with silhouettes as supervisions. In other words, we do not require 3D template priors, annotated keypoints, or multi-view images.
Our work is also related to self-supervised cross-instance correspondence learning, via landmarks , part segments , or canonical surface mapping . We utilize self-supervised co-parts segmentation to enforce semantic consistency, which was originally proposed purely for 2D images. The work of learns a mapping function that maps pixels in 2D images to a predefined category-level template in a self-supervised manner. However, it dose not use the learned correspondence for 3D reconstruction. We show that our work, despite having a focus on 3D reconstruction, outperforms at learning 2D to 3D correspondences as well.
Approach
One of the key elements for the CMR method to perform well is to exploit mannually annotated semantic keypoints for (i) precisely pre-computing the ground truth camera pose for each instance, and (ii) estimating a category-level 3D template prior. However, annotating keypoints is tedious, not well-defined for most object categories in the world and impossible to generalize to new categories. Thus, we propose a method within a more scalable, but challenging self-supervised setting without using manually annotated keypoints to estimate camera pose or a template prior.
Not surprisingly, simply taking out the keypoints supervision, as well as all the related information (i.e., the camera pose and the template prior) from the CMR network makes it unable to predict camera pose and shape correctly, as shown in Figure 2(c) and (d). This is due to the inherent ambiguity of hallucinating 3D meshes from only single-view 2D observations, where the model trivially picks a combination of camera pose and shape that yields the rendering that matches the given image and silhouette. Consider an extreme case, where the model predicts the front view for all instances, but is still able to match the image and silhouette observations by deforming each instance mesh accordingly.
In this section, we show the key to solving the “camera-shape ambiguity” is to make use of the semantic parts of object instances in both 3D and 2D. Specifically, we exploit the fact that (i) in the 2D space, the self-supervised co-part segmentation provides correct part segments for the majority of the object instances, even for those with large shape variations (see Figure 1(b)); and (ii) in the 3D space, semantic parts are invariant to mesh deformations, i.e., the semantic part label of a specific point on the mesh surface is consistent across all reconstructed instances of a category. We demonstrate that this semantic part invariance allows us to build a category-level semantic UV map, namely the canonical semantic UV map, shared by all instances, which in turn allows us to assign semantic part labels to each point on the mesh. By enforcing consistency between the canonical semantic map and an instance’s part segmentation in the 2D space, the camera-shape confusion can be largely resolved.
SCOPS is a self-supervised method that learns semantic part segmentation from a collection of images of an object category (see Figure 1(b)). The model leverages concentration and equivalence loss functions, as well as part basis discovery to output a probabilistic map w.r.t. the discovered parts, which are semantically consistent across different object instances. In Figure 10 (second row), we demonstrate several examples of semantic part segments predicted by the SCOPS. Although SCOPS successfully discovers all semantic parts for most instances, the shape and size of each part is not consistent across different instances, e.g., the part corresponding to the head of the bird in Figure 10 (second row) is too small. We discuss that, besides generalizing SCOPS for reconstructing objects, how our model is able to improve SCOPS in return, in Section 3.4.
Ideally, all instances should result in the same semantic UV map – the canonical semantic UV map for a category, regardless of shape differences of instances. This is because: (i) the semantic part invariance states that the semantic part labels assigned to each point on the mesh surface are consistent across different instances; and (ii) the mapping function that maps pixels from the UV space to the mesh surface is pre-defined and independent of deformations in the 3D space, such as face location or area changes. Thus, the semantic part labels of pixels in the UV map should also be consistent across different instances.
As mentioned above, because the self-supervisedly learned model only relies on images and silhouettes, which do not provide any semantic part information, the model suffers from the “camera-shape ambiguity” introduced in Section 1. Take row (i) in Figure 4 as an example. The model erroneously forms the wing tip in the reconstructed bird by deforming faces assigned as the “head part” (colored in red). This incorrect shape reconstruction, associated with an incorrect camera pose, however, can yield a rendering that matches the image and silhouette observation.
This ambiguity, although is not easy to spot by only comparing the rendered reconstruction image with the input image, however, can be identified once the semantic part label for each point on the mesh surface is available. One can tell that the reconstruction in row (i) of Figure 4 is wrong by comparing the rendering of the semantic part labels on the mesh surface and the 2D SCOPS part segmentation. Only when the camera pose and shape are both correct, will the rendering and the SCOPS segmentation be consistent, as shown in row (ii) in Figure 4. This observation inspires us to propose a probability and a vertex-based constraint that facilitate camera pose and shape learning by encouraging the consistency of semantic part labels in both 2D images and the mesh surface.
We empirically found the mean squared error (MSE) metric to be more robust than the Kullback–Leibler divergence for comparing two probability maps.
We also propose a vertex-based constraint to enhance semantic part consistency (Figure 5) by enforcing that 3D vertices assigned a part label , after being projected to the 2D domain with the predicted camera pose , align with the area assigned to that part in the input image:
where is the set of vertices on a learned category-level 3D template (see Section 3.2) with the part label , is the set of 2D pixels sampled from the part in the original input image and is the number of parts. Here we use the Chamfer distance, because the projected vertices and pixels with the same part label in the input image do not have a strictly one-to-one correspondence.
Note that, is a set of vertices on the category-level shape template as opposed to each instance reconstruction , since using results in a degenerate solution where the network only alters 3D shape to satisfy this vertex-based constraint, rather the camera pose. Instead, using drives the network towards learning the correct camera pose, in addition to shape.
2 Progressive Training
We train the framework in Figure 3 by a progressive training method based on two considerations: (a) building the canonical semantic UV map, introduced in Section 3.1, requires reliable texture flows to map the SCOPS from 2D images to the UV space. Thus the canonical semantic UV map can only be obtained after the reconstruction network is able to predict texture flow reasonably well, and (b) a canonical 3D shape template is desirable, since it speeds up the convergence of the network and also avoids degenerate solutions when applying the vertex-based constrain as introduced in Section 3.1. However, jointly learning the category-level 3D shape template and the instance-level reconstruction network leads to undesired trivial solutions. Thus, we propose an expectation-maximization (EM) style progressive training procedure below. In the E-step, we train the reconstruction network with the current template and canonical semantic UV map fixed, and in the M-step, we update the template and the canonical semantic UV map using the reconstruction network learned in the E-step.
In the E-step, we fix the canonical semantic UV map as well as the category-level template and train the reconstruction network mainly with the following objectives. (i) A negative IoU objective between the rendered and the ground truth silhouettes for shape learning. (ii) A perceptual distance objective between the rendered and the input RGB images for texture learning. (iii) The probability and vertex-based constraints introduced in Section 3.1 to resolve the “camera-shape ambiguity” under the self-supervised setting. (iv) A texture consistency constraint to facilitate accurate texture flow learning that will be introduced in Section 3.3. Other constraints is included in the appendix. Note that in the first E-step, the template is a sphere and the probability as well as vertex-based constraints are not involved.
In the M-step, we compute the canonical semantic UV map as introduced in Section 3.1 and learn a category-level template from scratch, i.e., from a sphere primitive. As far as we know, we are the first method that learns a category-level template from scratch. This is in contrast to existing methods , where the template is either a readily available instance mesh from the category or estimated from annotated keypoints . Jointly learning the shape template along with the reconstruction network does not guarantee a meaningful “mean shape” which encapsulates the most representative characteristics of objects in a category. Instead, we propose a feed-forward template learning approach: the template starts out as a sphere and is updated every training epochs by:
3 Texture Cycle Consistency Constraint
As shown in Figure 6 (Left), one issue with the learned texture flow is that the texture of 3D mesh faces with a similar color (e.g., black) can be incorrectly sampled from a single pixel location of the image. Thus we introduce a texture cycle consistency objective to regularize the predicted texture flow (i.e., 2D3D) to be consistent with the camera projection (i.e., 3D2D). As shown in Figure 6 (Right), considering the pixel marked with a yellow cross in the input image, it can be mapped to the mesh surface through the predicted texture flow along with the pre-defined mapping function introduced in Section 3. Meanwhile, its mapping on the mesh surface can be re-projected back to the 2D image by the predicted camera pose, as shown by the green cross in Figure 6 (Right). If the predicted texture flow conforms to the predicted camera pose, the yellow and green crosses would overlap, forming a cycle.
We note that while not targeting 3D mesh reconstruction directly, a similar intuition, but with a different formulation was also introduced in .
4 Better Part Segmentation via Reconstruction
By mapping the canonical UV map to the surface of each reconstructed mesh and rendering it with the predicted camera pose, we obtain psuedo “ground truth” segmentation maps as supervision for SCOPS training. We use the semantic consistency constraint in Section 3.1 as a measurement to select the reliable reconstructions with high semantic consistency (i.e., with low probability and vertex-based semantic consistency loss values) to train SCOPS with. The improved SCOPS can, in turn, provide better regularization for our mesh reconstruction network, forming an iterative and collaborative learning loop.
Experimental Results
We first introduce our experimental settings in Section 4.1, and present qualitative evaluations for the bird, horse, motorbike and car categories in Section 4.2. Quantitative evaluations and ablation studies for the contribution of each proposed module are discussed in Section 2 and Section 4.4, respectively.
We validate our method on both rigid objects, i.e., car and motorcycle images from the PASCAL3D+ dataset , and non-rigid objects, i.e., bird images from the CUB-200-2011 dataset , horse, zebra, cow images from the ImageNet dataset and penguin images from the OpenImages dataset .
2 Qualitative Results
Thanks to the self-supervised setting, our model is able to learn from a collection of images and silhouettes (e.g., horse and cow images and penguin images ), which cannot be achieved by existing methods that require extra supervisory signals.
We show the learned templates for the bird, horse, motorbike and car categories in Figure 7 and Figure 8, which capture the shape characteristics of each category, including the details such as the beak and feet of a bird, etc. We also visualize the canonical semantic UV map by showing the semantic part labels assigned to each point on the template surface. For instance, the bird meshes have four semantic parts – head (red), neck (green), belly (blue) and back (yellow) in Figure 7, which are consistent with the part segmentation predicted by the SCOPS method .
We show the results of 3D reconstruction from each single-view image in Figure 7 (b)-(d) and Figure 8 (b). Our model can reconstruct instances from the same category with highly divergent shapes, e.g., a thin bird in (b), a duck in (c) and a flying bird in (d). Our model also correctly maps the texture from each input image onto its 3D mesh, e.g., the eyes of each bird as well as fine textures on the back of the bird. Furthermore, the renderings of the reconstructed meshes under the predicted camera poses (2nd and 3rd columns in Figure 7 and Figure 8) match well with the input images in the first column, indicating that our model accurately predicts the original camera view.
In Figure 10, we visualize the results of improving SCOPS with our 3D reconstruction network as discussed in Section 3.4. Thanks to our learned canonical semantic UV map, the improved SCOPS method is able to predict the correct parts and accurately localizes them with a more precise size (head and neck parts in column 1,2,3).
3 Quantitative Evaluations
In this section, we quantitatively evaluate the reconstruction network in terms of shape, texture and camera pose prediction of non-rigid (bird ) as well as rigid (car ) objects. Our model cannot be quantitatively evaluated on other categories (e.g., horse, cows, penguins, etc.) due to a lack of ground truth keypoints, 3D meshes or camera poses in these datasets. Furthermore, since the ground truth textures and camera poses are also not available for the bird and car categories, we evaluate them through the task of keypoint transfer. Given a pair of source and target images of two different object instances from a category, we map a set of annotated keypoints from the source image to the target image by first mapping them onto the learned template and then to the target image. Each mapping can be carried out by either the learned texture flow or the camera pose, as explained below.
We first evaluate shape reconstruction on the bird category. Due to a lack of ground truth 3D shapes in the CUB-200-2011 dataset , we follow and compute the mask reprojection accuracy – the intersection over union (IoU) between rendered and ground truth silhouettes. As shown in Table 1, our model is able to achieve comparable if not better mask reprojection accuracy compared to CMR , which unlike our method is learned with additional supervision from semantic keypoints. This indicates that our model is able to predict 3D mesh reconstructions and camera poses that are well matched to the 2D observations.
Next, we evaluate shape reconstruction on the car category. Although PASCAL3D+ provides “ground truth” meshes (the most similar ones to the image in a mesh library), our reconstructed meshes are not aligned with these “ground truth” meshes since our self-suerpvised model is free to learn its own “canonical reference frame”. Thus, to quantitatively evaluate the intersection over union (IoU) between the two meshes, following CMR , we exhaustively search a set of scale, translation and rotation parameters that best align to the “ground truth” meshes. Our method achieves an IoU (0.62) that is comparable to CMR (0.64), even though the latter is trained with keypoints supervision.
To find the 3D template’s vertex that corresponds to an annotated 2D keypoint of a source image, we first render all 3D vertices using the source image’s predicted pose . Then, is the vertex whose 2D projection lies closest to the keypoint . Next, we render the point with the target image’s predicted pose and compare it to its ground truth keypoint to compute PCK. Figure 10 (b) demonstrates the keypoint transfer results by predicted camera pose. Table 1 shows that our model achieves favourable performance against the baseline method .
4 Ablation Studies
In this section, we discuss the contribution of each proposed module: (i) The semantic consistency constraint discussed in Section 3.1. (ii) The texture cycle consistency introduced in Section 3.3. (iii) The improved SCOPS method introduced in Section 3.4. We evaluate on the CUB-200-2011 dataset and use the mask reprojection accuracy as well as the keypoint transfer (via texture flow and via camera pose) accuracy discussed in Section 2 as our metrics.
As shown in Table 2 (b) vs. (d) our baseline model trained without semantic consistency constraint performs much worse at the keypoint transfer task than our full model, indicating this baseline model predicts incorrect texture flow and camera views. We note that this baseline model achieves better mask IoU because the model trained without any constraint is more prone to overfit to the 2D silhouette observations.
Our model trained without the texture cycle consistency constraint achieves worse performance (Table 2 (b) vs.(c)) at transferring keypoints using the predicted texture flow. This proves the effectiveness of the texture cycle consistency constraint in encouraging the model to learn better texture flow.
Our method improves SCOPS by segmenting parts more consistently in terms of their shape and size as shown in Figure 10. However, this is non-trivial to quantify numerically as the ground-truth segmentation labels for the parts are not available in all the dataset that we use . Instead, we indirectly measure the improvement by training two models, each of which uses the semantic part segmentation predicted either by the original or the improved SCOPS method. As shown in Table 2 (b) vs.(e), our keypoint transfer performance drops by 5.3% and 2.5% via texture flow and camera pose if we use the original SCOPS model. More qualitative visualizations of the improvement of SCOPS can be found in the appendix.
Conclusion
In this work, we learn a model to reconstruct 3D shape, texture and camera pose from single-view images, with only a category-specific collection of images and silhouettes as supervision. The self-supervised framework enforces semantic consistency between the reconstructed meshes and images and largely reduces ambiguities in the joint prediction of 3D shape and camera pose from 2D observations. It also creates a category-level template and a canonical semantic UV map, which capture the most representative shape characteristics and semantic parts of objects in each category, respectively. Experimental results demonstrate the efficacy of our proposed method in comparison to the state-of-the-art supervised category-specific reconstruction methods.
Appendix
In this appendix, we provide additional details, discussions, and experiments to support the original submission. We first discuss implementation details in Section 6. Then, we visualize the contribution of each module via ablation studies in Section 7. We further present more quantitative and qualitative results in Section 8 and Section 9. Finally, we describe failure cases and limitations of the proposed method in Section 10.
In the M-step (Section 3.2 in the submission), we update the template by decoding the averaged shape feature via the shape decoder. Instead of using all training samples to obtain the averaged feature, we select a subset of the training samples to form a set and compute the averaged feature for the samples in this set. In the following, we explain why and how to form this set used in Eq.(4) of the submission. Empirically we found that for several categories, there exist ambiguities that produce inconsistent mesh reconstructions, e.g., side-view images of horses could be reconstructed with their heads on either the left or the right side. Aggregating such instance meshes leads to incorrect estimation of the category-level template. To resolve this, we select a subset of reconstructed meshes whose viewpoints roughly match (e.g. horses with heads on the left side). To do so, from the meshes reconstructed for all the training images, we first choose the instance with the most “reliable” reconstruction results, i.e., the instance whose rendered silhouette has the largest intersection over union (IoU) with its corresponding ground truth silhouette, as an exemplar (e.g. a horse shape with its head on the left). We then use the top training samples with meshes that are most similar to the exemplar mesh to form the subset in Eq.(4) (e.g., all chosen horse samples have heads on the left). We measure the similarity between an individual instance mesh and the exemplar mesh by computing the IoU between their rendered silhouettes.
2 Network Architecture and Other Objectives
In addition to the objectives discussed in Section 3.2 of the submission, we further utilize a graph Laplacian constraint to encourage the reconstructed mesh surface to be smooth , and adopt an edge regularization to penalize irregularly-sized faces as in . More details can be found in .
where and are the reconstruction and discriminator networks, respectively. Figure 12 illustrates the adversarial objective.
3 Network Training
We train the reconstruction network with an initial learning rate of and gradually decay it by a factor of every 2000 iterations. The network is trained for two EM training rounds (each training round contains one E and M-step) on four NVIDIA Tesla V100 GPUs for two days. We found that two rounds of EM training are sufficient to generate high-quality reconstruction results. During the inference stage, the model takes 0.022 seconds to reconstruct a 3D mesh from a sized single-view image on a single NVIDIA Tesla V100 GPU. In Figure 13, we show the learned template shape as well as the semantic parts after the first (left figure) and second M-steps (right figure), where both the template shape and the semantic parts after the second M-step are better than the first.
Ablation Studies
We show the results of three baselines in Figure 16. The experimental settings for each are illustrated in Table 3 and are the following: (a) a basic model trained with only the texture cycle consistency constraint described in Section 3.3 in the submission, but without any other proposed modules, i.e., the category-level template, the semantic consistency constraint and the adversarial training; (b) learning the model in (a) together with the category-level template; and (c) learning the model in (b) with the additional semantic consistency constraint.
As shown in Figure 16, the basic model (a) reconstructs meshes that only appear plausible from the observed view to match the 2D supervision (images and silhouettes). It fails to generate plausible results for unobserved views, e.g., for all the 3 examples. On adding template shape learning (Section 3.2 in the submission) to (a), the model in (b) learns more plausible reconstruction results across different views. This is because it is easier for the model to learn residuals w.r.t a category-level template compared to w.r.t a sphere, to match the 2D observations. However, without semantic part information, the model still suffers from the “camera-shape ambiguity” discussed in Section 1 of the manuscript. For instance, the head of the template is deformed to form the tail and the wing’s tip in the first and second examples, respectively in Figure 16. By additionally including the semantic consistency constraint in the model (c), the network is able to reduce the “camera-shape ambiguity” and predict the correct camera pose as well as the correct shape. Furthermore, adding adversarial training introduces better reconstruction details, as shown in Figure 16 (d). For instance, the bird may have more than two feet without the adversarial training constraint as demonstrated in the third example in Figure 16.
In addition, we demonstrate the effectiveness of the texture flow consistency constraint by visualizing the keypoint transfer results in Figure 17. The model trained without this constraint performs worse than our full model, especially when the bird has a uniform color, e.g., the second and the last examples in Figure 17. Figure 17 also shows that the proposed method performs favourably against the baseline CSM method.
2 Ablation Studies on Semantic Consistency Constraints
We show an ablation study of the probability and vertex-based semantic consistency constraints in Table 4, where both constraints contribute to the reconstruction network.
More Quantitative Evaluations
For the keypoint transfer task, in Figure 15, We demonstrate the precision versus recall curve of our method (via texture flow) and of the CSM method on the CUB-200-2011 test dataset. Our method, even without the template prior, outperforms the baseline CSM method in terms of the Keypoint Transfer AP metric (APK, ).
More Qualitative Evaluations
We show more qualitative results for birds in Figure 18. We also show one application of our model to reconstruct 3D meshes of 2D bird paintings in Figure 11. Reconstruction of rigid objects (cars and motorbikes) is demonstrated in Figure 20, horses and cows in Figure 19, and penguins and zebras in Figure 21. Note that we use six semantic parts for the car category to encourage the SCOPS method to differentiate between the fronts and the sides of cars. For other objects, we use four semantic parts.
Failure Case and Limitations
Our work is the first to tackle the challenging self-supervised single-view 3D reconstruction task. Impressive as the results are, this challenging task is far from being solved. In this section, we discuss typical failure cases and limitations of the proposed method.
First, our method utilizes the SCOPS method to provide semantic part segmentation, and so it suffers when the semantic part segmentation is not accurate, as shown in the first row of Figure 15. Second, our model struggles to predict camera poses that are rare in the training dataset. For instance, the bird in the second row in Figure 15 presents a rare case where the camera is located very close to the bird, which is not correctly predicted by our model. Third, our model, including the semantic template, is trained in a fully data-driven way. It captures the major shape characteristics of each instance but ignores some details, e.g., the two wings of flying birds, and the legs of zebras or horses are not separated, as shown in Figure 18 and Figure 21. We leave all these failure cases and limitations to future works.