Implicit Mesh Reconstruction from Unannotated Image Collections
Shubham Tulsiani, Nilesh Kulkarni, Abhinav Gupta
Introduction
Inferring 3D structure of diverse categories of objects from images in the wild (see Figure 1) has been one of the long term goals in computer vision. Despite the decades of progress in computing, graphics and machine learning, we do not yet have systems that can infer the underlying 3D for objects in natural images. This is in stark contrast to the progress witnessed in related problems such as object recognition and detection where we have developed scalable methods that make accurate predictions for hundreds of categories for images in the wild. Why is there this disconnect between progress in 3D perception and 2D understanding, and what can we do to overcome it? We argue that the central bottleneck for 3D perception in the wild has been the strong reliance on 3D supervision, or the low expressivity of previous models (e.g., fixed 3D template). In this work, we aim to bypass these bottlenecks and present an approach that can learn inference of a deformable 3D shape using only the form of data that 2D recognition systems leverage – in the wild category-level image collections with (approximate) instance segmentations.
Given a single input image, our goal is to be able to infer the 3D shape, texture and camera pose for the underlying object. A scalable solution needs to have two ingredients: (a) 3D modeling which can handle instance variations and deformations of objects; (b) require minimal supervision to allow scalability across diverse categories. There have been recent attempts on both these axes. For example, a recent approach learns explicit 3D representations from image collections, and can handle instance variations and pose variations. But this approach crucially relies on 2D keypoint annotations to guide learning, thereby making it difficult to scale beyond a handful of categories. On the other hand, Kulkarni et al. learn pixel to surface mappings that are consistent with a global (articulated) 3D template, and show that this can help learn accurate prediction. This approach bypasses keypoint supervision, but cannot model any shape variations (fat vs thin bird) and additionally requires 3D part supervision for modeling articulations.
Our approach handles both, shape and pose variation by inferring implicit category-specific shape representations. To bypass the need for direct supervision, we also infer the corresponding texture and enforce reprojection consistency between our predictions and the available images and segmentations. Drawing inspiration from the work by Kulkarni et al. , we use a geometric cycle-consistency loss between global 3D and local pixel to surface predictions to derive additional learning signal from unannotated image collections. Leveraging a single 3D template shape per category as initialization, we learn the category-level implicit shape space from image collections. We show sample shape predictions obtained using our approach in Figure 1 and also visualize the 3D shape with predicted texture from the predicted and novel camera viewpoints. We observe that, despite the lack of direct supervision, our method effectively captures the shape variation across instances e.g. head bending down, length of tail. As illustrated in Figure 1, our approach is applicable across a diverse set of object categories, and the reliance on only image collections with approximate segmentation masks allows learning in settings where previous approaches could not.
Related Work
The recent success of deep learning has resulted in a number of learning based approaches for the task of single-view 3D inference. The initial approaches showed impressive volumetric inference results using synthetic data as supervision, and these were then generalized to other representations such as point clouds , octrees , or meshes . However, these approaches relied on ground-truth 3D as supervision, and this is difficult to obtain at scale or for images in the wild. There have therefore been attempts to relax the supervision required e.g. by instead using multi-view image collections . Closer to our setup, several approaches have also addressed the task of learning 3D inference using only single-view image collections, although relying on additional pose or keypoint annotations. While these approaches yield encouraging results, their reliance on annotated 2D keypoints limits their applicability for generic categories. Our work also similarly learns from image collections, but does so without using any supervision in the form of 2D keypoints or pose; and this allows us to learn from image collections of objects in the wild.
Learning 3D from Unannotated Image Collections.
Our motivation of learning 3D from unannotated image collections is also shared by some recent works . Nguyen-Phuoc et al. use geometry-driven generative modeling to learn 3D structure, but their approach only infers a volumetric feature capable of view synthesis. Unlike our work, this does not output a tangible 3D shape or texture for the underlying object. Closer to our approach, Kulkarni et al. predict explicit 3D representations by proposing a cycle-consistency loss between inferred 3D and learned pixelwise 2D to 3D mappings. However, their method only produces a limited 3D representation in the form of a rigid or articulated template, and cannot handle intra-instance shape variations e.g. fat vs thin bird. Our implicit shape representation allows capturing such variation in addition to the instance texture, and we generalize their cycle consistency loss for our representation.
Implicit 3D Representations.
Several recent works using implicit functions have shown impressive results on the tasks of 3D reconstruction. Unlike explicit representations (e.g. meshes, voxels, point clouds) these methods learn functions to parameterize a 3D volume or surface. Our representation is inspired by work from Groueix et al. which learns a mapping conditioned on the latent code from points on a 2D manifold to a 3D surface, and we further equip this with an implicit texture as in Oechsle et al. . However, all these prior approaches require settings with ground-truth 3D and texture supervision available. In contrast, our work leverages these representations in an unsupervised setup with corresponding technical insights to make learning feasible e.g. use of a ‘texture flow’ and incorporation of category-specific shape consistency.
Approach
Given an image of an object, we aim to infer its 3D shape, texture, and camera pose. Moreover, we want to learn this inference using only image collections with instance segmentations as supervisory signal. Towards this, our approach leverages implicit shape and texture representations, and enforces geometric consistency with available image to bypass the need of direct supervision.
Specifically, we leverage geometry-driven objectives to learn an encoder which predicts the desired properties from the input image – here denotes a weak perspective camera, and correspond to latent variables that instantiate the underlying shape and texture. We first describe in Section 3 the category-specific implicit shape representation pursued and then describe the proposed learning framework in Section 3.2. While we initially consider only shape inference, we show in Section 3.3 how our approach can also incorporate texture prediction.
Here and are both modeled as neural networks, and represent the implicit functions for a category-level mean shape and instance level deformation respectively. To overcome ambiguities and possible degenerate solutions when learning without supervision, we initialize for each category to match a manually chosen template 3D shape (see appendix for details and visualizations).
Almost all naturally occurring objects, as well as several man-made ones, exhibit reflectional symmetry. We incorporate this by constraining the learnt mean shape and deformations to be symmetric along the plane. For example, the deformation is constrained to be symmetric as follows (where denotes a reflection function):
Encouraging Locally Rigid Transforms.
In our formulation, the instance-specific deformation is required to capture any change from the category-level mean shape. This change in 3D structure can stem from intrinsic variation e.g. a bird can be thin or fat, or can be caused by articulation e.g. the same horse would induce different deformations if its head is bent down or held upright. As articulations can be viewed as local rigid transformations of the underlying shape, we encourage the learnt deformations to explain the variation using locally rigid transformations if possible. We note that under any rigid transform, the distance between two corresponding points remains unchanged, and incorporate a regularization that penalizes the mean of this variation across local neighborhoods. Denoting by a local neighborhood of , this objective is:
2 Learning 3D Inference via Reprojection Consistency
While we do not have direct supervision available for learning the shape and pose inference, we can nevertheless derive supervisory signal by encouraging our 3D predictions be geometrically consistent with the available image evidence. Concretely, we enforce that the predicted 3D shape, when rendered according to the predicted camera, matches the foreground mask, while specially emphasizing boundary alignment. Following a surface mapping consistency formulation by Kulkarni et al. , we also implicitly encourage semantically similar regions across instances to be ‘explained’ by consistent regions of the deformable shape space.
Given the inferred (implicit) 3D shape , we recover an explicit mesh by sampling the implicit function at a fixed resolution. We then use a differentiable renderer to obtain the foreground mask for this predicted mesh camera , and define a loss against the ground-truth mask : .
While this loss coarsely aligns the predictions to the image, it does not emphasize details such as long tails of birds, or animal legs. We therefore incorporate an objective as proposed by Kar et al. – points should project within the foreground, and contour points should have some projected 3D points nearby (see appendix for formulation).
Regularization via Pixel to Surface Mappings.
In our formulation, each point on the predicted 3D shape corresponds to a unique point on the surface of a unit sphere. To ensure that the inferred 3D is consistent across different instances in a category, we would ideally like to enforce that semantically similar regions across images are ‘explained’ by similar regions of the unit sphere e.g. the projecting on the horse head is similar across instances. However, we do not have access to any semantic supervision to directly operationalize this insight, but we can do so indirectly.
Leveraging Optional Keypoint Supervision.
While our goal is to learn 3D reconstruction with minimal supervision, our approach easily allows using semantic keypoint labels e.g. nose, tail etc. if available. For each semantic keypoint annotated in the 2D images, we first manually annotate a corresponding 3D keypoint on the category-level template mesh. This allows us to associate a unique spherical coordinate with each keypoint for a category. Given an input training image with annotated 2D keypoint locations , we can penalize the reprojection error.
Note that we only leverage this supervision in certain ablations, and that all other results (including visualized predictions) are obtained without using this additional signal i.e. only mask supervision.
Training Details.
Our reprojection and cycle consistency losses depend on both, the predicted camera and the inferred shape . Unfortunately, learning these together is susceptible to local minima where only a narrow range of poses are predicted. We follow suggestions from prior work to overcome this and predict multiple diverse pose hypotheses and their likelihoods, and minimize the expected loss. Additionally, as the inferred poses in initial training iterations are often inaccurate, the learned deformations are not meaningful. We therefore only allow training after certain epochs. We use a pretrained ResNet-18 based encoder as , and a simple 4 layer feedforward MLP to instantiate the shape and texture implicit functions.
3 Texture Prediction via Implicit Texture Flow
Towards capturing the appearance of the depicted object, we predict the texture that gets overlaid on the predicted (implicit) 3D shape, and learn this inference by enforcing consistency with the observed foreground pixels. As our 3D shape is parametrized via a deformation of a sphere, we simply need to associate each point on the sphere with a corresponding texture to induce a textured 3D shape.
To derive learning signal for this prediction, we follow Kanazawa et al. and differentiably render the predicted 3D shape with the implied texture and penalize a perceptual loss against the foreground pixels of the image. We additionally also incorporate a term encouraging the texture flow to sample from foreground pixels instead of background ones.
Experiments
We present empirical and qualitative results across a diverse set of categories, and leverage several existing datasets to obtain the required training data. Across all these datasets, we only use the images and annotated (or automatically obtained) segmentation masks for training, but use the additional available annotations e.g. keypoints or approximate 3D for evaluation. We download a representative template model (used to initialize ) for all the categories from .
(Aeroplane, Car). We use images from the PASCAL3D+ dataset, which combines images from PASCAL VOC and Imagenet . For the former set, a manually annotated foreground mask is available, whereas an automatically obtained one is used for the Imagenet subset. We follow the splits used in prior work , but unlike them, do not use the keypoint annotations for training. As annotations for approximate 3D models are available on this dataset, it allows us to empirically measure the reconstruction accuracy.
Curated Animate Categories
(Bird, Cow, Horse, Sheep). We use the CUB-200-2011 dataset for obtaining bird images with segmentation masks. For the other categories, we use the splits from which, similar to PASCAL3D+, combine images from VOC and Imagenet. Across all these categories, we have keypoint annotations available for images in the test set, and we use these for indirect evaluation of the quality and consistency of the inferred 3D.
Quadrupeds from Imagenet
(Lion, Bear, Elephant, and 20 others). We also apply our method on categories from Imagenet using automatically obtained segmentations . We use the images from Kulkarni et al. who (noisily) filter out instances with significant truncation and occlusion.
2 Evaluation using Semantic Keypoints
3 3D Reconstruction Accuracy
In Section 4.2, we compared our approach to previous works that learn using similar supervision, but only infer a constrained 3D representation. However, there have also been prior approaches which, similar to ours, allow more expressive 3D inference, but require additional keypoint or pose supervision. We compare our learned 3D inference to these (relatively) strongly supervised methods using the PASCAL3D+ dataset which has approximate 3D ground-truth available in the form manually selected templates.
We report in Table 2 the mean intersection over union (IoU) of the predicted 3D shape with the available ground-truth on held out test images. We compare our approach to a deformable model fitting , a volumetric prediction , and an explicit mesh inference approach – all of which rely on keypoint or pose labels for learning. We observe that across both the examined categories, our approach, using only foreground mask as supervision, yields competitive (and sometimes better) performance.
4 Learning from Unannotated Image Collections
To allow direct or indirect empirical evaluation, we so far only examined categories where annotated image collections are available. However, as we do not leverage these annotations for learning, we can go beyond this limited set of classes and apply our approach to generic categories from Imagenet. In particular, we consider several quadruped categories e.g. lion, bear, zebra etc., and train our approach with automatically obtained approximate instance segmentations using PointRend . We visualize our predictions in Figure 5 and also show additional random results in the supplementary. We see that our predictions can capture variation e.g. thin or fat body, head bending down, sitting vs standing etc., and that the inferred texture for visible and invisible regions is also meaningful.
5 Ablations
We ablate various terms in our learning objective using the semantic keypoint reprojection (PCK-P) and transfer (PCK-T) metrics. In particular, we examine whether incorporating consistency loss with pixelwise surface mappings is helpful, and if prioritizing locally rigid transforms and boundary alignment is beneficial. We report in Table 3 the mean accuracy for both evaluations across the four animate categories evaluated (bird, horse, cow, and sheep). We observe that all the components improve performance, though we note that some components e.g. encouraging rigidity or boundary alignment lead to more significant qualitative improvements.
Discussion
We presented an approach for learning implicit shape and texture inference from unannotated image collections. Although this enabled 3D prediction for a diverse set of categories, several challenges still remain towards being able to reconstruct thousands of object classes in generic images. As we model shapes via deformation of a (learned) template, our model does not allow for large or topological shape changes that maybe common in artificial object categories (e.g. chairs). Additionally, our reprojection based objectives implicitly assume that the object is largely unoccluded, and it would be interesting to generalize these objectives to allow partial visibility. Lastly, our results do not always capture the fine details or precise pose, and it may be desirable to leverage additional learning signal if available e.g. from videos. While there is clearly more progress needed to build systems that can reconstruct any object in any image, we believe our work represents an exciting step towards learning scalable and self-supervised 3D inference.
References
Appendix
We visualize the resulting meshes learned for some categories in Figure 6, and and the sphere and the resulting shapes to highlight correspondence. We observe that these capture the underlying 3D mesh well, but exhibit certain artifacts e.g. hind legs of cow (top right). While this initialization helps in learning, we note that the learned shape space and inferred deformations allow us to model shapes that significantly vary from this initial template e.g. fat vs thin bird, articulated animals. Please see the main text and additional visualizations below for examples.
Boundary Reprojection Consistency.
To encourage our inferred 3D shapes to match the foreground mask boundary, we adapt the objectives proposed by Kar et al. . Specifically, we encourage that the projected 3D points should lie inside the object, and that each point on the 2D boundary of the foreground mask should be close to some projected point(s).
Correspondence Transfer via 3D Inference.
As our category-specific implicit shape representation deforms a common sphere to yield the shape for any given instance, each mesh point is associated with a unique spherical coordinate. Given a predicted shape and pose for an image , we can therefore render a per-pixel spherical coordinate (akin to rendering a textured mesh) as shown in Figure 7. These per-pixel spherical mappings subsequently allow us to transfer semantics (e.g. keypoints) from a source image to a target image. For example, given a keypoint annotation in a source image, we can use to rendered spherical coordinate at that (or nearest) pixel as a query to find the corresponding point in a target image.