Pri3D: Can 3D Priors Help 2D Representation Learning?
Ji Hou, Saining Xie, Benjamin Graham, Angela Dai, Matthias Nießner
Introduction
In recent years, we have seen rapid progress in learning-based approaches for semantic understanding of 3D scenes, particularly in the tasks of 3D semantic segmentation, 3D object detection, and 3D semantic instance segmentation . Such approaches leverage geometric observations, exploiting the representation of points , voxels , or meshes to obtain accurate 3D semantics. These have shown significant promise towards realizing applications such as depth-based scene understanding for robotics, as well as augmented or virtual reality. In parallel to the development of such methods, the availability of large-scale RGB-D datasets , has further accelerated the research in this area.
One advantage of learning directly in 3D in contrast to learning solely from 2D images is that methods operate in metric 3D space; hence, it is not necessary to learn view-dependent effects and/or projective mappings. This allows training 3D neural networks from scratch in a relatively short time frame and typically requires a (relatively) small number of training samples; e.g., state-of-the-art 3D neural networks can be trained with around 1000 scenes from ScanNet. Our main idea is to leverage these advantages in the form of 3D priors for image-based scene understanding.
Simultaneously, we have seen tremendous progress on representation learning in the image domain, mostly powered by the success of recent contrastive learning based methods . The exploration in 2D representation learning heavily relies on the paradigm of instance discrimination, where different augmented copies of the same instance are drawn closer. Different invariances can be encoded from those low-level augmentations such as random cropping, flipping and scaling, as well as color jittering. However, despite the common belief that 3D view-invariance is an essential property for a capable visual system , there remains little study linking the 3D priors and 2D representation learning. The goal of our work is to explore the combination of contrastive representation learning with 3D priors, and offer some preliminary evidence towards answering an important question: can 3D priors help 2D representation learning?
To this end, we introduce Pri3D, which aims to learn with 3D priors in a pre-training stage and subsequently use them as initialization for fine-tuning on image-based downstream tasks such as semantic segmentation, detection, and instance segmentation. More specifically, we introduce geometric constraints to a contrastive learning scheme, which are enabled by multi-view RGB-D data that is readily available. We propose to exploit geometric correlations through implicit multi-view constraints between different images through the correspondence of pixels which correspond to the same geometry, as well as explicit correspondence of geometric patches which correspond to image regions. This imbues geometric knowledge into the learned representations of the image inputs which can then be leveraged as pre-trained features for various image-based vision tasks, particularly in the low training data regime.
We demonstrate our approach by pre-training on ScanNet under these geometric constraints for representation learning, and show that such self-supervised pre-training (i.e., no semantic labels are used) results in improved performance on 2D semantic segmentation, instance segmentation and detection tasks. We demonstrate this not only on ScanNet data, but also generalizing to improved performance on NYUv2 semantic segmentation, instance segmentation and detection tasks. Moreover, leveraging such geometric priors for pre-training provides robust features which can consistently improve performance under a wide range of amount of training data available. While we focus on indoor scene understanding in this paper, we believe our results can shed light on the the paradigm of representation learning with 3D priors and open new opportunities towards more general 3D-aware image understanding.
A first exploration of the effect of 3D priors for 2D image understanding tasks, where we demonstrate the benefit of 3D geometric pre-training towards complex 2D perception such as semantic segmentation, object detection, and instance segmentation.
A new pre-training approach based on 3D-guided view-invariant constraints and geometric priors from color-geometry correspondence, which learns features that can be transferred to 2D representations, complementing and improving image understanding across multiple datasets.
Related Work
Research in 3D scene understanding has recently been spurred forward with the introduction of larger-scale, real-world 3D scanned scene datasets . We have seen notable progress in development of methods for semantic segmentation , object detection , and instance segmentation in 3D. In particular, the introduction of sparse convolutional neural networks have presented a computationally-efficient paradigm producing state-of-the-art results in such tasks. Inspired by the developments in 3D scene understanding, we introduce learned geometric priors to representation learning for image-based vision tasks, leveraging a sparse convolutional backbone for 3D features used during pre-training.
In the past year, we have also seen new developments in 3D representation learning. PointContrast first showed that unsupervised, contrastive-based pre-training improves performance across various 3D semantic understanding tasks. Hou et al. introduces spatial context into 3D contrastive pre-training, resulting in improved performance in 3D limited annotation and data scenarios. Zhang et al. introduces a instance-discrimination-style pre-training approach that directly operates on depth frames. Our approach bridges these concepts into feature learning that can be transferred to 2D image understanding tasks.
D Contrastive Representation Learning.
Representation learning has driven significant efforts in deep learning; on the image domain, pre-training a network on a rich set of data has been shown to improve performance in fine-tuning for a smaller target dataset for various applications. In particular, the contrastive learning framework to learn representations from similar/dissimilar pairs of data has been demonstrated to show incredible promise . Notably, using an instance discrimination task in which positive pairs are created with data augmentation, MoCo shows that unsupervised pre-training can surpass various supervised counterparts in detection and segmentation tasks, and SimCLR further reduces the gap to supervised pre-training in linear classifier performance. Our approach leverages multi-view geometric information to augment contrastive learning and imbue robust geometric priors into learned feature representations.
Multi-Modality Learning
CLIP firstly proposes to train on images but with natural language supervision, and achieves significant results on zero-shot learning. BPNet proposes a bidirectional projection module to mutually leverage 2D and 3D information for semantic segmentation task. 3D-to-2D Distillation introduces additional 3D network in the training phase to embed 3D features for 2D semantic segmentation task. Existing works need to modify networks or add fusion modules in the training and/or inference phases. To this end, our method is more flexible as our pre-trained weights can be directly used like the ImageNet pre-trained model without any further modules or 3D/NLP data in the downstream tasks.
Correspondences Matching
Schmidt et al. advocates a new approach to learning visual descriptors for dense correspondence estimation for the re-localization purpose, e.g., in the SLAM context. Schuster et al. presents a robust, unified descriptor network leveraging stacked dilated convolutions (SDC) for larger receptive field to better estimate dense pixel matching. HumanGPS estimates dense correspondences between human images under arbitrary camera viewpoints and body poses. Existing works focus on 2D-2D correspondences matching problem itself. Our approach uses 2D-3D as well as 2D-2D view-invariant correspondences matching as pretext task to embed 3D priors for 2D downstream tasks.
Learning Representations from 3D Priors
In this section, we introduce Pri3D; our key idea is to leverage constraints from RGB-D reconstructions, now readily available in various datasets , to embed 3D priors in image-based representations. From a dataset of RGB-D sequences, each sequence consists of depth and color frames, and , respectively, as well as automatically-computed 6-DoF camera pose alignments (mapping from each camera space to world space) from state-of-the-art SLAM, all resulting in a reconstructed 3D surface geometry . We first revisit the simplest 3D priors, i.e. depth prediction . After seeing positive signals, we further explore more 3D priors. Specifically, we observe that multi-view constraints can be exploited in order to learn view-invariance without the need of costly semantic labels. In addition, we learn features through geometric representations given by the obtained geometry in RGB-D scans, again, without the need of human annotations. For both, we use state-of-the-art contrastive learning in order to constrain the multi-modal input for training. We show that these priors can be embedded in the image-based representations such that the learned features can be used as pre-trained features for purely image-based perception tasks; i.e., we can perform tasks such as image segmentation or instance segmentation on a single RGB image. An overview of our approach is shown in Figure 2.
In 2D constrative pre-training algorithms, a variety of data augmentations are used for finding positive matching pairs, such as MoCo and SimCLR . For instance, they use random crops as self-supervised constraints within the same image for positive pairs, and correspondences to crops from other images as negative pairs. Our key idea is that with the availability of 3D data for training, we can leverage geometric knowledge to provide matching constraints between multiple images that see the same points. To this end, we use the ScanNet RGB-D dataset which provides a sequence of RGB-D images with camera poses computed by a state-of-the-art SLAM method , and reconstructed surface geometry . Note that both the pose alignments and the 3D reconstructions were obtained in a fully-automated fashion without any user input.
For a given RGB-D sequence in the train set, our method then leverages the 3D data to finding pixel-level correspondences between 2D frames. We consider all pairs of frames from the RGB-D sequence. We then back-project frame ’s depth map to camera space, and transform the points into world space by . The depth values of frame are similarly transformed into world space. Pixel correspondences between the two frames are then determined as those whose 3D world locations lie within cm of each other (see Figure 3). We use the pairs of frames which have at least 30% pixel overlap, with overlap computed as number of corresponding pixels in both frames divided by total number points in the two frames. In total, we sample around 840k pairs of images from the ScanNet training data.
In the training phase, a pair of sampled images is input to a shared 2D network backbone. In our experiments, we use a UNet-style backbone with ResNet architecture as an encoder, but note that our method is agnostic to the underlying encoder backbone. We then consider the feature map from decoder of the 2D backbone, where its size is half of the input resolution. For each image in the pair, we use the aforementioned pixel-to-pixel correspondences which refer to the same physical 3D point. Note that these correspondences may have different color values due to view-dependent lighting effects but represent the same 3D world location; additionally, the regions surrounding the correspondences appear different due to different viewing angles. In this fashion, we treat these pairs of correspondences as positive samples in contrastive learning; we use all non-matching pixels as negatives. Non-matching pixels are also defined within the set of correspondences. For a pair of frames with pairs of correspondences as positive samples, we use all negative pairs (each of pixels from the first frame with each non-matching pixel from the second). Non-matching pixel-voxels are defined similarly but from a pair of frame and 3D chunk.
Between the features of matching and non-matching pixel locations, we then compute a PointInfoNCE loss , which is defined as:
where is the set of pairs of pixel correspondences, and represents the associated feature vector of a pixel in the feature map. By leveraging multi-view correspondences, we apply implicit 3D priors – without any explicit 3D learning, we imbue view-invariance in the learned image-based features.
2 Geometric Prior
In addition to multi-view constraints, we also leverage explicit geometry-color correspondences inherent to the RGB-D data during training. For an RGB-D train sequence, the geometry-color correspondences are given by associating the surface reconstruction with the RGB frames of the sequence. For each frame , we compute its view frustum in the world space. A volumetric chunk of is then cropped from the axis-aligned bounding box of the view frustum. We represent as a cm resolution volumetric occupancy grid from the surface. We thus consider pairs of color frames and geometric chunks .
From the color-geometry pairs , we compute pixel-voxel correspondences by projecting the depth values for each pixel in the corresponding frame into world space to find an associated occupied voxel in that lies within cm of the 3D location of the pixel.
During training, we leverage the color-geometry correspondences with a 2D network backbone and a 3D network backbone. We use a UNet-style architecture with ResNet encoder for the 2D network backbone, and a UNet-style sparse convolutional 3D network backbone. Similarly to view-invariant training, we also take the output from the decoder of 2D network backbone where its output size is half of the input resolution. We then use the pixel-voxel correspondences in for contrastive learning, with positives as all matching pixel-voxel pairs and negatives as all non-matching pixel-voxel pairs. We apply the PointInfoNCE loss (Equation 1) with as the 2D features of a pixel, and is the feature vector from its 3D correspondence, and the set of 2D-3D pixel-voxel correspondence pairs.
3 Joint Learning
We can leverage not only the view-invariant constraints and geometric priors during training, but also learn jointly from the combination of both constraints. We can thus employ a shared 2D network backbone and a 3D network backbone, with the 2D network backbone constrained by both view-invariant constraints and as the 2D part of the geometric prior constraint.
During training, we consider of overlapping color frames and as well as and which have geometric correspondence with , respectively. The shared 2D network backbone then processes and computes the view-invariant loss from Section 3.1. At the same time, and are processed by the 3D sparse convolutional backbone, with the loss (discussed in 3.2) relative to the features of and respectively. This embeds both constraints into the learned 2D representations.
Experimental Setup
Our approach aims to embed 3D priors into the learned 2D representation by leveraging our view-invariant and geometric prior constraints. In this section, we introduce our detailed experimental setup for pre-training with an RGB-D dataset and fine-tuning on downstream 2D scene understanding tasks.
As described in the previous section, our pre-training method leverages the pixel-to-pixel and geometry-to-color correspondences for view-invariant contrastive learning. The specific form of our pre-training objective requires a feature extractor capable of providing per-pixel or per-3D-point features for the backbone architecture, as the positive and negative matches are defined over 2D pixels or 3D locations.
Our meta-architectures for both view-invariant constraints and geometric priors are U-Nets with residual connections. The encoder part of the U-Net is a standard ResNet. For view-invariant learning with 2D image inputs, we use ResNet18 or ResNet50 as encoders. The decoder part of the U-Net architecture consists of convolutional layers and bi-linear interpolation layers. For learning geometric priors from 3D volumetric occupancy input, we use sparse convolutions , specifically a Residual U-Net-32 backbone implemented with MinkowskiEngine , using a cm voxel size.
Stage I: Pri3D encoder initialization.
We empirically found that for the pre-training phase, good initialization of the encoder network is critical to make learning robust. Instead of starting with random initialization, we initialize the encoder with network weights trained on ImageNet (i.e. we pre-train the network for pre-training). The whole pipeline can be seen as a two-stage framework. We note that our method aims to improve the general representation learning, thus is not tied to a specific learning paradigm (e.g. supervised pre-training or self-supervised pre-training). From this perspective, we can leverage supervised pre-training of ResNet encoders with ImageNet data for encoder initialization for pre-training. We name this model Pri3D.
Although the use of a supervised ImageNet pre-trained initialization is a common practice, for completeness we also evaluate Pri3D in an unsupervised pipeline without using ImageNet labels. Results suggest that Pri3D does not rely on any semantic supervision (e.g. ImageNet labels) to succeed, and still is able to achieve a substantial gain in this setup. We name this variant Unsupervised Pri3D. Further results of Unsupervised Pri3D are demonstrated in supplementary materials.
Stage II: Pri3D pre-training on ScanNet.
Our pre-training method is enabled by the inherent geometry and color information present in the RGB-D data sequences. For pre-training, we leverage the color image and geometric reconstructions provided by the automatic reconstruction pipeline of ScanNet ; note that we do not use the semantic annotations during pre-training. We also use a depth proxy loss by default in our experiments. Please refer to supplementary materials for details. ScanNet contains 2.5M images from 1513 ScanNet train video sequences. We regularly sample every frame without any other filtering (e.g., no control on viewpoint variation), and compute the set of overlapping pairs of frames that have pixel overlap, resulting in k frame pairs for which we compute their corresponding geometric chunks for each image, in order to apply both our view-invariant and geometric prior constraints.
Downstream Fine-tuning.
We evaluate our Pri3D models by fine-tuning them on a suite of downstream image-based scene understanding tasks. We use two datasets, ScanNet and NYUv2 , and the three tasks of semantic segmentation, object detection, and instance segmentation. As our pre-training dataset is ImageNet and ScanNet, fine-tuning on ScanNet represents a scenario of in-domain transfer—it would be interesting to know if the 3D priors can help with 2D representations for image-based tasks on the same dataset. We further evaluate the performance of Pri3D on the NYUv2 dataset which maintains different statistics. This represents a out-of-domain transfer scenario. For semantic segmentation tasks, we directly use the U-Net architecture for dense prediction. The encoder and decoder networks are both pre-trained with Pri3D. For instance segmentation and detection tasks, we use Mask-RCNN framework implemented in Detectron2 . Only the backbone encoder part is pre-trained.
Implementation details.
For pre-training, we use an SGD optimizer with learning rate 0.1 and batch-size of 64. The learning rate is decreased by a factor of 0.99 every 1000 steps, and our method is trained for 60,000 iterations. For MoCoV2 , we use the official PyTorch implementation. MoCoV2 is trained for 100 epochs with batch size 256. The fine-tuning experiments on semantic segmentation are trained with a batch size of 64 for 80 epochs. The initial learning rate is 0.01, with polynomial decay with power 0.9. All experiments are conducted on 8 NVIDIA V100 GPUs.
Baselines.
As we are using additional RGB-D data from ScanNet, it is important to benchmark our method against relevant baselines in order to answer the question: are 3D priors useful for 2D representation learning?
Supervised ImageNet Pre-training (IN). We use the ImageNet pre-trained weights provided in torchvision; this represents a widely adopted paradigm for image-based tasks. No ScanNet data is involved.
1-Stage MoCoV2 (MoCoV2-IN+SN). We train MoCoV2 on an expanded dataset that combines ImageNet with ScanNet. We explore two strategies: 1) Directly combining the two datasets with shuffled images and 2) mixing minibatches (sampling half images from ImageNet and the other half from ScanNet). In this case, we use ScanNet data but no 3D priors are considered.
2-Stage MoCoV2 (MoCoV2-supINSN). As we use supervised pre-training (IN) in our method as encoder initialization, for fair comparison, we also try one version with (supervised) IN as the encoder initialization, then add another stage to fine-tune MoCoV2 with randomly shuffled ScanNet images. In this case, we use ScanNet data but no 3D priors are used.
Trivial Correspondences. We use our framework but instead of learning from multi-view correspondences, we take one single-view image and create two copies by applying color space augmentations including: RGB jittering, random color dropping and Gaussian blur. Positive matches are defined on pixels at the same location. In this case, we use ScanNet data but no 3D priors are considered.
Depth Prediction. We use the single frame depth prediction as a pre-training task. Note, our approach also leverages the depth prediction as a proxy loss by default. In this case, we use ScanNet data and a simple 3D prior is considered.
Through above baselines, we aim to justify that Pri3D learns to embed 3D priors in 2D representations that lead to an improved downstream performance; it is nontrivial to achieve the goal, given the auxiliary RGB-D dataset.
Results
In this section, we present the results of our downstream fine-tuning results as well as relevant baselines mentioned in the previous section.
We use our pre-trained network weights learned with Pri3D, and fine-tune for 2D semantic segmentation, object detection, and instance segmentation tasks on ScanNet images, demonstrating the effectiveness of representation learning with 3D geometric priors. For fine-tuning, following the standard protocol in the ScanNet benchmark : we sample every 100 2D frames, resulting in 20,000 train images and 5,000 validation images.
2D Semantic Segmentation. We first show fine-tuning for semantic segmentation results in Table 1, in comparison with several baselines that also use ScanNet RGB-D data. We show the applicability of our approach with a standard ResNet50 backbone and a smaller ResNet18 backbone.
Comparing to just training the semantic segmentation model from scratch on downstream dataset ( with ResNet50), all pre-training methods help significantly, even just using the ImageNet pre-training. This confirms the common belief in computer vision that a good 2D representation is essential for good performance on the target task. Several baselines, when adding the ScanNet RGB-D data, also works reasonably well, but not much better than the naive ImageNet Pre-training baseline. This suggests that simply adding the ScanNet data into the representation learning pipeline does not necessarily lead to better results. Our Pri3D variants, including the view-invariant contrastive learning, geometry-color correspondence based contrastive learning and the combination of the two, provides substantially better representation quality that leads to improved semantic segmentation performance. We note that our method has a major performance boost ( absolute mIoU) even compared with the ImageNet Pre-training results. We believe this is an encouraging result and represents a practical use case as ImageNet pre-trained networks are often readily available.
Moreover, we evaluate our approach under limited data scenarios in Figure 5. Our Pri3D pre-training shows an even larger gap when using a small subset of the training images, again compared to the strong ImageNet pre-training baseline. With only 20% of the training data, we are able to recover and of the finetuning performance when using 100% training data, with ResNet50 and ResNet18 backbone respectively.
2D Object Detection and Instance Segmentation. To demonstrate that Pri3D is generalizable for different image-based tasks, we show results on fine-tuning for object detection in Table 2 and instance segmentation in Table 3. For both tasks, we observe similar behavior to the semantic segmentation counterpart. All pre-training methods bring substantial improvement over training from scratch, but Pri3D models stand out and yield more gain compared to ImageNet Pre-training alone (+3.2% and +2.8% AP@0.5 for instance segmentation and detection, respectively). We note that for this set of experiments, we only transfer the encoder weights, discarding the decoder weights in the U-Net architecture for pre-training. This resembles similar practice in language domains (e.g. BERT ) and shows that the main gain of Pri3D is better encoder representations.
To demonstrate our method is agnostic to semantic segmentation backbones, we further show results with PSPNet and DeepLabV3/DeepLabV3+ in Table 4. Pri3D (Ours) consistently outperforms the baseline across different backbone choices.
2 NYUv2
We show that our method learns transferable features across datasets. With Pri3D pre-trained on ScanNet RGB-D data, we explore fine-tuning on NYUv2 for downstream 2D tasks. The NYU-Depth V2 dataset is comprised of video sequences from a variety of indoor scenes, recorded by Microsoft Kinect RGB-D sensors. It contains 1449 densely labeled pairs of aligned RGB and depth images. We use the official split: 795 images for training, 654 images for test. Similar to ScanNet, we also evaluate on 3 popular downstream tasks, 2D semantic segmentation, object detection, and instance segmentation. Table 5 shows the semantic segmentation performance on NYUv2.
We show the semantic segmentation fine-tuning performance on NYUv2 in Table 5; the object detection fine-tuning results in Table 6; and the instance segmentation fine-tuning results in Table 7. The experimental setup is similar to the ScanNet downstream fine-tuning counterpart, and we use supervised ImageNet pre-trained weights for encoder initialization of all methods. For all three tasks, we observe improved performance over different baselines such as training from scratch, training with ImageNet pre-trained weights, and MoCoV2-style pre-training on additional ScanNet data. Compared to the ImageNet pre-training baseline, we achieve a margin of AP@0.5 for instance segmentation, mIoU for semantic segmentation (ResNet50 backbone) and AP@0.5 for object detection.
Conclusion
We have introduced Pri3D, a new method for representation learning for image-based scene understanding tasks. Our core idea is to incorporate 3D priors in a pre-training process whose constraints are applied under a contrastive loss formulation. We learn view-invariant and geometry-aware representations by leveraging multi-view and image-geometry correspondence from existing RGB-D dataset. We show that this results in significant improvement compared to 2D-only pre-training. With limited training data available, we outperform the semantic segmentation baselines by 11.9% on ScanNet. We hope our results can shed light on the the general paradigm of representation learning with 3D priors and open up new opportunities towards 3D-aware image understanding.
Acknowledgments This work was supported by a TUM-IAS Rudolf Moßbauer Fellowship, the ERC Starting Grant Scan2CAD (804724), the German Research Foundation (DFG) Grant Making Machine Learning on Static and Dynamic 3D Data Practical, a Google Research Grant, and the Bavarian State Ministry of Science and the Arts as coordinated by the Bavarian Research Institute for Digital Transformation (bidt).
References
Appendix A Data-Efficient Learning
We plot the curves across different percentages of used training data for data-efficient learning similarly to . We demonstrate that our pre-training algorithm generalizes with different backbones. We show data-efficient learning curves on semantic segmentation task in ScanNet with a ResNet18 backbone in Figure 6. Please refer to main paper for results with a ResNet50 backbone.
To demonstrate that our pre-training algorithm generalizes well across different downstream tasks in limited data scenarios, we further plot data-efficient learning curves on the 2D object detection task on ScanNet in Figure 7, as well as data-efficient learning curves on the 2D instance segmentation task in Figure 8. We use Mask-RCNN with a ResNet50 backbone for both tasks.
Appendix B Additional Qualitative Visualizations
We additionally show more visualizations of 2D semantic segmentation on the ScanNet and NYUv2 datasets (see Figure 9). We can achieve notably improved segmentation results by using our pre-trained weights than ImageNet pre-trained weights.
Appendix C Generalization Across Datasets.
We conduct additional experiments on more diverse datasets for pre-training (ScanNet, MegaDepth , KITTI ) and downstream tasks including COCO and outdoor data (Cityscapes and KITTI) (see Table 8). We still observe a consistent improvement with our pre-training methods across different datasets and tasks. On COCO, the gap is not as significant as it in other scenarios, due to the drastic domain gap between COCO and our 3D dataset for pre-training.
Appendix D What The Contrastive Loss Learns
To demonstrate our pre-training algorithm has learned meaningful features, we show some visualizations of correspondences matching. We first use pre-trained weights to make a forward pass inference on a pair of images to get features for each pixel. For each pixel in one image, we do nearest neighbour search to find its closed corresponding pixel in anther image in feature space. We draw the matching in this way in Figure 10.
Appendix E Unsupervised Pre-training Pipeline
We experiment with network weights trained with self-supervision on ImageNet for encoder initialization. This is also a two-stage pipeline but without using any semantic labels in either stage. Even though the use of a supervised ImageNet pre-trained initialization is a common practice, for completeness we also evaluate Pri3D in an unsupervised pipeline without using ImageNet labels; we demonstrate the experimental results of Unsupervised Pri3D in Table 9. Results suggest that Pri3D does not rely on any semantic supervision (e.g. ImageNet labels) to succeed, and still is able to achieve a substantial gain in this setup. To clarify the baselines:
Unsupervised ImageNet Pre-training (MoCoV2-IN). We use MoCoV2 ImageNet pre-trained weights. No ScanNet data is involved.
2-Stage MoCoV2 (MoCoV2-unsupINSN). We start with MoCoV2-IN as the encoder initialization, but add another stage to finetune MoCoV2 with randomly shuffled ScanNet images. In this case, we use ScanNet data but no 3D priors are used.
Appendix F Depth Prediction As A Proxy Loss
We explore the influence of 3D priors on 2D tasks starting with the simplest 3D prior, i.e., depth prediction . Depth prediction indicates a positive signal but not strong enough. We then try different other stronger 3D priors, and further use depth prediction as a proxy loss by default. Moreover, we show an ablation study on the depth proxy loss in Table 10.
Appendix G Limitations
While our approach demonstrates the promising effect of learning 3D priors for 2D representation learning, there are various limitations. For instance, joint 2D and 3D pre-training, in contrast to our current 3D-based constraints only, would likely provide the most informative signal for representation learning for downstream tasks. Additionally, our current 3D-based pre-training leverages indoor scene data from ScanNet, and we would expect further generalizability by augmentation with data from other environments, such as outdoor scene data (e.g., ).