P4Contrast: Contrastive Learning with Pairs of Point-Pixel Pairs for RGB-D Scene Understanding

Yunze Liu, Li Yi, Shanghang Zhang, Qingnan Fan, Thomas Funkhouser, Hao Dong

Introduction

The goal of our work is to learn without human supervision how to extract dense features from RGB-D data that are useful for 3D scene understanding tasks such as semantic segmentation and object detection. In the field of 2D visual understanding, a popular approach to learn such useful feature representations is contrastive learning – where feature extractors are optimized to discriminate instances of a dataset via a contrastive loss in the latent space . For example, SimCLR uses constrastive learning to pretrain models for image classification. Recently, PointContrast has explored the possibility of applying self-supervised contrastive learning for high-level 3D scene understanding, with the pretext task of pulling together positive point pairs generated by transforming a point cloud with rigid transformations. The discovery is very encouraging: self-supervised pre-training on large datasets of 3D scans improves the performance of downstream 3D scene understanding tasks across different datasets.

In this paper we investigate how to make full use of contrastive learning to train dense feature extractors directly for RGB-D scans. At first glance, the problem seems quite straight forward – simply create point-pixel pairs via appending raw RGB values from 2D pixels to 3D points and use the same contrastive strategy as in PointContrast (Figure 1 (a)). Surely, contrastive learning with RGB and 3D points together should produce better features than with either modality alone. However, we find this is not the case in practice. We observe that this simple extension fails to leverage both the color and geometry input effectively. Additional RGB input in the PointContrast framework brings very marginal performance improvement, sometimes even a performance degradation.

An alternative approach is to design models to extract features separately from RGB and depth channels and train them with a cross-modal contrastive loss , as is shown in Figure 1 (b). These approaches force the models to extract discriminating features from both 2D color and 3D geometry modalities , but it does not leverage the synergies between the two modalities to extract features better than that could be learned from either alone. Indeed, in most cases, the model trained on the weaker modality regresses to the features learned by the stronger one, sometimes even making the features worse.

To address these problems, in this paper, we propose contrastive learning from RGB-D data with Pairs of Point-Pixel Pairs, as is shown in Figure 1 (c). That is, we pretrain a model, called P4Contrast, that extracts deep features of point-pixel pairs and aims to pull together positive pairs where: 1) the point and pixel within each point-pixel pair come from the same RGB-D observation, and 2) the two point-pixel pairs in a pair are in correspondence with one another. The key motivation behind this approach is that working with pairs of point-pixel pairs provides more flexibility for creating hard negatives than the previous approaches. Like PointContast, we can create negatives by picking pairs of point-pixel pairs which are not true correspondences. Like cross-modal contrastive embedding, we can create negatives by replacing the RGB or 3D point within one or both of the two point-pixel pairs. We can further create other hard negatives with combinations of these strategies.

Training a network to discriminate positive examples from all these types of negative examples encourages learning stronger features. Since the model must make a decision about whether the two components of a point-pixel pair are from the same observation, it cannot be lazy and learn strong features for only one of the two modalities. Since both components of the point-pixel pair are processed by the same network and used to discriminate correspondences between point-pixel pairs, it must learn synergistic relationships between the modalities for extracting useful features.

To evaluate this approach experimentally, we use P4Contrast to pretrain a deep backbone on a large-scale RGB-D scene dataset, ScanNet . Then we finetune the backbone on target datasets for downstream scene understanding tasks including semantic segmentation on ScanNetV2 , semantic segmentation on 3RScan and 3D object detection on SUN RGB-D . We find that P4Contrast significantly boosts the performance of previous state-of-the-art methods for all three tasks (+2.4 mIoU on ScanNetV2 semantic segmentation, +4.4 mIoU on 3RScan semantic segmentation and +2.0 mAP on SUN RGB-D 3D object detection).

Our key contributions are three-fold. First, we formulate dense RGB-D representation learning as a point-pixel pair contrastive learning problem and present a novel pretext task to encourage RGB-D information fusion. Second, we design a unified contrastive learning solution, P4Contrast, covering the pretraining objective design, the deep learning backbone construction, and data augmentation strategies. Third, we demonstrate the efficacy of P4Contrast on three large-scale RGB-D scene understanding benchmarks and we provide extensive ablation studies to validate our designs.

Related Work

RGB-D Fusion for Semantic Scene Understanding. Semantic scene understanding (SSU) involves a number of fundamental vision tasks, such as object detection , semantic segmentation , and pose estimation . To learn better visual representations for SSU, usually both color and geometry information (RGB-D) is leveraged. The specific RGB-D data representation and fusion approaches matter for either 2.5D or 3D scene understanding tasks. For 2.5D SSU, RGB and depth information are mostly formed as 2D images and encoded by popular 2D neural networks. The color and geometry features are combined through either early fusion , middle fusion , or late fusion . For 3D SSU, some recent works lift both RGB and depth into the 3D space represented as a 6-channel point cloud (RGBXYZ), which is directly processed by the 3D network backbones for joint RGB-D feature extraction . The more common approach is to represent RGB as the color image and depth as the point cloud, which are processed by 2D and 3D backbones individually first for later RGB-D feature fusion . In this paper, we aim for 3D SSU, and argue that RGB and depth information should be complementary to each other regardless of the specific data representation and the accompanying network backbones. Thus we propose to fuse color and geometry information in both 2D and 3D backbones to learn better visual representations for scene semantics.

Contrastive Self-supervised Representation Learning. Contrastive learning (CL) is a representative method of Self-supervised learning (SSL) that has gained increasing attention and demonstrated promising results . Most CL methods are instance-level, aiming to learn an embedding space where samples from the same instance are pulled closer and samples from different instances are pushed apart . Recently, Chen et al. proposed a simple framework for CL (SimCLR) with larger batch sizes and extensive data augmentation, achieving comparable results with supervised learning. SimCLR requires a large minibatch size to achieve superior performance, which is computationally prohibitive. MoCo improves the efficiency of CL by storing representations using a queue that is independent of minibatch size. More recently, CL based on prototypes has shown promising performance . While CL has been actively explored for 2D image understanding, few works have been done for 3D scene understanding. Very recently, PointContrast initiates the efforts through presenting a CL framework to learn dense point cloud features. However, it is not specifically designed for RGB-D scans. When color information is available, it simply treats the RGB values as additional features of 3D points, failing to leverage both the color and geometry input effectively.

Self-Supervised Multimodal Learning. Extensive studies on multimodal learning are dedicated to modelling the multiple modalities and their complex interactions with the aim at leveraging complementary information present in multimodal data and yielding more robust predictions . The multimodal feature fusion can be typically categorized as early , late , and hybrid fusion . Recently, several methods have been developed to learn cross-modal embedding in a self-supervised way . Tian et al. proposed multiview CL to learn a representation that aims to maximize mutual information between different views of the same scene but is otherwise compact . Mahendran et al. learned pixels embeddings to enable the similarity between their embeddings matches that between their optical flow vectors . Existing contrastive RGB-D representation learning methods mostly focus on cross-modal embedding . Their feature extractors still consume single modality input but emphasize more on the correlated feature between RGB images and 3D point clouds. Those existing works are mostly tested on single objects or for low-level tasks such as registration.

Method

In this section, we introduce our self-supervised contrastive learning pipeline. Since our work is closely related to PointContrast , we first provide a brief review in Section 3.1. To cope with the multimodal input in RGB-D scans and learn representations to better fuse color and geometry information, we then introduce our novel self-supervised pretraining solution, P4Contrast, in Section 3.2, which covers the pretext task formulation as well as the loss design. We detail the network architecture that can leverage both 2D and 3D context during feature learning in Section 3.3, we discuss our data augmentation strategy in Section 3.4 and talk about the dataset we use for pretraining in Section 3.5.

PointContrast is a contrastive learning framework that can learn descriptive dense and local geometric features on 3D point cloud. Through simple self-supervised pretraining on a large set of real 3D scenes, useful representations could be learned to boost the performance of a range of high-level 3D scene understanding tasks, such as 3D semantic segmentation and 3D object detection. A key observation is that high-level semantic scene understanding tasks require not only global but also local geometric features, making directly contrasting point cloud instances and extracting global representations insufficient. PointContrast instead contrasts pairs of points. The pretext task requires minimizing the feature distance between corresponding points from two views of a point cloud and maximizing the feature distance between unmatched points.

PointContrast uses the Sparse Residual U-Net (SR-UNet) as the network backbone that is designed for 3D point cloud processing. It makes full use of geometry information without emphasizing the role of color information in RGB-D scene understanding. In fact, under the PointContrast framework, simply treating colors as additional channels of 3D points does not bring much performance boost. The above observation motivates our work, P4Contrast, which learns to better fuse color and geometry signals in a novel contrastive learning framework.

2 Contrasting Pairs of Point-Pixel Pairs as a Pretext Task

Contrasting pairs of point-pixel pairs The common contrastive learning frameworks leverage anchor, positive and negative samples. The objective is to bring closer anchor-positive pairs and pull apart anchor-negative pairs. Following this generic idea, PointConstrast treats points as the sample representation and contrasts pairs of points. In this paper, we propose to replace the pure geometric points with point-pixel pairs as the sample representation which brings the geometric and color knowledge together, and hence contrasts pairs of point-pixel pairs.

To be specific, given an RGB-D scan (g, c), we firstly apply data transformations to obtain two versions (g1\textbf{g}^{1}, c1\textbf{c}^{1}) and (g2\textbf{g}^{2}, c2\textbf{c}^{2}) of it. This provides anchor point-pixel pairs (gi1,ci1)(\textbf{g}_{i}^{1},\textbf{c}_{i}^{1}) and positive point-pixel pairs (gi2,ci2)(\textbf{g}_{i}^{2},\textbf{c}_{i}^{2}). The negative point-pixel pairs can be naively computed via using all the non-positive ones {(gj2,cj2)∣j≠i}\{(\textbf{g}_{j}^{2},\textbf{c}_{j}^{2})|j\neq i\}. Then the goal of contrastive objective is to extract local features that support correspondence computation. However, we argue this is not an ideal solution while handling the multi-modal RGB-D inputs, since either geometric or color feature can be already strong enough to support the correspondence computation, and hence the feature extraction network could safely ignore the other information (color/geometry) without violating the objective much, which phenomenon can be observed in many different scenarios .

Disturbed point-pixel pairing The above observation indicates that a contrastive objective forcing the feature extractor to pay attention to both color and geometry input is required for better representation learning. Toward this end, we introduce disturbed point-pixel pairing to prepare the negative samples. For an anchor point-pixel pair (gi1,ci1)(\textbf{g}_{i}^{1},\textbf{c}_{i}^{1}), in addition to its unmatched point-pixel pairs {(gj2,cj2)∣j≠i}\{(\textbf{g}_{j}^{2},\textbf{c}_{j}^{2})|j\neq i\}, we further break the correspondences between points and pixels to generate a new set of negative samples {(gj2,cd(j)2)}\{(\textbf{g}_{j}^{2},\textbf{c}_{d(j)}^{2})\}, where d(⋅)d(\cdot) is a disturbing function satisfying d(j)≠jd(j)\neq j. In this case, both the correspondence between geometric positions of two transformed scenes and the binding between geometric and color features are broken to form our negative point-pixel pair. Notice we do not exclude partially negative pairs (gi2,cd(i)2)(\textbf{g}_{i}^{2},\textbf{c}_{d(i)}^{2}) from the negative set, meaning a point-pixel pair is only positive when both the point and pixel parts are positive. This strategy forces the underlying network to extract meaningful features from both geometry and color information. We train the feature extractor through optimizing a contrastive loss over the point-pixel pairs. To be specific, we minimize the distance between anchor and positive point-pixel pairs while maximizing the distance between anchor and negative point-pixel pairs. We visualize this idea in Figure 2.

Hardness of partially negative point-pixel pairs Disturbing point-pixel pairing breaks the correspondences between points and pixels and introduces negative samples that could force the feature extractor to jointly depict geometry and color. Among the disturbed pairs, we find the hardness of partially negative pairs playing a big role on the representation learning process. Here we consider an anchor pair (gi1,ci1)(\textbf{g}_{i}^{1},\textbf{c}_{i}^{1}) corresponding to point gi1\textbf{g}_{i}^{1} and a negative pair (gi2,cd(i)2)(\textbf{g}_{i}^{2},\textbf{c}_{d(i)}^{2}) whose point part is positive while the pixel part corresponds to another point gd(i)1\textbf{g}_{d(i)}^{1}. In order to discriminate (gi1,ci1)(\textbf{g}_{i}^{1},\textbf{c}_{i}^{1}) from (gi2,cd(i)2)(\textbf{g}_{i}^{2},\textbf{c}_{d(i)}^{2}), the feature representation for the anchor pair can not ignore the color information ci1\textbf{c}_{i}^{1}.

Due to spatial smoothness, the difficulty of discriminating (gi1,ci1)(\textbf{g}_{i}^{1},\textbf{c}_{i}^{1}) from (gi2,cd(i)2)(\textbf{g}_{i}^{2},\textbf{c}_{d(i)}^{2}) usually increases as point gi1\textbf{g}_{i}^{1} and gd(i)1\textbf{g}_{d(i)}^{1} get closer in the 3D space. Therefore we intuitively define the hardness of a partially negative point-pixel pair (gi2,cd(i)2)(\textbf{g}_{i}^{2},\textbf{c}_{d(i)}^{2}) regarding an anchor pair (gi1,ci1)(\textbf{g}_{i}^{1},\textbf{c}_{i}^{1}) as 1∣∣gi1−gd(i)1∣∣2\frac{1}{||\textbf{g}_{i}^{1}-\textbf{g}_{d(i)}^{1}||_{2}}. The partially negative pairs are especially challenging yet important negative samples. Making them too hard, say re-pairing a point with the colored pixel from a nearby point, could confuse the feature extractor. On the other hand, the point-pixel incompatibility within a very easy pair may soften the impact on the feature extractor to discriminate the color and geometry information, therefore contributing less to the representation learning process. These require a strategy properly controlling the hardness of partially negative pairs during learning.

Progressive hardness increasing We adopt a strategy of progressive hardness increasing while preparing the partially negative pairs for contrastive learning. To be specific, at training iteration kk, we sample a random index jj satisfying 1∣∣gi1−gj1∣∣2>clip(h(k),ϵ)\frac{1}{||\textbf{g}_{i}^{1}-\textbf{g}_{j}^{1}||_{2}}>\text{clip}(h(k),\epsilon) and j≠ij\neq i, where h(⋅)h(\cdot) is a linearly growing function, clip(⋅,ϵ)\text{clip}(\cdot,\epsilon) clips an input value making sure the output is smaller than ϵ\epsilon. We re-pair gi2\textbf{g}_{i}^{2} with cj2\textbf{c}_{j}^{2} as a partially negative pair of the anchor pair (gi1,ci1)(\textbf{g}_{i}^{1},\textbf{c}_{i}^{1}). Intuitively speaking, the lower bound for the hardness of partially negative pairs is gradually increasing along the learning process. We visualize the partially negative pairs with different hardness as colored point clouds in Figure 3. It can be seen that partially negative pairs with higher hardness have a higher correlation with the anchor point-pixel pairs, increasing the difficulty of discriminating anchor point-pixel pairs.

Loss design Inspired by the PointInfoNCE loss proposed in , we use a PairInfoNCE loss to contrast different point-pixel pairs. The loss encourages an anchor point-pixel pair to be similar to its positive pair and dissimilar to many negative pairs. Specifically, our loss Lc\mathcal{L}_{c} is defined as follows:

where fijv\textbf{f}_{ij}^{v} denotes the extracted feature of point-pixel pair (giv,cjv),v=1,2(\textbf{g}_{i}^{v},\textbf{c}_{j}^{v}),v=1,2 and d(⋅)d(\cdot) is the disturbing function breaking the correspondences between points and pixels. Notice the negative samples contain both disturbed point-pixel pairs and undisturbed ones. d(⋅)d(\cdot) is designed to increase the hardness of partially negative pairs progressively during training.

3 Learning Backbone Capturing 2D-3D Context

To learn the dense RGB-D feature representations, a deep learning backbone is required to consume the input point-pixel pairs (gi\textbf{g}_{i}, ci\textbf{c}_{i}). Note the point and pixel is the representation of geometry and color information, which can be brought into the same domain in the form of either 3D colored point cloud or 2D RGB-D image through camera projection. Therefore the point-pixel pair (gi\textbf{g}_{i}, ci\textbf{c}_{i}) can be encoded by both the popular 2D and 3D convolutional network backbones. 2D convolutional backbones leverage 2D context where the distance of point-pixel pairs is determined by the color pixel ci\textbf{c}_{i} while in 3D convolutional backbones the distance of point-pixel pairs is computed using the geometric point gi\textbf{g}_{i}. The difference of the underlying metric spaces influences how features get aggregated. We conject that 2D and 3D context should be complementary to each other and a good deep learning backbone should not restrict itself to either one. We therefore present our 2D-3D context contrastive learning backbone, which combines a 3D and a 2D backbone as is shown in Figure 4.

Contrasting with 3D context We use a Sparse Residual U-Net (SR-UNet) as our 3D backbone which is designed in and also leveraged by PointContrast. The backbone treats a 3D point cloud as a set of sparse voxels and leverages sparse 3D convolutions in a 34-layer U-Net architecture . Each input point is equipped with a 6-channel feature including its 3D location and RGB color. The SR-UNet consumes the 6-channel input points and densely generates per-point feature representations.

Contrasting with 2D context We use FuseNet as our 2D backbone, which is originally designed for RGB-D image semantic segmentation. FuseNet contains a RGB network and a depth network, both with 2D convolutional encoder-decoder structures. The RGB network and the depth network first extract their own features at the front end of FuseNet. Then sparse fusion is used to fuse their features and the output is a dense feature map.

Contrasting with 2D-3D context Our overall backbone simply combines the 3D SR-UNet and the 2D FuseNet by concatenating their corresponding output features. Notice our backbone combines both early fusion happening inside 2D/3D backbones, and late fusion merging the outputs of 2D and 3D backbones, therefore could be categorized as a hybrid fusion backbone.

4 Data augmentation

Data augmentation is an important component in contrastive learning which allows generating anchor, positive and negative examples and extracting representations invariant to certain noise in the input. As contrastive learning becomes more and more popular, what are suitable data augmentations has also attracted many interests , mostly restricted to 2D images though. PointContrast as the first work leveraging contrastive learning for dense and local 3D feature learning, proposes a 3D data augmentation strategy containing two steps: rendering a scene point cloud into a depth view and applying random rigid transformations plus scaling afterwards. We discover through experiments that this augmentation strategy is quite vulnerable to mode collapsing (features collapsed in space) with the SR-UNet backbone as is shown in Figure 5. This will restrict the benefits of pretraining for downstream semantic understanding tasks. After experimenting with different 3D data augmentation strategies, we end up using only point jittering to generate different versions of a scene point cloud. Though losing rotation and scale invariance, we find this allows very robust training against mode collapsing, producing more uniformly distributed features which has been argued to be beneficial to downstream tasks previously . As for RGB image views, we add Gaussian noise to an RGB image to generate different versions of it , which allows us to easily maintain the point-pixel correspondences.

5 Dataset for Pretraining

Following PointContrast, we use ScanNet as the dataset for pretraining. ScanNet is a large-scale indoor scene dataset containing ~1500 RGB-D scans. Each scan contains a reconstructed point cloud and a sequence of RGB images with the ground truth point-pixel correspondences known. We firstly pretrain the SR-UNet and FuseNet backbone separately and then fuse their output features together. While training SR-UNet, we use colored point cloud to represent a ScanNet scene. As described in Section 3.4, we jitter point positions and colors with Gaussian noise to generate anchor, positive and negative examples. While training FuseNet, we sample paired RGB and depth images from a scene. Again we augment RGB and depth images with Gaussian noise.

6 Implementation Details

For ScanNet and 3RScan semantic segmentation task, we train the model with one V100 GPU for 60,000 iterations. Batch size is 16. We use SGD+momentum optimizer with an initial learning rate 0.8. We use polynomial learning rate scheduler with a power factor of 0.9. We set the weight decay as 0.0001 and the voxel size is 2.5cm. Above hyper parameters are mostly consistent with PointContrast.

In the experiment of SUN RGB-D 3D object detection, our settings are consistent with ImVoteNet. We train the model on one 2080Ti GPU for 180 epochs. The initial learning rate is 0.001 and the tower weights are 0.3,0.3,0.4 respectively. We sample 20,000 points from each scene and the voxel size is 5cm.

Experiment

Following PointContrast, we adopt a supervised fine-tuning strategy to evaluate how well the learned representations could transfer to downstream tasks. Specifically, we first pretrain our 2D-3D backbone on ScanNet dataset with the proposed pretraining objective and data augmentation strategies. Then we use the pretrained weights as initialization and further refine them for target downstream tasks. The performance gain will be a good indicator measuring the quality of learned features. In this section, we cover three popular RGB-D scene understanding tasks: semantic segmentation on ScanNetV2 , 3D object detection on SUN RGB-D and semantic segmentation on 3RScan in Section 4.1, 4.2 and 4.3 respectively. We also demonstrate the benefits of representation learning when the training set is small in Section 4.4. In addition, we provide extensive ablation studies to validate our design choices in Section 4.5.

We first conduct experiments to see whether self-supervised pretraining on ScanNet dataset can help semantic scene understanding on its own. With P4Contrast we can explore the synergetic signals between color and geometry and learn representations robust to noise and potentially different from what obtained with a direct supervised training. Therefore, there is a good reason to believe our self-supervised pretraining should improve semantic segmentation even when the source dataset for pretraining and target dataset for evaluation are the same. We follow the official data split of ScanNet with 1,201 training scenes and 312 validation scenes. We pretrain on both the training and validation set and conduct supervised finetuning using only the training set. We evaluate the semantic segmentation on the validation set covering all the 20 semantic categories considered in the official benchmark. Mean IoU(mIoU) is used as the evaluation metric. We compare our method with baselines trained from scratch without pretraining, as well as PointContrast which treats point-pixel pairs as individual “points”. To better understand the contribution from the pretraining objective, we factorize the backbone difference and provide a variation of our approach using the SR-UNet 3D backbone only (P4Contrast(3D Context)). We compare our full approach, P4Contrast(2D-3D Context), with all above methods and report the results in Table 1. Notice we also highlight the input leveraged by different methods where “Geo” denotes point cloud only and “Geo+RGB” adds 2D color images.

We notice in , comparisons with baseline methods on ScanNetV2 semantic segmentation task do not factorize the backbone difference. In , the from scratch training baseline leverages a SR-UNet with 3D convolution kernel of size 5 while PointContrast leverages a SR-UNet with 3D convolution kernel of size 3, making the comparisons unfair. To factorize the influence from backbones as well as some other hyperparameters, we use the officially released code from PointContrast, build the 3D backbone of P4Contrast upon that while keeping the hyperparameters unchanged as much as possible. In addition, for a more complete and fair comparison, we re-run the PointContrast code as well as the from scratch training baseline using backbones with 3D convolution kernels of both size 3 and 5. We use PointContrast1 to denote results from re-running the officially released code. Possibly due to some hyperparameter change, we find the released code of PointContrast fails to reproduce results reported in the paper. Since our framework is largely built upon the released version PointContrast1, we can still fairly compare our results with PointContrast1.

We have several discoveries from Table 1. Using backbones with 3D convolution kernels of size 5 as an example, PointContrast has improved the segmentation mIoU by 1.1%1.1\% over the from-scratch-training baseline when only point cloud is used as input. When color information is also available, the improvement drops to 0.5%0.5\%. Appending color values as additional features to 3D points does not introduce much benefit in the PointContrast framework, with only a 0.3%0.3\% mIoU improvement. The comparison between P4Contrast(3D context) and PointContrast1 directly validate the efficacy of contrasting pairs of point-pixel pairs v.s contrasting “points” containing color features since they use the same SR-UNet backbone. The improvement from 72.7%72.7\% to 73.6%73.6\% confirms that with our novel pretraining objective design, we can better leverage the color information to boost the segmentation performance on RGB-D scenes. With our full approach whose backbone leverages both 2D and 3D context information, we can achieve a 2.4%2.4\% mIoU improvement over the baseline trained from scratch. Similar observations apply to the setting using backbones with 3D convolution kernels of size 3, where the segmentation mIoU is 0.4%−1.1%0.4\%-1.1\% higher than with kernel size 5.

2 Fine-tuning on SUN RGB-D 3D object detection

We then focus on another popular RGB-D scene understanding task, 3D object detection. We also switch to a different target dataset, SUN RGB-D , to test how pretraining on ScanNet scenes could benefit high-level understanding of single-view RGB-D scans. SUN RGB-D contains a training set with ~5K single-view RGB-D scans and a test set with ~5K scans. The scans are annotated with amodal 3D oriented bounding boxes for objects from 37 categories. Following the standard protocol , we only evaluate and report the results on 10 most common object categories. Our 2D-3D backbone supports pixel/point-wise label prediction but does not directly output 3D bounding boxes, therefore we need to modify our backbone for detection. In analogy to PointContrast modifying and finetuning on VoteNet , we adapt our backbone based on the current state of the art RGB-D 3D object detection architecture, ImVoteNet . ImVoteNet backbone leverages PointNet++ to process depth point cloud and uses Faster R-CNN pretrained on COCO train2017train2017 dataset to process input images, which aligns well with our 3D convolutional backbone and 2D convolutional backbone. To also leverage the power of supervised pretraining from COCO, we only pretrain the PointNet++ backbone on ScanNet in this case but with both RGB and point cloud inputs using the objective from P4Contrast. We compare with several baseline methods including VoteNet , PointContrast and ImVoteNet in Table 2. We use mAP@0.25 as our evaluation metric.

It is worth mentioning that we use the official code from the authors of ImVoteNet as our detection backbone. After re-running the code, we observe a 1.9%1.9\% mAP drop compared with the numbers reported in . We confirm with the authors and find this is due to changes to the codebase and hyperparameters in the released version. For a fair comparison, we report these numbers as ImVoteNet2 in Table 2 since our method pretrains ImVoteNet2 with P4Contrast. VoteNet and PointContrast do not fully explore how to leverage images to boost the performance of RGB-D 3D object detection, their performance is outperformed by ImVoteNet2 by an obvious margin. Using PointContrast to fine-tune ImVoteNet, we observe very limited improvement from 61.5 to 61.8 which still lagsbehind our proposed P4Contrast by a significant margin. After pretrained with P4Contrast on ScanNet, we obtain a 2.0%2.0\% gain on mAP@0.25. This again validates that P4Contrast could encourage effective fusion of color and geometry signals and our pretrained representations generalizes well across datasets.

3 Fine-tuning on 3RScan Semantic Segmentation

To demonstrate the efficacy of P4Contrast on more benchmarks, we conduct another supervised fine-tuning experiment on 3RScan dataset focusing on the semantic segmentation task. 3RScan is a large-scale real indoor scene dataset which features 1482 3D reconstructions of 478 naturally changing indoor environments, containing a training set with 385 scans and a test set with 47 scans. In comparison to scans from ScanNet which are captured with a Structure Sensor, scans in 3RScan are captured with Google Tango with different noise patterns and scan qualities. Such differences bring extra challenges while transferring the pretrained representation. As we will show, P4Contrast still does a good job to improve the downstream 3RScan semantic segmentation task. We evaluate and report the results on the predefined 27 categories in 3RScan dataset, and adopt mean IoU (mIoU) as the evaluation metric. The numerical results are reported in Table 3.

Compared to PointContrast, P4Contrast (3D context) improves the mIoU from 38.8 to 40.8. As these two approaches share the same network backbone and only differs in their strategy to process the input color and geometry information, the improvement validates the effectiveness of contrasting pairs of point-pixel pairs. After applying the full pipeline of P4Contrast, our performance is further upgraded to 41.7, and achieves 4.4 mIoU improvement over the train-from-scratch baseline.

4 Fine-tuning on Small Training Set

Representation learning has exhibited powerful transfer capability that enables pretraining on a large-scale dataset and fine-tuning on a different target set. Good representations should be able to bring even larger performance gains when the target fine-tuning dataset is small in scale. To validate this, we reduce the amount of labeled data while fine-tuning for ScanNet semantic segmentation task. Different from Section 4.1, we only finetune the pre-trained network on 10% of the training set but still test on the whole validation set. Again, we compare with the train from scratch baseline as well as the PointContrast method using mIoU as the evaluation metric. We show the experimental results in Table 4. As expected, the full pipeline of P4Contrast is still able to yield better performance over all the other competitors. Moreover, we observe an mIoU improvement of 4.5 over the train from scratch baseline when only using 10% of the training set, which almost doubles 2.4 mIoU improvement when using 100% of the training set.

5 Analysis Experiments and Discussions

In this section, we provide more analysis to provide an in-depth understanding of our framework. We use ScanNetV2 semantic segmentation as the target downstream task throughout the section and mIoU is used as the evaluation metric. We use SR-UNet as our 3D convolutional backbone and FuseNet as our 2D convolutional backbone. The 3D convolution kernels used in SR-UNet is of size 5.

Multi-context or single-context PointContrast restrict itself to a 3D convolutional backbone which can only capture 3D context while loses the dense 2D patterns in RGB-D input. We instead propose to contrast with both 2D and 3D context through a backbone combining 2D convolution from FuseNet and 3D convolution from SR-UNet. To validate this design choice, we experiment with three different backbones: FuseNet with just 2D context; SR-UNet with just 3D context; and FuseNet+SR-UNet combining 2D and 3D context. In addition to P4Contrast, we also adapt PointContrast by changing the backbone it uses. In all the experiments we feed both RGB and point cloud inputs to the backbone (colored point cloud for 3D convolution and RGB-D image for 2D convolution). The results are shown in Table 5. For both PointContrast-like framework and P4Contrast, using 2D-3D context has the best segmentation performance. P4Contrast outperforms PointContrast-like framework consistently by over 1%1\% mIoU using all three backbones, indicating its better capability of fusing color and geometry input.

The influence of pair hardness We introduce disturbed point-pixel pairing in P4Contrast, which constructs partially negative point-pixel pairs. These challenging negative pairs encourage the learned representation to focus on both geometry and color inputs. However, directly pretraining the backbone with very hard partially negative pairs could confuse the network and get training stuck. On the other hand, always training with easy partially negative pairs will not be able to squeeze the power of disturbed point-pixel pairing fully. To demonstrate this, we experiment with three different hardness scheduling strategies. One is our proposed strategy, progressive hardness increasing, which raises the hardness lower bound of partially negative pairs following a clipped linear function clip(h(k),ϵ)\text{clip}(h(k),\epsilon) as training iteration number kk increases, where clip(⋅,ϵ)\text{clip}(\cdot,\epsilon) sets any values larger than ϵ\epsilon to ϵ\epsilon. Another is to always use h(0)h(0) as the hardness lower bound. The last is to always use ϵ\epsilon as the hardness lower bound. We use “Progressive”, “Easy” and “Hard” to represent these three strategies respectively in Table 6. As can be seen, our progressive hardness increasing strategy leads to the best performance.

What data augmentation to use We have mentioned that the data augmentation strategy leveraged in PointContrast is quite vulnerable to mode collapsing in our experiment as is shown in Figure 5. That augmentation strategy mainly consists of random rigid transformation to the input point cloud. To figure out the best augmentation strategy for P4Contrast, we conduct a range of controlled experiments. Our base strategy transforms an input RGB-D scan by adding Gaussian noise to both the color image and the 3D point cloud. We test different variations including: no RGB image augmentation, add point cloud rotation augmentation, add point cloud translation augmentation, random point cloud scaling, add point cloud flipping augmentation and multi-view rendering. The results are shown in Table 7. We find our RGB image augmentation is quite helpful and removing it will hurt the performance. In addition, adding either rotation, scaling, translation or flipping to augment point clouds will hurt the representation learning process. Using the multi-view rendering strategy, we achieves a mIoU of 72.8, which is still not as good as our data augmentation strategy. These validate that our simple data augmentation strategy which involves only data jittering is quite effective empirically. It would be an interesting direction to further explore the best data augmentation strategy as well as the underlying principle for specific downstream tasks in the future.

Cross-modal embedding v.s. P4Contrast A popular representation learning framework for RGB-D input is to extract features separately from RGB and depth channels and train them with a cross-modal contrasive loss . However as we mentioned previously, such approaches usually fail to leverage the synergies between color and geometry inputs to extract features better than using a single modality alone. To demonstrate this, we design a point and pixel level contrastive learning framework using cross-modal contrastive loss and compare with our P4Contrast. To be specific, we use SR-UNet to consume an input point cloud and extract per-point representations. We use FuseNet without the depth fusion layers to consume an input RGB image and extract per-pixel representations. Contrastive loss is then applied between point and pixel representations to bring together corresponding point-pixel pairs and push apart unmatched point-pixel pairs. After pretraining, we concatenate the representations of corresponding point-pixel pairs for downstream tasks. On the ScanNetV2 semantic segmentation task, this cross-modal embedding framework achieves 72.4%72.4\% mIoU, the same as PointContrast1 in Table 1 using only point cloud input and much worse than P4Contrast with a 2.2%2.2\% mIoU drop. This again justifies our design of P4Contrast and confirms the deficiency of cross-modal embedding as a proper RGB-D representation learning framework.

Fusion architecture study With our backbone, we essentially fuse color and geometry signals multiple times. We early fuse color and geometry by combining RGB values and point coordinates together at the input of SR-UNet. We also have a feature level fusion between color and geometry inside FuseNet. By concatenating the output of FuseNet and SR-UNet, we conduct another late fusion. Therefore our backbone can be treated as a hybrid fusion architecture. We compare it with two naive solutions. One is an early fusion strategy with just SR-UNet as the backbone consuming 3D point clouds equipped with color features. The other is a late fusion strategy where we first learn per-point representations from 3D point cloud and learn per-pixel representations from RGB images separately using PointContrast-like frameworks. And then we concatenate the features of corresponding point-pixel pairs as a fused feature representation. To learn per-pixel representations from just RGB images, we simply adapt the FuseNet backbone by removing the depth-fusion layers. Late fusion can not model the cross modal correlations very well and thus only achieves 72.8%72.8\% mIoU on ScanNetV2, with just 0.6%0.6\% mIoU improvement compared with the baseline trained from scratch. Jointly considering both color and geometry inputs and contrasting pairs of point-pixel pairs with an early fusion strategy already improves the learned representations a lot and achieves 73.6%73.6\% mIoU. With the hybrid fusion strategy, we can learn the best multi-modal feature representation, achieving 74.6%74.6\% mIoU.

Conclusion

This paper proposes constrasting “pairs of point-pixel pairs” as a new method for self-supervised representation learning for points in RGB-D scans. The key idea is to train using hard negatives with disturbed correspondences between RGB and 3D points within the same RGB-D observation, as well as between different observations. This approach encourages the network to learn a representation that encodes salient features of the RGB, 3D point cloud, and their synergies, which leads to pretrained features that outperform others on two RGB-D scene understanding benchmarks. This result is very encouraging, and suggests future work investigating its applications to other multi-modal domains.

Acknowledgement

We would like to thank Haoqi Yuan, Zihan Jia, and Mingdong Wu for their help in establishing the initial foundation for the project. This work was supported by the Key-Area Research and Development Program of Guangdong Province(2019B121204008) and the Center on Frontiers of Computing Studies (7100602567).

References