Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection
Vibashan VS, Poojan Oza, Vishal M. Patel
Introduction
In recent years, object detection has seen tremendous advancements due to the rise of deep networks . The major contributor to this success is the availability of large-scale annotated detection datasets , as it enables the supervised training of deep object detector models. However, these models often have poor generalization when deployed in visual domains not encountered during training. In such cases, most works in the literature follow the Unsupervised Domain Adaptation (UDA) setting to improve generalization . Specifically, UDA methods aim to minimize the domain discrepancy by aligning the feature distribution of the detector model between source and target domain . To perform feature alignment, UDA methods require simultaneous access to the labeled source and unlabeled target data. However in practical scenarios, the access to source data is often restricted due to concerns related to privacy/safety, data transmission, data proprietary etc. For example, consider a detection model trained on large-scale source data, that performs poorly when deployed in new devices having data with different visual domains. In such cases, it is far more efficient to transmit the source-trained detector model (500-1000MB) for adaptation rather than transmitting the source data (10-100GB) to these new devices . Moreover, transmitting only source-trained model alleviates many privacy/safety, data proprietary concerns as well . Hence, adapting the source-trained model to the target domain without having access to source data is essential in the case of practical deployment of detection models. This motivates us to study Source-Free Domain Adaptation (SFDA) setting for adapting object detectors (illustrated in Fig. 1).
The SFDA is a more challenging setting than UDA. Specifically, on top of having no labels for the target data, the source data is not accessible during adaptation. Therefore, most SFDA methods for detection consider training with pseudo-labels generated by source-trained model . However, these pseudo-labels are noisy due to domain shift and the training on noisy pseudo-label is sub-optimal solution . In order to avoid such a scenario, we need to consider two challenges of the SFDA training, 1) Effectively distill target domain information into source-trained model and 2) Enhancing the target domain feature representations.
A critical challenge is improving the features of the target domain data. Consider Fig. 2, which shows object proposals for an image from FoggyCityscapes , predicted by a detector model trained on Cityscapes . Here, all the proposals have Intersection-over-Union0.9 with respective ground-truth bounding boxes and each proposal is assigned a prediction with a confidence score. Noticeably, the proposals around the bus instance have different predictions, e.g., car with 18, truck with 93, and bus with 29 confidence. This indicates that the pooled features are prone to a large shift in the feature space for a small shift in the proposal location. This is because, the source-trained model representations would tend to be biassed towards source data, resulting in weak representation for the target data. To this end, we utilize the Contrastive Representation Learning (CRL) framework and design a novel contrastive loss to enhance the feature representations of the target domain.
CRL has been shown to learn high-quality representations from images in an unsupervised manner . CRL methods achieve this by forcing representations to be similar under multiple views (or augmentations) of an anchor image and dissimilar to all other images. In classification, the CRL methods assume that each image contains only one object. On the contrary, for object detection, each image is highly likely to have multiple object instances. Furthermore, the CRL training also requires large batch sizes and multiple views to learn high-quality representations, which incurs a very high GPU/memory cost, as detection models are computationally expensive. To circumvent these issues, we propose an alternative strategy which exploits the architecture of the detection model like Faster-RCNN . Interestingly, the proposals generated by the Region Proposal Network (RPN) of a Faster-RCNN essentially provide multiple views for any object instance as shown in Fig. 3 (a). In other words, the RPN module provides instance augmentation for free, which could be exploited for CRL, as shown in Fig. 3 (b). However, RPN predictions are class agnostic and without the ground-truth annotations for target domain, it is impossible to know which of these proposals would form positive (same class)/negative pairs (different class), which is essential for CRL. To this end, we propose a Graph Convolution Network (GCN) based network that models the inter-instance relations for generated RPN proposals. Specifically, each node corresponds to a proposal and the edges represent the similarity relations between the proposals. This learned similarity relations are utilized to extract information regarding which proposals would form positive/negative pairs and are used to guide CRL. By doing so, we show that such graph-guided contrastive representation learning is able to enhance representations for the target data.
Our contributions are summarized as follows:
We investigate the problem of source-free domain adaptation for object detection and identify some of the major challenges that need to be addressed.
We introduced an Instance Relation Graph (IRG) framework to model the relationship between proposals generated by the region proposal network.
We propose a novel contrastive loss which is guided by the IRG network to improve the representations for the target data.
The effectiveness of the proposed method is evaluated on multiple object detection benchmarks comprising of visually distinct domains. Our method outperforms existing source-free domain adaptation methods and many unsupervised domain adaptation methods.
Related works
Unsupervised Domain Adaption. Unsupervised domain adaptation for object detection was first explored by Chen et al. . Chen et al. proposed adversarial-based feature alignment for a Faster-RCNN network at image and instance level to mitigate the domain shift. Later, Saito et al. proposed a method that performs strong local feature alignment and weak global feature alignment based on adversarial training. Instead of utilizing an adversarial-based approach, Khodabandeh et al. proposed to mitigate domain shift by pseudo-label self-training on the target data. Self-training using pseudo-labels ensures that the detection model learns target representation. Later, Kim et al. proposed an image-to-image generation based adaptation strategy where given source and target domain, the proposed method generates target like source images. The generated target-like images are then used to train the detection model; as a result, the detection network learn target features. Recently, Hsu et al. explored domain adaptation for one-stage object detection, where he utilized a one-stage object detection framework to perform object center-aware training while performing adversarial feature alignment. There exists multiple UDA work for object detection ; however, all these works assume you have access to labeled source and unlabeled target data.
Source-Free Domain Adaptation. In a real-world scenario, the source data is not often accessible during the adaptation process due to privacy regulations, data transmission constraints, or proprietary data concerns. Many works have addressed the source-free domain adaptation (SFDA) setting for classification , 2D and 3D object detection and video segmentation tasks. First for the classification task, the SFDA setting was explored by Liang et al. proposed source hypothesis transfer, where the source-trained model classifier is kept frozen and target generated features are aligned via pseudo-label training and information maximization. Following the segmentation task Liu et al. proposed a self-supervision and knowledge transfer-based adaptation strategy for target domain adaptation. For object detection task, proposed a pseudo-label self-training strategy and proposed self-supervised feature representation learning via previous models approach.
Contrastive Learning. The huge success in unsupervised feature learning is due to contrastive learning which has attributed to huge improvement in many unsupervised tasks . Contrastive learning generally learns a discriminative feature embedding by maximizing the agreement between positive pairs and minimizing the agreement with negative pairs. In . in batch of an image, an anchor image undergoes different augmentation and these augmentations for that anchor forms positive pair and negative pairs are sampled from other images in the given batch. Later, in exploiting the task-specific semantic information, intra-class features embedding is pulled together and repelled away from cross-class feature embedding. In this way, learned a more class discriminative feature representation. All these works are performed for the classification task, and these methods work well for large batch size tasks . Extending this to object detection tasks generally fails as detection models are computationally expensive. To overcome this, we exploit graph convolution networks to guide contrastive learning for object detection.
Graph Convolution Neural Networks (GNNs). Graph Convolution Neural Networks was first introduced by Gori to process the data with a graph structure using neural networks. The key idea is to construct a graph with nodes and edges relating to each other and update node/edge features, i.e., a process called node feature aggregation. In recent years, different GNNs have been proposed (e.g., GraphConv , GCN , each with a unique feature aggregation rule which is shown to be effective on various tasks. Recent works in image captioning , scene graph parsing etc. try to model inter-instance relations by IoU based graph generation. For these applications, IoU based graph is effective as modelling the interaction between objects is essential and can be achieved by simply constructing a graph based on object overlap. However, the problem araises with IoU based graph generation when two objects have no overlap and in these cases, it disregards the object relation. For example, see Fig. 3 (a), where the proposals for the left sidecar and right sidecar has no overlap; as a result, IoU based graph will output no relation between them. In contrast for the CRL case, they need to be treated as a positive pair. To overcome these issues, we propose a learnable graph convolution network to models inter-instance relations present within an image.
Proposed method
Background. UDA considers labeled source and unlabeled target domain datasets for adaptation. Let us formally denote the labeled source domain dataset as , where denotes the source image and denotes the corresponding ground-truth, and the unlabeled target domain dataset as, , where denotes the target image without the ground-truth annotations. In contrast, the SFDA setting considers a more practical scenario where the access to the source dataset is restricted and only a source-trained model and the unlabeled target data are available during adaptation.
Mean-teacher based self-training. Self-training adaptation strategy updates the model on unlabeled target data using pseudo labels generated by the source-trained model. The pseudo labels are filtered through confidence threshold and the reliable ones are used to supervise the detector training . More formally, the pseudo label supervision loss for the object detection model can be given as:
To this end, we utilize mean-teacher which consists of student and teacher networks with parameters and , respectively. In the mean-teacher, the student is trained with pseudo labels generated by the teacher and the teacher is progressively updated via Exponential Moving Average (EMA) of student weights. Furthermore, motivated by semi-supervised techniques , the student and teacher networks are fed with strong and weak augmentations, respectively and consistency between their predictions improves detection on target data. Hence, the overall student-teacher self-training based object detection framework updates can be formulated as:
where is the student loss computed using the pseudo-labels generated by the teacher network. The hyperparameters and are student learning rate and teacher EMA rate, respectively. Although the student-teacher framework enables knowledge distillation with noisy pseudo-labels, it is not sufficient to learn high-quality target features, as discussed earlier. Hence, to enhance the features in the target domain, we utilize contrastive representation learning.
Contrastive Representation Learning (CRL). SimCLR is a commonly used CRL framework, which learns representations for an image by maximizing agreement between differently augmented views of the same sample via a contrastive loss. More formally, given an anchor image , the SimCLR loss can be written as:
where is the batch size, and are the features of two different augmentations of the same sample , whereas represents the feature of batch sample , where . Also, indicates a similarity function, e.g. cosine similarity. Note that, in general the CRL framework assumes that each image contains one category . Moreover, it requires large batch sizes that could provide multiple positive/negative pairs for the training .
2 Graph-guided contrastive learning
To overcome the challenges discussed earlier, we exploit the architecture of Faster-RCNN to design a novel contrastive learning strategy as shown in Fig. 4. As we discussed in Sec. 1, RPN by default, provides augmentation for each instance in an image. As shown in Fig. 3, cropping out the RPN proposals will provide multiple different views around each instance in an image. This property can be exploited to learn contrastive representation by maximizing the agreement between proposal features for the same instance and disagreement of the proposal features for different instances. However, RPN predictions are class agnostic and the unavailability of ground truth boxes for target domain makes it difficult to know which proposals belong to which instance. Consequently, for a given proposal as an anchor, sampling positive/ negative pairs become a challenging task. To this end, we introduce an Instance Relation Graph (IRG) network that models inter-instance relations between the RPN proposals. IRG then provides pairwise labels by inspecting similarities between two proposals to identify positive/negative proposal pairs.
Graph Convolution Network (GCN) is an effective way to understand the relationship and propagate information between the nodes . The proposed IRG network utilizes GCN to learn the relationship between the RPN proposals. Let us denote IRG as , where is nodes and is edges of the graph network. The nodes in corresponds to RoI features extracted from RPN proposals and edges encodes relationship between the and the proposals. We then aim to learn relation matrix , to find the relationship between the RPN proposals. Both the student and teacher networks share the IRG network for modeling relationships between object proposals.
Nodes. The nodes in IRG represent features of the RPN proposals obtained from RoI feature extractor. The nodes in are denoted as , where is the feature of the instance. Here, is the total number of RPN proposals. We set to 300 for both teacher and student. The teacher pipeline has input with weak augmentations; thus, the teacher RPN proposals are better and more consistent than strongly augmented student RPN. Hence, we use teacher RPN proposals to extract RoI features and construct IRG for both student and teacher networks.
Edges. The edges in the graph are denoted as , where is the edge of the and nodes, denoting the relation of corresponding instances in the feature space and can be formally represented as:
where, and are learnable function helps to model relation between nodes.
where denotes the Kullback–Leibler divergence, denotes softmax operator. Therefore, minimizing supervises the IRG network which inturn learns the instance relation matrix ().
2.2 Graph Contrastive Loss (GCL)
Instance pairwise labels. In order to utilize the contrastive loss, we need to understand the relation of the given anchor proposal with other RPN proposals to form positive/negative pairs. As mentioned earlier, this relation matrix () is obtained from the IRG network, which learns how proposals are related to each other. For instance pairwise label generation, let us consider proposal instances and and it’s corresponding learned relation between them, . Now, one can obtain positive/negative pairs by simply setting a threshold on normalized where the would indicate that they are highly related, forming a positive pair and vice versa for the negative pairs. The pairwise labels between and proposal instances, denoted as , can be given as:
where is a hyper parameter. Thus, for a given anchor proposal we obtain its corresponding positive and negative proposal pairs from .
Instance pairwise logits. As shown in Fig. 5, the RoI features are projected as key and query inorder to model better correlation among the RoI features. For given RoI features, we obtain key, query and pairwise logits as follows:
where and are linear layer weights and , and are key, query and instance pairwise logits. To this end, the contrastive loss can be computed from the instance pairwise logits () and instance pairwise labels ().
Contrastive loss. Considering any proposal as an anchor, where , let us define a set consisting of all the samples excluding the anchor as . Further, using pairwise labels from , we can create a positive pair set defined as . For given proposal, the Graph Contrastive Loss (GCL) can be calculated as:
By training with the proposed loss , the student network is encouraged to learn high-quality feature representations on the target domain. We show that it improves the detector’s performance by conducting experimental analysis in Sec. 4. Note that GCL is used only to update the student network parameters, whereas the teacher network parameters are updated via EMA.
3 Overall objective
So far, we have introduced an Instance Relation Graph (IRG), Graph Distillation Loss (GDL), and Graph Contrastive Loss (GCL) to effectively tackle the source free domain adaptation problem for detection. Then overall objective of our proposed SFDA method is formulated as:
Experiments and Results
To validate the effectiveness of our method, we compare our model performance with existing state-of-the-art UDA and SFDA methods on four different domain shift scenarios; 1) Adaptation to adverse weather, 2) Real to artistic, 3) Synthetic to real, and 4) Cross-camera. Note that in UDA we have access to both source and target domain data. However, in SFDA, we have access only to source-trained model and not the source domain data for adaptation.
Following the SFDA setting , we adopt Faster-RCNN with ImageNet pre-trained ResNet50 as the backbone. In all of our experiments, the input images are resized with a shorter side to be 600 while maintaining the aspect ratio and the batch size to 1. The source model is trained using SGD optimizer with a learning rate of 0.001 and momentum of 0.9 for 10 epochs. For the proposed framework, the teacher network EMA momentum rate is set equal to 0.9. In addition, the pseudo-labels generated by the teacher network with confidence greater than the threshold =0.9 are selected for student training. We utilize the SGD optimizer to train the student network with a learning rate of 0.001 and momentum of 0.9 for 10 epochs. We report the mean Average Precision (mAP) with an IoU threshold of 0.5 for the teacher network on the target domain during the evaluation.
2 Quantitative comparison
Description. Given a model trained on clear weather condition, we aim to perform adaptation to images in adverse weather conditions like fog/haze etc. The Cityscapes consist of 2,975 training images and 500 validation images with 8 object categories: person, rider, car, truck, bus, train, motorcycle and bicycle. The FoggyCityscapes consist of images that are rendered from the Cityscapes dataset by integrating fog and depth information. To this end, a model trained on Cityscapes is adapted to FoggyCityscapes without having access to the Cityscapes dataset.
Results. Table 1 provides the quantitative comparison with the existing UDA and SFDA methods for CityscapeFoggyCityscapes adaptation scenario. From Table 1, we can infer that the proposed method outperforms most of the existing UDA methods such as SWDA , InstanceDA , and CategoricalDA . However, compared MeGA-CDA and Unbiased DA methods, our proposed method produces a competitive performance with a drop of 2.5-3.5 mAP. But it is worth noting that, these method make use of labelled source data during adaptation whereas our proposed method only has access to source-trained model. Furthermore, compared with existing SFDA methods, SFOD and HCL , the proposed method provides improvement of 4.3 mAP and 3.2 mAP, respectively. We also compared with mean-teacher self-training baseline to show that adding the proposed GCL loss is able to enhance the features representation on the target domain, providing an improvement of 3.5 mAP.
2.2 Realistic to artistic data adaptation
Description. Here, we consider adaptation to dissimilar domains , where a model trained on the real-world images is aimed to perform adaptation towards artistic domain. We consider the model trained on the Pascal-VOC dataset and adapt to two target domains, namely, Clipart and Watercolor . The Clipart dataset contains 1K unlabeled images and has the same 20 categories as Pascal-VOC. The Watercolor consists of 1K training and 1K testing images with six categories.
Results. The PASCAL-VOCClipart adaptation results are reported in Table 4. Our method outperforms the existing UDA methods such as ADDA and DANN by a margin of 4.7 mAP and 0.3 mAP, respectively. Moreover, the PASCAL-VOCWatercolor adaptation results are reported in Table 3. Even in this case, our method outperforms the state-of-the-art UDA methods such as SWDA and Net by 2.6 mAP and 1.5 mAP, respectively. Furthermore, for both Clipart and Watercolor adaptation scenarios, our method consistently outperforms in every category compared with pseudo-label self-training (PL) and mean-teacher baseline.
2.3 Synthetic to real-world adaptation
Description. The cost of generating and labeling synthetic data is low compared to real-world data. Hence, it makes sense to train a detector on synthetic images and transfer the knowledge to real-world data. However, the style shift between synthetic to real domain makes it challenging. Here, we consider such scenario where we adapt a model trained on the synthetic data, Sim10K , to a real-world data, Cityscapes under SFDA condition, i.e., synthetic data are not available while adapting the model to the real-world images. The model is trained on 10,000 training images of Sim10k rendered by the Grand Theft Auto gaming engine. The target Cityscapes dataset consists of 2,975 unlabeled training images and 500 validation images.
Results. We report the results of Sim10KCityscapes in Table 2. Note that even though we adapt for only the car category, the proposed GCL training strategy is able to get discriminative positive pairs for different cars and improve the feature representations through contrastive training. Our proposed method outperforms existing UDA method like Cycle DA , Unbiased DA etc. by considerable margin. Under SFDA setting, the proposed method produces state-of-the-art performance by improving 1 mAP compared to SFOD .
2.4 Cross-camera adaptation
Description. In real-world scenarios, the target domain data is captured by a camera with configurations different from the source data. To emulate this cross-camera conditions, we consider a model trained on source, KITTI dataset , is adapted to target, Cityscapes . The KITTI dataset consists of 7,481 training images, which is used to get the source-trained detector model. The model is then adapted to the target domain dataset, i.e., Cityscapes.
Results. KITTICityscapes results are reported in Table 2. Our method outperform existing state-of-the-art UDA methods like Cycle DA , MeGA CDA and Unbiased DA by considerable margin. Further in SFDA setting, the proposed method produce state-of-the-art performance by improving around 1.1 mAP compared to SFOD.
3 Ablation analysis
We study the impact of the proposed GCL and IRG network by performing an in-depth ablation analysis on CityscapesFoggyCityscapes adaptation scenario.
Quantitative analysis. The results for CityscapesFoggyCityscapes ablation experiments are reported in Table 5. In Table 5, the first three experiments are performed to analyze the effect of various combinations of weak and strong augmentation for a mean-teacher framework in an SFDA setting. More precisely, we input the student and teacher network with Weak-Weak (WW), Strong-Strong (SS) and Strong-Weak (SW) augmented images, respectively. These three experiments show that strong-weak (SW) produces consistent and improved results compared to other variations. This is due to mutual learning between student and teacher networks, where student trains on strong augmentation leading to robust prediction and the teacher supervise the student by good pseudo-labels predicted from the weak augmented images. Furthermore, minimizing the discrepancy between instance relation graph network of student and teacher framework ensures consistency between student and teacher graph proposal feature representations. Subsequently, addition of graph distillation loss enhances the model performance from 34.3 mAP to 35.9 mAP. Finally, utilizing graph-guided contrastive learning on the proposal features further helps the model learn high-quality representations, resulting in an increase in performance by 1.9 mAP on the target domain.
Qualitative analysis. In Fig. 6, we show the relation matrix for the RoI features before and after it is processed by IRG. For better visualizations, we consider 25 out of 300 RoI features. It can be observed that relation between the proposals are poorly defined and IRG network is able to improve these relations through graph-based feature aggregation.
Conclusion
In this work, we presented a novel approach for source-free domain adaptive detection using graph-guided contrastive learning. Specifically, we introduced a contrastive graph loss to enhance the target domain representations by exploiting instance relations. We propose an instance relation graph network built on top of a graph convolution network to model the relation between proposal instances. Subsequently, the learned instance relations are used to get positive/negative proposal pairs to guide contrastive learning. We conduct extensive experiments on multiple detection benchmarks to show that the proposed method efficiently adapts a source-trained object detector to the target domain, outperforming the state-of-the-art source-free domain adaptation and many unsupervised domain adaptation methods.