Towards Online Domain Adaptive Object Detection

Vibashan VS, Poojan Oza, Vishal M. Patel

Introduction

The ability to train deep network models on large-scale annotated datasets has accelerated the progress for multiple computer vision tasks such as classification , segmentation , and detection . Despite this success, these models have limited generalization capabilities . Specifically, the model performance drops when the test data (target domain) is sampled from a different distribution than that of the training data (source domain) . For example, when a model is deployed in real-world applications such as autonomous navigation, it could encounter images with weather-based degradations, camera artifacts, etc., unknown during training.

Unsupervised Domain Adaptation (UDA) methods are generally employed to improve model generalization under domain shift condition. Existing UDA methods assume that both labelled source data and unlabeled target data are available during adaptation. This scenario is often not feasible in current real-world applications, as the labelled source data is often restricted due to privacy regulations, data transmission constraints, or proprietary data concerns. To overcome this drawback, recently, some works have explored Source-Free Domain Adaptation (SFDA) setting, where a source-trained model is adapted towards the target domain without requiring access to the source data. However, in both UDA and SFDA settings, adaptation is performed in an offline manner where the model is first adapted towards the target domain and then deployed in real-world applications. In addition, it is often impossible to have prior knowledge about the target domain in most real-world applications. In other words, the deployed model could encounter a diverse set of target domains and offline adaptation to every distribution shift would be infeasible. Therefore, we propose a unified adaptation framework which utilizes a source-trained detector and adapts to the target domain in both offline and online manner.

In recent years, few works have explored various test-time adaptation settings where adaptation is performed during test-time . Wang proposed a fully test-time adaptation strategy which performs entropy minimization during test-time and only updates the model batch-norm parameters for the classification task. However, extending TENT to detection framework has two critical drawbacks: 1) TENT use a very large batch size during test-time adaptation, which is not feasible during real-time deployment as images arrive one by one sequentially. 2) Updating only the batch norm parameter of a network with batch size 1 essentially degrades the model performance . Although existing test-time adaptation settings are closer to online-SFDA settings, they are not suitable for adapting a detection model during real-world deployment. To overcome these issues, we explore an Online Source-Free Domain Adaptation (Online-SFDA)setting, where a model is adapted to any distribution shifts encountered during deployment in an online manner with batch size 1. Fig. 1 illustrate the online source-free domain adaptation setting for detection and its differences against the other adaptation settings.

Source-free domain adaptive object detection is relatively new and a more challenging setting than UDA. Existing SFDA methods for detection adapt to the target domain by training on the pseudo-labels generated by the source-trained model. Due to domain shifts, these generated pseudo-labels are noisy and training a model on top of them would lead to noise overfitting . To alleviate these issues, we employ a mean-teacher framework where the student model is supervised using pseudo-labels generated by the teacher network and the teacher network is slowly updated via the exponential moving average (EMA) of student weights. Therefore, the student network is trained on consistent pseudo-labels leading to less overfitting and the teacher network is a gradual ensemble of target adapted student weights . However, this strategy is inefficient in learning two critical aspects required for optimal online adaptation: 1) They fail to learn robust target feature representation, 2) They fail to fully exploit the online target samples. Hence, we propose a novel memory module and a contrastive loss to fully utilize online target samples and learn robust target feature representation.

Contrastive Learning (CL) aims to learn high-quality features from unlabeled data by forcing similar object instances to stay close and push dissimilar ones apart in an unsupervised manner. This is especially useful for online-SFDA as source-labelled data are unavailable during adaptation. Existing CL methods are designed for classification tasks where they operate on image-level features and require multiple image views (or augmentations) to learn robust feature representation. Consequently, obtaining these large sets of views through input augmentations is computationally expensive for adapting detector models. However in detector models, it is possible to obtain different views for an object in an input image without heavy input augmentations. More precisely, the detector provides multiple object proposals generated by Region Proposal Network (RPN), which in turn provides multiple cropped views around the object instance at different locations and various scales. Therefore, applying CL loss on RPN cropped views guides the model to learn object-level feature representation on the target domain. Note this CL loss is used to supervise the student network, where the object-level features are obtained from the student RoI features. However to perform contrastive learning, these student RoI features require positive and negative pairs. To this end, we propose MemXfromer, a cross-attention transformer-based memory module where items in the memory record prototypical patterns of the continuous target distribution. The proposed MemXfromer solves two important problems for online adaptation: 1) store the target distribution during online adaptation, which are utilized for future adaptation. 2) stored temporal ensemble of target representations provides positive and negative pairs to guide the contrastive learning process. Further, we introduce a cross-attention based read and write technique which models better target distribution and provides strong positive and negative pairs for contrastive learning. Note that the proposed method is not only suitable for online adaptation but also for offline adaptation. In a nutshell, this paper makes the following contributions:

To the best of our knowledge, this is the first work to consider both online and offline adaptation settings for detector models.

We propose a novel unified adaptation framework which makes the detector models robust against online target distribution shifts.

We introduce the MemXformer module, which stores prototypical patterns of the target distribution and provides contrastive pairs to boost contrastive learning on the target domain.

We consider multiple detection benchmarks for experimental analysis and show that the proposed method outperforms existing UDA, and SFDA methods for both online and offline settings.

Related works

Unsupervised domain adaptation. Existing unsupervised domain adaptation methods can be categorized into three groups based on adversarial training , self-training and image-to-image translation . The first domain adaptive object detection was studied in , where they followed an adversarial-based strategy to perform feature alignment at both image-level and instance-level to mitigate the domain shift. Later, Saito proposed an adversarial-based strategy where strong alignment of the local features and weak alignment of the global features. Kim , introduced an image-to-image translation-based where multiple target domain images are created by stylizing the labelled source images. Multiple discriminators are used to performing adversarial alignment to reduce domain discrepancy by utilizing these target-styled source images. In , a pseudo-label based training strategy was formulated to counter noise in pseudo-labels to perform robust training of object detectors on the target domain. However, all these works assume to have access to labelled source data and unlabeled target data during adaptation, and they operate in an offline setting.

Source-free domain adaptation. In the source-free domain adaptation setting, we have a source-trained model which adapts to the target domain without having access to source data. Multiple works have addressed the source-free domain adaptation (SFDA) setting for classification , segmentation and object detection tasks. In detail, for classification task proposed a self-supervised method to learn target domain representation via information maximization. Further for segmentation and object detection , the proposed methods are based on pseudo-label self-training to learn target-specific representation. However, similar to existing UDA works, these SFDA methods operate in an offline setting. Thus, we explore online adaptation, which is a more practical way to tackle domain shifts for real-world applications.

Online adaptation. Sun proposed a Test-time training (TTT) strategy, where a model is trained on source data along with an auxiliary task (eg: rotation prediction) which is utilized during test-time to fine-tune the model on target test distribution. The major drawback of this adaptation strategy is training an auxiliary task along with source training just to perform adaptation during test-time is not a feasible solution and effective solution for real-world application. Later, Wang proposed a fully test-time adaptation setting, where the given source trained model adapts to the target domain by entropy minimization during test-time in an online manner by entropy minimization. In this way, Tent adapts to the target domain with test-time loss. Here, the major limitation of is a requirement of a large batch size during test-time adaptation, which is not feasible during real-time deployment as images arrive one by one sequentially. Although existing test-time adaptation settings provide close resemblance to online-SFDA settings, these test-time settings are not suitable for adapting a detection model during real-world deployment. Therefore in this work, we explore both online and offline adaptation settings for the object detection tasks.

Contrastive representation learning. Contrastive representation learning has shown huge progress towards unsupervised feature learning. The standard way of formulating contrastive learning for an anchor is by pulling together the feature embedding of anchor’s positive pairs and pushing apart from the anchor’s negative pair . These positive and negative pairs are formed by augmenting the anchor image and sampling from the input batch of images. Thus, for a given anchor, the positive pair are augmented anchor images and the negative pairs are other images from the batch . On top of this, by exploiting the task-specific label information, performed contrastive learning in a supervised manner. Nonetheless, all these tasks require a large batch size to perform contrastive learning effectively and it is not feasible to have more than one image during online adaptation. Thus, we propose a memory-based contrastive learning framework suitable for adapting object detectors during deployment in an online manner.

Proposed method

The online-SFDA setting considers a source-trained model with parameters Θsrc\Theta_{src} and adapts to any target distribution shifts during real-world deployment as illustrated in Fig. 1. Let us consider a stream of online target data denoted as T={x1,x2,..,xn}\mathcal{T}=\{x_{1},x_{2},..,x_{n}\}, where xnx_{n} is the nthn^{th} online sample. Since these samples arrive sequentially, the model gets adapted to each sample and the adapted weights are used for future online samples. Specifically, the model parameters during adaptation on the nthn^{th} sample xnx_{n}, i.e. Θsrc(n)\Theta_{src}^{(n)}, are initialized with the model parameters updated through online adaptation of previous xn−1thx_{n-1}^{th} sample. To summarize, online-SFDA performs continuous online adaptation, i.e., adaptation will be continued as long as there is a stream of data and necessity.

Student-teacher training. In online-SFDA, the model parameters need to be continuously updated in an online unsupervised manner. Consequently, the model risks forgetting the original hypothesis learned through supervised source training . To overcome this, prior works have employed a student-teacher framework. Specifically, the student parameters (Θstd\Theta_{std}) are adapted to the target domain by minimizing the detection loss supervised through the teacher-generated pseudo-labels. The adapted student parameters are then transferred to the teacher parameters (Θtch\Theta_{tch}) via Exponential Moving Average (EMA). This can be formally written as:

Contrastive Learning (CL). SimCLR is a commonly used CL framework which learns representations for an image by maximizing agreement between differently augmented views of the same sample. For given an anchor image xix_{i}, the SimCLR loss can be written as:

where NN is the batch size, ziz_{i} and zjz_{j} are the features of two different augmentations of the same sample xix_{i}, whereas zlz_{l} represents the feature of the lthl^{th} batch sample xlx_{l}, where l≠il\neq i. Also, sim⁡(⋅,⋅)\operatorname{sim}(\cdot,\cdot) indicates a similarity function, e.g. cosine similarity. Note that in general, the CRL framework assumes that each image contains one category/object . Moreover, it requires large batch sizes that could provide multiple positive/negative pairs for training . In contrast for object detection, each image will have multiple objects and a large batch size or multiple views are computationally not feasible. Hence, existing CRL methods are more suited for classification tasks.

Though existing contrastive learning methods like SimCLR are exceptional at learning high-quality representations, they are more suitable for the classification task. For detection, these CL methods require large batch size and heavy input augmentation, which are computationally expensive to apply for online parameter updates (discussed in Sec. 1). Therefore, we utilize a computationally efficient memory-based approach to make contrastive learning feasible and effective for online model updates. The proposed online-SFDA strategy is illustrated in Fig. 2.

where the cross-attention map StS_{t} is a 2D matrix of size Nm×NfN_{m}\times N_{f} and sti,js_{t}^{i,j} represents how jthj^{th} memory items is related to ithi^{th} teacher RoI features. We utilize this cross-attention map StS_{t} and VtV_{t} to update jthj^{th} memory item using following equation:

where F(.)F(.) is L2L_{2} norm. Therefore, using attention-based weighted average and global memory bank update for each online sample makes the MemXformer effectively store and model the target distribution.

where the cross-attention map SsS_{s} is a 2D matrix of size Nm×NfN_{m}\times N_{f} and given ithi^{th} student RoI features as query, the stis_{t}^{i} thth row presents NlN_{l} memory items attention score. Therefore, given ithi^{th} student RoI features as query, we generate its corresponding positive pair by attention guided weighted sum of most similar memory items. Thus, utilizing the cross-attention map SsS_{s} and considering memory items as value Vm={mj}j=1NlV_{m}=\{m^{j}\}_{j=1}^{N_{l}}, we compute the strong positive pair for ithi^{th} student RoI features using following equation:

where Ps={psi}i=1Nf\mathcal{P}_{s}=\{p^{i}_{s}\}_{i=1}^{N_{f}} corresponds to set of strong positive pair for student RoI features Fs\mathcal{F}_{s}. In detail, the retrieved positive pairs are temporal ensembles of the prototypical target distribution, which gives more information regarding the online target distribution shifts. This essentially guides contrastive learning to model the target distribution.

Negative Pair Mining. As explained earlier from MemXformer read operation, we obtain a set of strong positive pairs for a given student RoI feature. These strong positive pairs are essentially an ensemble of most similar memory items. However, these ensembled similar memory items also contain dissimilar memory items but are scaled with less attention weights. This restricts the contrastive learning capability to effectively model the target domain representation. To mitigate the dissimilar item’s effect on CL, we propose negative pair mining. Specifically in negative pair mining, given a student RoI feature as query and cross-attention map SsS_{s}, we mine the least similar 10%\% of the memory items and label them as negative pairs Ns={min}i=1Ns\mathcal{N}_{s}=\{m_{i}^{n}\}_{i=1}^{N_{s}}. As a result, by performing negative pair mining, we obtain NsN_{s} negative samples for one positive sample, where NsN_{s} is top 10%\% of least similar memory items.

Memory contrastive loss. Given student RoI feature fsif_{s}^{i} as anchor, utilizing MemXformer Read operation and negative pair mining we obtain strong positive Ps\mathcal{P}_{s} and negative pairs Ns\mathcal{N}_{s} from MemXformer. Therefore, given an image xnx_{n} with student RoI feature Fs\mathcal{F}_{s}, the MemCLR loss is calculated as: LMemCLR(xn)=−log⁡{1∣Fs∣∑i∈Fsexp⁡(fsi⋅psi)exp⁡(fsi⋅psi)+∑n∈Nsexp⁡(fsi⋅mn)},\mathcal{L}_{\text{MemCLR}}({x}_{n})=\\ -\log\left\{\frac{1}{\left|\mathcal{F}_{s}\right|}\sum_{i\in\mathcal{F}_{s}}\frac{\operatorname{exp}(f_{s}^{i}\cdot p_{s}^{i})}{\operatorname{exp}(f_{s}^{i}\cdot p_{s}^{i})+\sum_{n\in\mathcal{N}_{s}}\operatorname{exp}(f_{s}^{i}\cdot{m}^{n})}\right\}, Therefore, minimizing the MemCLR loss guided by strong positive and negative pairs enhance the student model to learn better target representation in an online-SFDA setting.

Overall loss. We illustrate our overall architecture for online source-free domain adaptation in Fig. 2. The proposed method utilizes a global memory bank to perform memory-based contrastive learning to robustify the representations under varying target distribution shifts. Therefore, the overall online-SFDA loss for any online sample xnx_{n} can be calculated as:

Experiments and Results

To validate the proposed method, we consider four domain shift scenarios where the source train model is adapted to the unlabelled target domain, typically used for comparison in UDA and SFDA literature. Specifically, we evaluate the proposed method with the existing UDA, SFDA and Test-time works under four domain shifts, 1) clear-weather to foggy-weather, 2) real to artistic, 3) synthetic to real, and 4) cross-camera adaptation. Note that, to show the effectiveness of our proposed approach, we evaluate both online and offline settings. Specifically, the offline setting follows the standard SFDA setting. The source-trained model is adapted towards the target domain using an unlabelled target train-set for multiple iterations and evaluated on the target test-set. Whereas in the online setting, the model is adapted towards the target domain in an online manner where the target test samples are seen only once and finally evaluated on the target test-set. This essentially simulates the real-world scenario where you see the target samples only once and adaptation needs to be continuous.

For the Online adaptation setting, we adopt Faster-RCNN with ResNet50 as the backbone pre-trained on ImageNet . In all of our experiments, the input images are resized with a shorter side to be 600 pixels while maintaining the aspect ratio. We set the batch size to 1 for all experiments. For the student-teacher framework, the weight momentum update parameter α\alpha of the EMA for the teacher model is set equal to 0.99. The pseudo-labels generated by the teacher network with confidence greater than the threshold TT=0.9 are selected for student training. We utilize an SGD optimizer to train the student network with a learning rate of 0.001 and momentum of 0.9 for both online and offline training. The Global Memory Bank contains NmN_{m} memory items, which are set to 1024. Further, the source model is trained using an SGD optimizer with a learning rate of 0.001 and momentum of 0.9 for 10 epochs. We report the mean Average Precision (mAP) with an IoU threshold of 0.5 for the teacher network on the distribution-shifted target domain test data during the evaluation.

When the source-trained models are deployed in real-world applications such as autonomous navigation, they are likely to encounter data from multiple weather conditions such as fog, haze, etc. In most cases, the deployed detector models would be trained for clear weather conditions. We propose to formulate this as an online adaptation problem, as it is difficult to pre-determine what kind of weather conditions will occur. Subsequently, we update the detector model in an online manner to adapt to any weather shifts the model might observe after deployment. To evaluate the proposed method under such conditions, we experiment on Cityscapes →\rightarrow FoggyCityscapes dataset. Here, we have a detection model trained on the Cityscapes dataset consisting of 2,975 normal weather images and 500 test images with 8 object categories: person, rider, car, truck, bus, train, motorcycle and bicycle. During inference, images from FoggyCityscapes are sequentially sent and the object detection model is adapted in an online manner to improve generalization on foggy/hazy weather. Table 1 provides the comparison of the proposed FTTA method with the state-of-the-art UDA, SFDA, and O-SFDA methods for Cityscape→\rightarrowFoggyCityscapes adaptation scenario. From Table 1, we can infer that UDA and SFDA methods operate in an offline manner, where as O-SFDA operates in an online manner. Firstly, in the online setting our proposed method outperforms existing UDA methods such as SWDA , MTOR and InstanceDA by a considerable margin. However, compared to MeGA-CDA and Unbiased DA our proposed method produces competitive performance with a drop of 3-4 mAP. Note that these UDA methods have access to labelled source data, whereas under the SFDA setting, the proposed model only has access to source-trained model. Furthermore, the proposed method outperforms SFDA methods like SFOD and HCL by 1.7 and 0.6 mAP, respectively. Secondly, when compared to the Test-time adaptation based methods such as Tent , our best-performing model surpasses it by a huge margin of by 3.0 mAP. Therefore, for Cityscape→\rightarrowFoggyCityscapes adaptation scenario, our proposed method produces state-of-the-art results in both online and offline SFDA settings.

1.2 Synthetic to real world adaptation

Collecting and annotating detection data is computationally intensive, where on top of assigning a category, one needs to add bounding boxes to every object location in the image. On the other hand, creating a synthetic dataset through simulation is much less computation-intensive and generates annotations for free. Hence, training a detector model on a synthetically generated dataset makes sense and then deploying it in real-world conditions. However, stylistic/appearance differences between real and synthetic data limit such deployment due to performance issues. Here, we formulate it as an online adaptation problem to update a synthetic data trained model on the real-world test data. Particularly, we consider a source model trained on Sim10k on 10,000 training images with 58,701 bounding boxes of car category, rendered by the gaming engine Grand Theft Auto. For real-world test data we use the Cityscapes validation set for online model adaptation. In Table 2, we report Sim10K→\rightarrowCityscapes adaptation results on the existing UDA, SFDA, and O-SFDA methods. In an offline setting, compared to the existing UDA works such as DAFaster , SWDA and RobustDA , the proposed method outperforms all of them by a considerable margin. Furthermore, when compared to SFOD the proposed method is better by 0.7 mAP. In an online setting, compared to Tent , our proposed method outperforms it by 4.0 mAP. Therefore, our proposed is able to perform well under synthetic to real-world domain shifts.

1.3 Cross-camera adaptation

In most real-world applications, it is assumed that both training and test data would be collected using a camera with the same parameters. However, the camera parameters are often different, which causes the collected images to have different appearances, such as radial distortions, tangential distortions, etc. This can cause the model to perform poorly due to changes in the camera parameters. Hence, to tackle any such camera distortions, we formulate the problem as an online adaptation problem and show that the proposed approach succeeds in generalizing to such cases. Here, we have access to only the source model, trained on the KITTI dataset with 7,481 training images with bounding boxes for the car category. To emulate cross-camera scenario, we consider online adaptation on the Citsycapes validation set containing 500 images. We report the results of the cross-camera adaptation experiment in Table 2. Similar to Sim→\rightarrowCityscapes adaptation even for Kitti→\rightarrowCityscapes adaptation, we show similar performance improvements compared to UDA, SFDA and O-SFDA methods. Specifically, in the O-SFDA setting, the proposed method outperforms Tent by 5.6 mAP. Thus, our proposed method is able to model the cross-camera domain shifts effectively.

1.4 Real to artistic adaptation

Here, we evaluate the proposed method for the case where there is a concept shift in during inference. By concept shift, we refer to the case where there is a complete change in the object, e.g., going from real-world to artistic images. Unlike previous scenarios where the objects go through stylistic/appearance changes, the entire concept of an object is different, e.g., a real-world car vs a cartoon car . We show that even in this challenging scenario, the proposed approach is able to improve model generalization through online updates. We consider a model trained on the Pascal-VOC data which adapts to test set of Watercolor . Specifically, the Watercolor consists of 1K training and 1K testing images with six categories.We compare PASCAL-VOC→\rightarrowWatercolor results with the existing methods in Table 3. From Table 3, we can infer that the proposed method outperforms most of the existing UDA methods and SFDA methods in offline settings. Further, in the online setting, when compared to TENT the proposed method is able to outperform by a significant margin. This demonstrates the capability of the proposed method to generalize even for both online and offline settings.

2 Ablation analysis

Quantitative analysis. The Cityscapes→\rightarrowFoggyCityscapes ablation experiment results are reported in Table 4 for the offline-SFDA setting. We first consider a student-teacher offline update baseline which, compared to the source-only baseline, provides significant improvements. To have a fair comparison, we also consider utilizing supervised contrastive loss for offline updates. In particular, we utilize predictions provided by student-teacher training as label information needed for applying the supervised contrastive loss over object proposals. Denoted as SupCon in Table 4, the addition of supervised contrastive learning further improves the performance by 1.3 mAP. However, the proposed memory-based contrastive learning outperforms the supervised contrastive learning by 1.4 mAP, indicating the utility of the proposed method to learn better target representations. Finally, we analyze the performance of the proposed method by varying global memory bank capacity from 256 to 1024 memory items. As shown in Table 4, memory-based contrastive loss with 1024 memory items performs the best when compared to 256 and 512 memory items. Further, note that our model takes around 1 second to perform online adaptation for one sample.

Qualitative analysis. Fig. 4 shows t-SNE visualization for source-only, student-teacher training and the proposed method for the Cityscapes→\rightarrowFoggyCityscapes online-SFDA setting. The t-SNE visualizations are created from the RoI features extracted from the predictions for 500 test images. Due to the distribution shift, the features are dispersed for the source-only baseline and classification boundaries are weak. With the help of student-teacher training, the model learns better classification boundaries, resulting in better quantitative performance. However, the features in the student-teacher training have a large variance and do not have compact features. Whereas the proposed method has even better classification boundaries and learns compact features for each category, resulting in a more robust model. Further qualitative comparison is performed to analyze the effect of the order of input sequence during online adaptation is shown in Fig. 5. Multiple experiments with changing the order of input sequence are conducted and corresponding performance mean and variance is plotted in Fig. 5. We can observe from the variance that the order of input sequence does not much affect the model’s performance. Further, we can observe the model performance increase as it encounters more test samples during online adaptation, showing the MemXformer effectiveness in exploiting online target distribution. Note that, in online adaptation, the test samples are seen only once and adaptation happens in an unsupervised manner.

Conclusion

In this work, we introduced a practical domain adaptation setting for the object detection task, which is feasible for real-world settings. Particularly, we proposed a novel unified adaptation framework which makes the detector models robust against online target distribution shifts. Further, We introduce the MemXformer module, which stores prototypical patterns of the target distribution and provides contrastive pairs to boost the contrastive learning on the target domain. We conducted extensive experiments on multiple detection benchmark datasets and compared existing unsupervised domain adaptation, source-free domain adaptation and test-time adaptation methods to show the effectiveness of the proposed approach for both online and offline adaptation of object detection models. We also analyzed multiple aspects of the proposed method in ablation experiments and identified increasing the online adaptation speed further is a potential directions for future research.

References