Decoupled Adaptation for Cross-Domain Object Detection
Junguang Jiang, Baixu Chen, Jianmin Wang, Mingsheng Long
Introduction
The object detection task has aroused great interest due to its wide applications. In the past few years, the development of deep neural networks has boosted the performance of object detectors [33; 15; 41]. While these detectors have achieved excellent performance on the benchmark datasets [11; 31], object detection in the real world still faces challenges from the large variance in viewpoints, object appearance, backgrounds, illumination, image quality, etc. Such domain shifts have been observed to cause significant performance drop . Thus, some work uses domain adaptation to transfer a detector from a source domain, where sufficient training data is available, to a target domain where only unlabeled data is available [8; 43]. This technique successfully improves the performance of the detector on the target domain. However, the improvement of domain adaptation in object detection remains relatively mild compared with that in object classification.
The inherent challenges come from three aspects. Data challenge: what to adapt in the object detection task is unknown. Instance feature adaptation in the object level (Figure 1(a)) might confuse the features of the foreground and the background since the generated proposals may not be true objects and many true objects might be missing (Figure 5). Global feature adaptation in the image level (Figure 1(b)) is likely to mix up features of different objects since each input image of detection has multiple objects. Local feature adaptation in the pixel level (Figure 1(c)) can alleviate domain shift when the shift is primarily low-level, yet it will struggle when the domains are different at the semantic level. Architecture challenge: while the above adaptation methods introduce domain discriminators and gradient reverse layers into the detector architecture to encourage domain-invariant features, the discriminability of features might get deteriorated [6; 5], which will greatly influence the localization and the classification of the detectors. Besides, where to place these modules in the detection architecture has a great impact on the final performance but is a little tricky. Therefore, the scalability of these methods to different detection architectures is not so satisfactory. Task challenge: object detection is a multi-task learning problem, consisting of both classification and localization. Yet previous adaptation algorithms mainly explored the category adaptation, and it’s still difficult to obtain an adaptation model suitable for different tasks at the same time.
To overcome these challenges, we propose a general framework – D-adapt, namely Decoupled Adaptation. Since adversarial alignment directly on the features of the detector might hurt its discriminability (architecture challenge), we decouple the adversarial adaptation from the training of the detector by introducing a parameter-independent category adaptor (see Figure 1(d)). To tackle the task challenge, we introduce another bounding box adaptor that’s decoupled from both the detector and the category adaptor. To tackle the data challenge, we propose to adjust the object-level data distribution for specific adaptation tasks. For example, in the category adaptation step, we encourage the input proposals to have IoUThe Intersection-over-Union between the proposals and the ground-truth instance. close to or to better satisfy the low-density separation assumption, while in the bounding box adaptation step, we encourage the input proposals to have IoU between and to ease the optimization of the bounding box localization task.
The contributions of this work are summarized as three-fold. (1) We introduce D-adapt framework for cross-domain object detection, which is general for both two-stage and single-stage detectors. (2) We propose an effective method to adapt the bounding box localization task, which is ignored by existing methods but is crucial for achieving superior final performance. (3) We conduct extensive experiments and validate that our method achieves state-of-the-art performance on four object detection tasks, and yields 17% and 21% relative improvement on Clipart1k and Comic2k.
Related Work
Domain adaptation is proposed to overcome the distribution shift across domains. In the classification setting, most of the domain adaptation methods are based on Moment Matching or Adversarial Adaptation. Moment Matching methods [50; 36] align distributions by minimizing the distribution discrepancy in the feature space. Taking the same spirit as Generative Adversarial Networks , Adversarial Adaptation [12; 37] introduces a domain discriminator to distinguish the source from the target, then the feature extractor is encouraged to fool the discriminator and learn domain invariant features. However, directly applying these methods to object detection yields an unsatisfactory effect. The difficulty is that the image of object detection usually contains multiple objects, thus the features of an image can have complex multimodal structures [20; 58; 5], making the image-level feature alignment problematic [58; 20].
Generic domain adaptation for regression.
Most domain adaptation methods designed for classification do not work well on regression tasks since the regression space is continuous with no clear decision boundary . Some specific regression algorithms are proposed, including importance weighting or learning invariant representations [40; 38]. RSD defines a geometrical distance for learning transferable representations and disparity discrepancy proposes an upper bound for the distribution distance in the regression problems. Yet previous methods are mainly tested on simple tasks while this paper extends domain adaptation to the object localization tasks.
Domain adaptation for object detection.
DA-Faster performs feature alignment at both image-level and instance-level. SWDA proposes that strong alignment of the local features is more effective than the strong alignment of the global features. Hsu et al. carries out center-aware alignment by paying more attention to foreground pixels. HTCN calibrates the transferability of feature representations hierarchically. Zheng et al. proposes to extract foreground regions and adopts coarse-to-fine feature adaptation. ATF introduces an asymmetric tri-way approach to account for the differences in labeling statistics between domains. CRDA and MCAR use multi-label classification as an auxiliary task to regularize the features. However, although the auxiliary task of outputting domain-invariant features to fool a domain discriminator in most aforementioned methods can improve the transferability, it also impairs the discriminability of the detector. In contrast, we decouple the adversarial adaptation and the training of the detector, thus the adaptors could specialize in transfer between domains, and the detector could focus on improving the discriminability while enjoying the transferability brought by the adaptors.
Self-training with pseudo labels.
Pseudo-labeling , which leverages the model itself to obtain labels on unlabeled data, is widely used in self-training. To generate reliable pseudo labels, temporal ensembling maintains an exponential moving average prediction for each sample, while the mean-teacher averages model weights at different training iterations to get a teacher model. Deep mutual learning trains a pool of student models with supervisions from each other. FixMatch uses the model’s predictions on weakly-augmented images to generate pseudo-labels for the strongly-augmented ones. Unbiased Teacher introduces the teacher-student paradigm to Semi-Supervised Object Detection (SS-OD). When some image-level labels exist, the performance can be further improved by encoding correlations between coarse-grained and fine-grained classes , employing noise-tolerant training strategies , or learning a mapping from weakly-supervised to fully-supervised detectors in SS-OD. Recent works [21; 25; 26] utilize self-training in cross-domain object detection and take the most confident predictions as pseudo labels. MTOR uses the mean teacher framework and UMT adopts distillation and CycleGAN in self-training. However, self-training suffers from the problem of confirmation bias [1; 4]: the performance of the student will be limited by that of the teacher. Although pseudo labels are also used in our proposed D-adapt, they are generated from adaptors that have independent parameters and different tasks from the detector, thereby alleviating the confirmation bias of the overly tight relationship in self-training.
Proposed Method
In supervised object detection, we have a labeled source domain , where is the image, is the bounding box coordinates, and is the categories. The detector is trained with , which consists of four losses in Faster RCNN : the RPN classification loss , the RPN regression loss , the RoI classification loss and the RoI regression loss ,
In cross-domain object detection, there exists another unlabeled target domain that follows different distributions from . The objective of is to improve the performance on .
To deal with the architecture challenge mentioned in Section 1, we propose the D-adapt framework, which has three steps: (1) decouple the original cross-domain detection problem into several sub-problems (2) design adaptors to solve each sub-problem (3) coordinate the relationships between different adaptors and the detector.
Since adaptation might hurt the discriminability of the detector, we decouple the category adaptation from the training of the detector by introducing a parameter-independent category adaptor (see Figure 1(d)). The adaptation is only performed on the features of the category adaptor, thus will not hurt the detector’s ability to locate objects. To fill the blank of regression domain adaptation in object detection, we need to perform adaptation on the bounding box regression. Yet feature visualization in Figure 6(c) reveals that features that contain both category and location information do not have an obvious cluster structure, and alignment might hurt its discriminability. Besides, the common category adaptation methods are also not effective on regression tasks , thus we decouple category adaptation and the bounding box adaptation to avoid their interfering with each other. Section 3.2 and 3.3 will introduce the design of category adaptor and box adaptor in details. In this section, we will assume that such two adaptors are already obtained.
To coordinate the adaptation on different tasks, we maintain a cascading relationship between the adaptors. In the cascading structure, the later adaptors can utilize the information obtained by the previous adaptors for better adaptation, e.g. in the box adaptation step, the category adaptor will select foreground proposals to facilitate the training of the box adaptor. Compared with the multi-task learning relationship where we need to balance the weights of different adaptation losses carefully, the cascade relationship greatly reduces the difficulty of hyper-parameter selection since each adaptor has only one adaptation loss. Since the adaptors are specifically designed for cross-domain tasks, their predictions on the target domain can serve as pseudo labels for the detector. On the other hand, the detector generates proposals to train the adaptors and higher-quality proposals can improve the adaptation performance (see Table 5 for details). And this enables the self-feedback relationship between the detector and the adaptors.
For a good initialization of this self-feedback loop, we first pre-train the detector on the source domain with . Using the pre-trained , we can derive two new data distributions, the source proposal distribution and the target proposal distribution . Each proposal consists of a crop of the image We use uppercase letters to represent the whole image, lowercase letters to represent an instance of object., its corresponding bounding box , predicted category and the class confidence . We can annotate each source-domain proposal with a ground truth bounding box and category label , similar to labeling each RoI in Fast RCNN , and then use these labels to train the adaptors. In turn, for each target proposal , adaptors will provide category pseudo label and box pseudo label to train the RoI heads,
Note that our D-adapt framework does not introduce any computational overhead in the inference phase, since the adaptors are independent of the detector and can be removed during detection. Also, D-adapt does not depend on a specific detector, thus the detector can be replaced by SSD , RetinaNet , or other detectors.
2 Category Adaptation
The goal of category adaptation is to use labeled source-domain proposals to obtain a relatively accurate classification of the unlabeled target-domain proposals . Some generic adaptation methods, such as DANN , can be adopted. DANN introduces a domain discriminator to distinguish the source from the target, then the feature extractor tries to learn domain-invariant representations to fool the discriminator, which will enlarge the decision boundaries between classes on the unlabeled target domain. However, the above adversarial alignment might fail due to the data challenge – the input data distribution doesn’t satisfy the low-density separation assumption well, i.e., the Intersection-over-Union of a proposal and a foreground instance may be any value between 0 and 1 (see Figure 2(a)) and explicit task-specific boundaries between classes hardly exist, which will impede the adversarial alignment . Recall that in standard object detection, proposals with IoU between and will be removed to discretize the input space and ease the optimization of the classification. Yet it can hardly be used in the domain adaptation problem since we cannot obtain ground truth IoU for target proposals.
To overcome the data challenge, we use the confidence of each proposal to discretize the input space, i.e., when a proposal has a high confidence being the foreground or background, it should have a higher weight in the adaptation, and vice versa (see Figure 2(b)). This will reduce the participation of proposals that are neither foreground nor background and improve the discreteness of the input space in the sense of probability. Then the objective of the discriminator is,
where both the feature representation and the category prediction are fed into the domain discriminator (see Figure 2(c)). This will encourage features aligned in a conditional way , and thus avoid that most target proposals aligned to the dominant category on the source domain. The objective of the feature extractor is to separate different categories on the source domain and learn domain-invariant features to fool the discriminator,
where is the cross-entropy loss, is the trade-off between source risk and domain adversarial loss. After obtaining the adapted classifier, we can generate category pseudo label for each proposal .
3 Bounding Box Adaptation
Following RCNN , we adopt a class-specific bounding-box regressor, which predicts the bounding box regression offsets, for each of the foreground classes, indexed by . On the source domain, we have the ground truth category and bounding box label for each proposal, thus we use the smooth loss to train the regressor,
where is the regression prediction, is ground truth category, is the ground truth bounding box offsets calculated from and . However, it’s hard to obtain a satisfactory regressor with on the target domain due to the domain shift.
Inspired by the lastest theory , we propose an IoU disparity discrepancy method. As shown in Figure 3(a), we train a feature generator network which takes proposal inputs, and two regressor networks and which take features from . The objective of the adversarial regressor network is to maximize its disparity with the main regressor on the target domain while minimizing the disparity on the source domain to measure the discrepancy across domains. Then the objective of the adversarial regressor is
Note that on the source domain is only defined on the box corresponding to the ground truth category and that on the target domain is only defined on the box associated with the predicted category . Equation 6 guides the adversarial regressor to predict correctly on the source domain while making as many mistakes as possible on the target domain (Figure 3(b)). Then the feature extractor is encouraged to output domain-invariant features to decrease domain discrepancy,
where is the trade-off between source risk and adversarial loss. After obtaining the adapted regressor, we can generate box pseudo label for each proposal .
Experiments
Following six object detection datasets are used: Pascal VOC , Clipart , Comic , Sim10k , Cityscapes and FoggyCityscapes . Pascal VOC contains categories of common real-world objects and images. Clipart contains 1k images and shares categories with Pascal VOC. Comic2k contains 1k training images and 1k test images, sharing categories with Pascal VOC. Sim10k has images with bounding boxes of car categories, rendered by the gaming engine Grand Theft Auto. Both Cityscapes and FoggyCityscapes have training images and validation images with 8 object categories. Following , we evaluate the domain adaptation performance of different methods on the following four domain adaptation tasks, VOC-to-Clipart, VOC-to-Comic2k, Sim10k-to-Cityscapes, Cityscapes-to-FoggyCityscapes, and report the mean average precision (mAP) with a threshold of .
2 Implementation Details
Stage 1: Source-domain pre-training. In the basic experiments, Faster-RCNN with ResNet-101 or VGG-16 as backbone is adopted and pre-trained the on the source domain with a learning rate of for 12k iterations.
Stage 2: Category adaptation. The category adaptor has the same backbone as the detector but a simple classification head. It’s trained for iterations using SGD optimizer with an initial learning rate of , momentum , and a batch size of for each domain. The discriminator is a three-layer fully connected networks following DANN . is kept for all experiments. is when and otherwise.
Stage 3: Bounding box adaptation. The box adaptor has the same backbone as the detector but a simple regression head (two-layer convolutions networks). The training hyper-parameters (learning rate, batch size, etc.) are the same as that of the category adaptor. is kept for all experiments. The input of the bounding box adaptor (the crops of objects) will be twice larger than the original predicted box, so that the bounding box adapter could access more location information.
Stage 4: Target-domain pseudo-label training. The detector is trained on the target domain for iterations, with an initial learning rate of and reducing to exponentially.
The adaptors and the detector are trained in an alternative way for iterations. We perform all experiments on public datasets using a 1080Ti GPU. Code is available at https://github.com/thuml/Decoupled-Adaptation-for-Cross-Domain-Object-Detection.
3 Comparison with State-of-the-Arts
We first show experiments on dissimilar domains using the Pascal VOC Dataset as the source domain and Clipart as the target domain. Table 1 shows that our proposed method outperforms the state-of-the-art method by points on mAP. Figure 4 presents some qualitative results in the target domain. We also compare with Unbiased Teacher , the state-of-the-art method in semi-supervised object detection, which generates pseudo labels on the target domain from the teacher model. Due to the large domain shift, the prediction from the teacher detection model is unreliable, thus it doesn’t do well. In contrast, our method alleviates the confirmation bias problem by generating pseudo labels from different models (adaptors).
We also use Comic2k as the target domain, which has a very different style from Pascal VOC and a lot of small objects. As shown in Table 2, both image-level and instance-level feature adaptation will fall into the dilemma of transferability and discriminability, and do not work well on this difficult dataset. In contrast, our method effectively solves this problem by decoupling the adversarial adaptation from the training of the detector and improves mAP by 7.0 compared with the state-of-the-art.
Adaptation from synthetic to real images.
We use Sim10k as the source domain and Cityscapes as the target domain. Following , we evaluate on the validation split of the Cityscapes and report the mAP on car. Table 4 shows that our method surpasses all other methods.
Adaptation between similar domains.
We perform adaptation from Cityscapes to FoggyCityscape and report the results* denotes this method utilizes CycleGAN to perform source-to-target translation. in Table 4. Note that since the two domains are relatively similar, the performance of adaptation is already close to the oracle results.
4 Ablation Studies
In this part, we will analyze both the performance of the detector and the adaptors. Denote be the number of proposals of class predicted as class , be the total number of proposals of class , and be the number of classes (including the background), then we use to measure the overall performance of the category adaptor. We use the intersection-over-union between the predicted bounding boxes and the ground truth boxes, i.e., , to measure the performance of the bounding box adaptor. All ablations are performed on VOC Clipart and the iteration is kept 1 for a fair comparison.
Table 7(a) show the effectiveness of several specific designs mentioned in Section 3.2. Among them, the weight mechanism has the greatest impact, indicating the necessity of the low-density assumption in the adversarial adaptation. To verify this, we assume that the ground truth IoU of each proposal is known, and then we select the proposal with IoU greater than a certain threshold when we train the category adaptor. Table 5 shows that as the IoU threshold of the foreground proposals improves from to , the accuracy of the category adaptor will increase from to , which shows the importance of the low-density separation assumption.
Ablation on the bounding box adaptation.
Table 7(b) illustrates that minimizing the disparity discrepancy improves the performance of the box adaptor and bounding box adaptation improves the performance of the detector in the target domain.
Ablation on the training strategy with pseudo-labels.
In Equation 2, losses are only calculated on the regions where the proposals are located, and those anchor areas overlapping with the proposals are ignored. Here, we compare this strategy with the common practice in self-training – filter out bounding boxes with low confidence, then label each proposal that overlaps with these boxes. Although the category labels of these bounding boxes are also generated from the category adaptor, the accuracy of these generated proposals is low (see Table 7(c)). In contrast, our strategy is more conservative and both the on the proposals and the final mAP of the detector are higher.
5 Analysis
Figure 5 gives the percent of error of each model on VOCClipart following . The main errors in the target domain come from: Miss (ground truth regarded as backgrounds) and Cls (classified incorrectly). Loc (classified correctly but localized incorrectly) errors are slightly less, but still cannot be ignored especially after category adaptation, which implies the necessity of box adaptation in object detection. Category adaptation can effectively reduce the proportion of Cls errors while increasing that of Loc errors, thus it is reasonable to cascade the box adaptor after the category adaptor. Bounding box adaptation can reduce the proportion of Loc errors, revealing its effectiveness.
Feature visualization.
We visualize by t-SNE in Figures 6(a)-6(b) the representations of task VOC Comic2k (6 classes) by category adaptor with and category adaptor with . The source and target are well aligned in the latter, which indicates that it learns domain-invariant features. We also extract box features from the detector and get Figure 6(c)-6(d). We find that the features of the detector do not have an obvious cluster structure, even on the source domain. The reason is that the features of the detector contain both category information and location information. Thus adversarial adaptation directly on the detector will hurt its discriminability, while our method achieves better performance through decoupled adaptation.
Discussion and Conclusion
Our method achieved considerable improvement on several benchmark datasets for domain adaptation. In actual deployment, the detection performance can be further boosted by employing stronger adaptors without introducing any computational overhead since the adaptors can be removed during inference. It is also possible to extend the D-adapt framework to other detection tasks, e.g., instance segmentation and keypoint detection, by cascading more specially designed adaptors. We hope D-adapt will be useful for the wider application of detection tasks.
Acknowledgements
This work was supported by the National Megaproject for New Generation AI (2020AAA0109201), National Natural Science Foundation of China (62022050 and 62021002), Beijing Nova Program (Z201100006820041), and BNRist Innovation Fund (BNR2021RC01002).
References
Appendix A More Experiment Results
As shown in Tables 7, our method also applies to the one-stage detector RetinaNet , which improves the mAP by on VOC Clipart. The proposed D-adapt framework also surpasses both image-level1(b) and feature-level1(c) alignment as well as their combination by a considerable margin.
Results on VOC→→\rightarrowWaterColor.
As shown in Table 8, D-adapt also achieves strong performance on WaterColor dataset.
Ablations on the decouple strategy.
Further, we discuss whether the decoupling of different adaptors is useful.
In our original implementation, the input distributions of different adaptors are completely different. In the category adaptation step, we encourage the input proposals to have IoU close to or to better satisfy the low-density separation assumption. In the bounding box adaptation step, we encourage the input proposals to have IoU between and to ease the optimization of the bounding box localization task.
If different adaptors are coupled, they must share the same input distribution. Table 9 shows that only sharing the input distributions will greatly damage their respective performance. Note that different adaptors still have independent architectures. And we can conclude that the decoupling of different adaptors is quite crucial.
Ablation on bounding box adaptor.
Table 10 shows that the gain brought by box adaptation is consistent, for example when , it can still improve the mAP from to .
Appendix B More visualization results.
Figure 7-10 gives more qualitative results on Faster RCNN.