Multi-view Adversarial Discriminator: Mine the Non-causal Factors for Object Detection in Unseen Domains

Mingjun Xu, Lingyun Qin, Weijie Chen, Shiliang Pu, Lei Zhang

Introduction

The problem of how to adapt object detectors to unknown target domains in real world has drawn increasing attention. Traditional object detection methods are based on independent and identically distributed (i.i.d.) hypothesis, which assume that the training and testing datasets have the same distribution. However, the target distribution can hardly be estimated in real world and differs from the source domains, which is coined as domain shift . And the performance of object detection models will sharply drop when facing the domain shift problem.

Domain adaptation (DA) is proposed to deal with the domain shift problem, which enables the model to be adapted to the target distribution by aligning features extracted from the source and unlabeled target domains. However, the requirement of target domain datasets still limits the applicability of DA methods in reality. Domain generalization (DG) goes one step further, aiming to train a model from single or multiple source domains that can generalize to unknown target domains.

Although lots of DG methods have been proposed in the image classification field, there are still some unresolved problems. In our opinion, the common features extracted by previous DG methods are still not pure enough. The main reason is that through a single-view domain discriminator in DAL, only the significant domain style information can be removed, while some implicit and insignificant non-causal factors in source domains may be absorbed by the feature extractor as a part of common features. This has never been noticed. This implies the multi-mode structure of data and single-view domain discriminator cannot fully interpret the data. There is a piece of evidence to support our claim.

To confirm our suspicions on the domain discriminator, we designed a validation experiment. As is shown in Fig. 1, we use DANN model with DAL strategy to train a common feature extractor. When domain classifier converges, we freeze feature extractor and re-train domain classifier with a newly added residual block . We observe an interesting phenomenon: when re-trained with the newly added residual block, the domain classifier loss continues to decline. That is, some domain-specific information still exists. This phenomenon confirms our claim that in existing DG, DAL cannot explore and remove all domain specific features. This is because domain classifier only observes significant domain-specific feature in a single-view, while insignificant domain specific features in one view (space) can be significant in other views (latent spaces).

Based on the former experiment, we propose that mining common features through DAL in single-view on a limited number of domains is insufficient. By using traditional DAL, only the primary style information w.r.t. domain labels can be removed. Here we analyse this problem from the perspective of causality. As shown in Fig. 2, in a limited number of domains, the common features still contain non-causal factors such as light color, illumination, background, etc., which is expressed as the orange arrows in the figure. And such insignificant non-causal factors observed from one view may still be significant uninformative features in other latent spaces (views). So a natural idea is to explore and remove the implicit non-causal information from multiple views and purify the common features for generalizing to unseen domains.

In order to remove the potential non-causal information, we rethink the domain discriminator in DAL and propose a multi-view adversarial domain discriminator (MAD) that can observe the implicit insignificant non-causal factors. In our life, in order to get the whole architecture of an object, we often need to observe it from multiple views/profiles. A toy example is shown in Fig. 3 (left part). When we observe the Penrose triangle from one specific view, we might misclassify it as a triangle, ignoring that it might also appear to be L from another perspective. Following this intuition, we construct a Multi-View Domain Classifier (MVDC) that can discriminate features in multiple views. Specifically, we simulate multi-view observations by mapping the features to different latent spaces with auto-encoders , and discriminate these transformed features via multi-view domain classifiers. By mining and removing as many non-causal factors as possible, MVDC encourages the feature extractor to learn more domain-invariant but causal factors. We conduct an experiment based on MVDC and show the heatmaps from different views in Fig. 3 (right part), which verifies our idea that different noncausal factors can be unveiled in different views.

Although the Multi-View Domain Classifier can remove the implicit non-causal features in principle, it still implies a sufficient diversity of source domains during training. So we further design a Spurious Correlation Generator (SCG) to increase the diversity of source domains. Our SCG generates non-causal spurious connections by randomly transforming the low-frequency and extremely high-frequency components, as points out that in the spectrum of images, the extremely high and low frequency parts contain the majority of domain-specific components.

Combining MVDC and SCG, the Multi-view Adversarial Discriminator (MAD) is formalized. Cross-domain experiments on six standard datasets show our MAD achieves the SOTA performance compared to other mainstream DGOD methods. The contributions are three-fold:

1. We point out that existing DGOD work focuses on extracting common features but fails to mine and remove the potential spurious correlations from a causal perspective.

2. We propose a Multi-view Adversarial Discriminator (MAD) to eliminate implicit non-causal factors by discriminating non-causal factors from multiple views and extracting domain-invariant but causal features.

3. We test and analyze our method on standard datasets, verifying the effectiveness and superiority of our method.

Related Work

Object detection is a critical problem in computer vision, aiming to locate and classify the specified instances in specific images. Modern object detection methods can be divided into two categories: one-stage methods and two-stage methods . However, traditional object detection methods suffer from domain shifts in practical applications. In order to alleviate the performance degradation caused by domain shift, lots of domain adaptive object detection (DAOD) methods are presented . DAOD methods are trained with labeled source domains and unlabeled target domains, and alleviate the domain shift problem by DAL. The DAOD methods can be divided into two parts: adversarial-based methods and reconstruction-based methods. For the former, the domain adversarial learning structure is introduced to align feature maps by . For the latter, firstly uses CycleGAN to generate pseudo samples that are similar to the target domain from the source domain samples. DAOD methods still have problems in real-world applications. On the one hand, they still require additional effort to collect unlabeled target domain datasets, which is expensive and even impossible. On the other hand, they cannot guarantee the causality of features. We hope to find domain-invariant but causal features that are more robust for unseen target domains.

2 Domain Generalization

Domain generalization has been studied for a long time in the image classification field. Existing domain generalization methods can be divided into the following three categories. First, domain augmentation methods aim to increase the diversity of source domains by transferring images to new domains. augment source domains in image-space. and perform augmentation on the frequency spectrum. Second, representation learning methods aim to extract domain-invariant representation from source domains. firstly adopts the idea of DAL in domain adaptation for domain generalization. Third, there are also learning strategies like which firstly adopts meta-learning for domain generalization, following the idea of enabling the network learn how to learn domain-invariant components from different domains.

3 Causal Mechanism

Methods based on causal mechanisms consider that the prediction based on statistical dependence is unreliable, because the statistical correlations contain both spurious non-causal correlations and causal correlations. For example, smoking, yellow teeth and lung cancer are closely related. Nevertheless, only smoking is the causal factor of lung cancer. To improve the generalization of methods, they try to mine these invariant causal correlations. In recent years, solving DG problems by finding causal factors is gaining more and more attention. Some methods attempt to obtain invariant causal mechanisms . Meanwhile, other methods try to recover causality characteristics . Existing methods focus on looking for invariant causal factors. However, we argue that one should pay more attention to exploring the potential non-causal spurious correlations, because the domain-invariant representations learned by traditional DAL are often biased towards one view, as Fig. 3 shows. We propose to purify the domain-invariant features by removing implicit non-causal factors from multiple views in DAL.

Proposed MAD Approach

Existing DG methods learn common features with conventional DAL among finite domains . However, such common features extracted are often not pure due to the implicit non-causal factors. As discussed before, we propose a Multi-view Adversarial Discriminator (MAD) to explore and remove potential spurious correlations and encourage the model to extract purer domain-invariant but causal features. As is shown in Fig. 4, our MAD contains two new parts. First, a Spurious Correlation Generator (SCG) module is designed to increase the diversity of source domains and make the potential non-causal factors more significant. Second, a Multi-View Domain Classifiers (MVDC) module is designed to identify the non-causal factors for both image and instance levels, such that the domain adversarial learning is more sufficient and the non-causal factors are richer in different views, which instructs the feature extractor to ignore them. To summarize, SCG explores and exposes the potential non-causal factors, while MVDC discriminates and removes them.

We make the following definitions to formalize the domain generalization problem. The source domain is denoted as Ds={Xs,Ys}D_{s}=\{X_{s},Y_{s}\}. The feature extractor f(⋅)f(\cdot) can extract the features S=f(Xs)S=f(X_{s}) from input images XsX_{s}. Feature SS contains causal and non-causal factors {scau,snon}\{s_{cau},s_{non}\} in the finite source domains. Intuitively, not all common components scoms_{com} are causal factors, but domain private components spris_{pri} are non-causal factors, and there is,

The non-causal factors snons_{non} are supposed to obey Gaussian distribution, i.e., snon∼N(μ,σ2)s_{non}\sim\mathcal{N}(\mu,\sigma^{2}) .

2 Spurious Correlations Generator

where u,vu,v denotes the position of the spectrum and r(RL,RH)r(R_{L},R_{H}) denote the cut-off frequency of low and high frequency. We then randomize this non-causal factor SS according to a Gaussian distribution as RG(S)=S⋅(1+N(0,1))R_{G}(S)=S\cdot(1+\mathcal{N}(0,1)). Finally, we get the augmented image x^\hat{x} with potential non-causal factors by adopting the Inverse Discrete Cosine Transform F′(⋅)\mathscr{F}^{{}^{\prime}}(\cdot) to the augmented spectrum. Our spurious correlations generator can be expressed as:

3 Multi-View Domain Classifier

DAL is a standard method to extract the common feature of different domains, which minimizes the A−Distance\mathcal{A}-Distance of the extracted features between different domains. DAL is a minimax optimization problem between feature extractor F\mathcal{F} and the ideal domain classifier:

where H\mathcal{H} denotes a hypothesis set of all possible domain classifiers, h(⋅)h(\cdot) is one of the domain classifiers in H\mathcal{H}, and e(⋅)e(\cdot) denotes encoders which map feature to divers latent spaces. A single hh depends on the most discriminative domain private features, so it ignores the insignificant domain specific components of the features and incorrectly takes such non-causal components as common features.

We, therefore, propose to improve the sensitivity of the domain classifier to potential non-causal factors by extending DAL to more views. Specifically, our MVDC can map features into multiple latent spaces with encoders eie_{i} and then discriminate features in each space with an independent domain classifier hih_{i}. These domain classifiers encourage the feature extractor F\mathcal{F} to ignore the implicit non-causal factors and learn domain-invariant but causal features.

3.2 Classifier Structure

Fig. 6 shows the structure of one branch of the MVDC, which represents one of the multiple views to observe the features. The complete structure of MVDC contains MM branches for image-level features and MM branches for instance-level features respectively. Each branch of MVDC contains an auto-encoder and a classifier in structure.

The encoder and decoder are the basic network structure of each branch, which map features into different latent spaces to show different profiles of the feature. The encoder part aims to compress the features and map them into different latent spaces. Then the latent features are fed into an independent domain classifier. Meanwhile, in order to ensure the semantic content invariance of features, the latent features are mapped back into the original space through a subsequent decoder.

To explore the non-causal factors hidden in the whole image and each instance, we make different designs on multi-view domain classifiers for image-level and instance-level respectively. For the image-level, we focus on the global non-causal factors of an image, such as the illumination, color and background texture. These global non-causal factors are similar across the image, so we use the convolutional layers to construct the encoder and decoder. In each branch, we use dilated convolution with different dilation rates to extract different non-causal factors of domains. For the instance-level, we use fully connected layers to mind more semantic non-causal factors like the camera angle of each instance.

3.3 Loss Function

For training the object detector with MAD approach, several loss functions are introduced in the following text.

First, a reconstruction loss is used to ensure that the semantics of features are not changed by the encoder. The mapped feature e(s)e(s) should contain the semantic information required to reconstruct the original feature ss. Only if the semantic information is guaranteed to be complete, the subsequent domain classifier is meaningful. So we use MSE loss to constrain the distance between the original feature ss and the reconstructed feature g(e(s))g(e(s)). The reconstruction loss can be described as:

Second, the adversarial domain classifier loss is used to ensure the mapped features are domain distinguishable (inner optimization) and domain confused (outer optimization). We use Cross-Entropy loss to adversarially train the K-domain classifiers in total MM branches. And the domain label in kthk^{th} domain is denoted as yky_{k}.

The third constraint is the most critical view-different loss, which ensures the auto-encoders to map features into diverse latent spaces. Therefore, we propose to enlarge the feature difference between latent spaces (views), such that the insignificant non-causal factors become significant. So we construct the following MSE loss of each feature pair from MM different latent spaces.

The fourth constraint is used to ensure the consistency of the results in these 2M2M branches of two levels. For each pair of image-level and instance-level branches, we adopt l2l_{2} distance between the average value of image-level predictions pi(u,v)p_{i}^{(u,v)} and each instance-level prediction pj,np_{j,n} as the consistency constraint. We suppose the feature map of each image contains ∣I∣|I| pixels and NN instances in total. Then, the consistency loss of the whole model can be described as:

The MVDC loss in both image level and instance level can be presented as:

Then, we obtain the overall loss of MAD by trade-off the object detector loss and MVDC loss with λ\lambda as:

Experiments

We adopt seven cross-domain object detection benchmark datasets, which will be introduced below. Cityscapes dataset mainly contains daytime scenery in the streets and Foggy Cityscapes and Rain Cityscapes are datasets of synthesized images with different weather conditions based on the depth information from the Cityscapes. SIM10k dataset contains rendered images of rendered 3D models. KITTI is an autonomous driving dataset. PASCAL VOC dataset is collected from the real world. The BDD100k dataset is a large-scale dataset for autonomous driving. We abbreviate {SIM 10k, Cityscapes, Foggy Cityscapes, Rain Cityscapes, BDD100k, KITTI, PASCAL VOC} as {S, C, F, R, B, K, V} respectively in the following text.

2 Experimental Setup

Implementation Details. First, to verify the effectiveness of our MAD method, we conduct cross-test experiments on {C, F, R, B}, which means we train a model on one of these datasets and test it on the rest datasets. For each source and target pair, we only calculate the result on the intersection of their label space. To uniform the annotation styles of these datasets, we regard labels {motor, motorcycle and motorbike} as {motor} and labels {bike and bicycle} as {bike}.

Second, to verify the superiority of our MAD method, we compare with the existing DGOD methods like MLDG , DIDN , FACT and FSDR . Several DAOD methods such as DAF , SW-DA , SC-DA , MTOR , GPA are also compared under the task from cityscapes to foggy cityscapes. We train a total of 10 epochs. In training process, we set the initial learning rate to 0.002, and start to attenuate the learning rate to 0.0002 at the 7th epoch to make the model converge better. In our experiments, we train the models with MindSpore and PyTorch frameworks. Our code is available at github.com/K2OKOH/MAD. Mean average precisions (mAP) with a IoU threshold of 0.5 is reported.

Baseline. We build our method on the basis of FasterRCNN framework with vgg16 pre-trained on ImageNet as the backbone and adopt the Stochastic Gradient Descent (SGD) as the optimization method.

3 Results and Discussion

The results in Tab. 1 show that our method can achieve better results in most cross-domain scenarios. Trained with limited number of source domains, our SCG method can add non-causal factors in more directions to the existing images, which better simulate the potential target domain distribution. This makes our method superior to MLDG, FACT and FSDR that extract features over finite known source domains or fixed augmented domains. Comparing the single-view DANN with our MAD, we can also find that our MAD performs better, which shows that our Multi-View domain classifier can further help mine and remove non-causal factors from the simulated target distribution.

In Tab. 2, we compare our MAD with mainstream DG and DA methods. Under both single-source and multi-source DG settings, our method has the best generalization ability among DG methods and exceeds Multi-Source methods in most categories. Furthermore, our target-free MAD can even surpass some of the domain adaptive methods, which are trained with unlabeled target images.

We further conduct experiments on the common categories car in six datasets (C, F, R, S, K, V and B) to verify the domain generalization ability of our MAD. As shown in Tab. 3, our method also performs the best in most unseen target domains.

As is shown in Fig. 7, we also perform visualization of feature distribution via t-SNE under the task from C to F. (a) shows the feature distribution of cars in datasets C and F extracted with the original Faster RCNN model, in which the difference of distribution between domains is clear. (b) shows the feature distribution of the same datasets and category extracted by DANN, from which we can see that DANN can align the distribution of different domains in a single view. However, as shown in (c), the aligned feature distributions by DANN are still separated with multi-view discriminators by our MVDC, which means that DANN can only remove the significant non-causal factors and the remained insignificant non-causal factors are still domain discriminative. Compared to (c), (d) shows the feature distribution in multi-view extracted by our MAD, and we can see that our MAD can indeed map features into different spaces and well-align different domains under each view.

The multi-view adversarial discriminator is substantially orthogonal to the computer vision tasks and thus can also be applicable to DG-based image classification tasks. Therefore, we conduct single-source DG experiments on the widely used PACS and VLCS datasets, and compared with ERM and DANN frameworks, as shown in Tab. 4. The results show the effectiveness of our MAD.

4 Ablation Study

We conducted an ablation study on our MAD methods to verify the validity of each part. Our method can be divided into four parts in total, namely spurious correlations generator (SCG), image-level and instance-level multi-view domain classifier (IMG, INS) and the consistency constraints (CST). We study the contribution of each part by adding them sequentially and observing the change in mAP performance. We train MAD on domain C and test it on other domains F, R, B to conduct ablation experiments.

Tab. 5 reflects the effectiveness of each part of our MAD. By introducing SCG, potential spurious correlations are injected into the network. The MVDC consisting of three submodules (IMG, INS, CST) further mines and removes insignificant spurious correlations in the domains. Specifically, the image-level adversarial submodule (IMG) eliminates overall non-causal factors, and the instance-level submodule (INS) eliminates the semantic non-causal factors in each instance. The consistency loss (CST) ensures the consistency of the domain discriminators in two stages.

5 Hyper-parameters Analysis

We tested two hyper-parameters of our MAD method.

First, the number MM of views is the key hyper-parameter in MAD. More views lead to better performance, but too many auto-encoders increase model complexity with diminishing marginal effect. As we can see in Fig. 8 (a), we found that performance improved until M=5M=5 and then converged with further views until M=8M=8. Thus, we set M=3M=3 as a balance between performance and cost.

Second, the trade-off coefficient λ\lambda of the domain adversarial loss in Eq. 10 is used to balance the main task of object detection and the MVDC part. We take several values from 0.05 to 0.2 for testing. As can be seen from Fig. 8 (b), we set λ=0.1\lambda=0.1 in MAD for all the experiments.

Conclusion

This paper analyzes the problem of domain adversarial learning (DAL) from the perspective of causal mechanisms. We point out that existing DG methods fail to remove potential non-causal factors implied in common features, because DAL is biased by the single-view nature of the domain discriminator. To overcome this problem, we propose a Multi-view Adversarial Discriminator (MAD) to learn domain-invariant but causal features. Our MAD includes an SCG that generates potential spurious correlations to diversify the source domains and an MVDC that constructs multi-view domain classifiers to remove implicit non-causal factors in latent spaces. Finally, MAD purifies the domain-invariant features and the causality is augmented. Extensive experiments on benchmarks for cross-domain object detection verify the generalization ability to unseen domains.

This work was partially supported by National Natural Science Fund of China (62271090), National Key R&D Program of China (2021YFB3100800), Chongqing Natural Science Fund (cstc2021jcyj-jqX0023), CCF Hikvision Open Fund (CCF-HIKVISION OF 20210002), CAAI-Huawei MindSpore Open Fund, and Beijing Academy of Artificial Intelligence (BAAI).

References