Causality-inspired Single-source Domain Generalization for Medical Image Segmentation

Cheng Ouyang, Chen Chen, Surui Li, Zeju Li, Chen Qin, Wenjia Bai, Daniel Rueckert

I Introduction

Deep learning based medical image segmentation approaches usually achieve state-of-the-art performance when being trained and tested on datasets from a single domain, i.e. from identically distributed training and testing data. However, in practice, deep learning models perform less well when the testing data is drawn from a different distribution than the training data (i.e. a different domain) . The discrepancy between training and testing domains is termed as the domain shift . In medical image segmentation, the most notorious source of domain shift is the differences in image acquisition (imaging modalities, scanning protocols, or device manufacturers) . This type of domain shift is therefore termed as the acquisition shift . We argue that the performance deterioration under acquisition shift can be attributed to the following two mechanisms: the shifted domain-dependent features and the shifted-correlation effect. Shifted domain-dependent features: Domain-dependent features include intensities and textures, which constitute image appearance. Deep networks are susceptible to shifts in intensity/texture . This is in contrast to human annotators: they can easily find correspondence of the same anatomical structure across domains , usually by looking at the shape that is domain-invariant and intuitively causal to human-defined segmentation masks, compared with intensity/texture. Shifted-correlation effect: Due to a confounder (i.e. a “third” variable that spuriously correlates two variable-of-interests) , objects in the background might be correlated but not causally related to the object-of-interests . The network might take these objects in the background as clues for recognizing the object-of-interests . For example, in , a model that recognizes pneumonia in X-ray images is actually looking at the hospital mark in the background, which correlates with pneumonia due to the confounder: data selection bias. These correlations are often detrimental under domain shift. This is because decision rules based on these correlations may break in the shifted domain: the correlated objects in the background may disappear, or it may not co-shift in the same way as the object-of-interests. In the above example, the model fails on real-world images where hospital mark do not correlate to pneumonia.

To mitigate domain shift, previous attempts include unsupervised domain adaptation (UDA) and multi-source domain generalization (MDG) approaches . Unfortunately, UDA or MDG may not always be practical: they rely on training data from the target domain or from multiple source domains, which are often unavailable due to cost or privacy concerns. UDA also requires expertise for fine-tuning on target data, incurring difficulties to its deployment in the real world.

A more practical setting is single-source domain generalization: to train a deep learning model to be robust against domain shifts, using training data from only one source domain. Since no examples of the target domain are available, we resort to bottom-up approaches that are built on the above causal analysis of acquisition shift. We aim to 1) steer the network towards shape information which is domain-invariant and is intuitively causal to segmentation results; 2) to immunize the segmentation model against the shifted-correlation effect, by removing the confounder that spuriously correlates objects in the background and the object-of-interests during training.

Learning causal features and removing confoundings usually require intervention: fixing the variable-of-interest while incorporating other variables in a fair way . This is similar to randomized controlled trials. To mitigate acquisition shift, a straightforward intervention is to train the model with images of a fixed cohort of patients that are taken under all possible acquisition processes. Since this is unrealistic, we resort to data-augmentation-based intervention that incorporates possible acquisition shifts via simulation.

In this work, we propose a causality-inspired data augmentation approach for single-source domain generalization. It exposes the network to synthetic acquisition-shifted training samples that incorporate shifts in intensity/texture and shifted correlations. Specifically, to efficiently synthesize diverse appearances (intensities and textures) without losing generality, we employ shallow convolutional networks with random weights that are sampled at each training iteration to augment images. As stable decision rules can hardly be formed on constantly-varying intensities/textures, the network would resort to domain-invariant features such as shapes. To remove the confounder that leads to the shifted-correlation effect, we first reveal that the image acquisition process naturally confounds objects in the background and the object-of-interests, in terms of their appearances. We then design a practical method for simulating and independently sampling the appearances of potentially confounded objects durining training. This is achieved by applying different appearance transformations in a spatially-variable manner, with the help of pseudo-correlation maps computed using unsupervised algorithms. The overall approach is used as additional stages following standard data augmentations. It is therefore generic to architectures of segmentation networks. In summary, we make the following contributions:

We investigate single-source domain generalization problem for cross-domain medical image segmentation from a causal view. We propose a simple and effective causality-inspired data augmentation approach.

We propose 1) global intensity non-linear augmentation (GIN) technique that efficiently transforms images to have diverse appearances via randomly-weighted shallow convolutional networks; 2) interventional pseudo-correlation augmentation (IPA) technique that removes the confounder that leads to the shifted-correlation effect. This is realized by independently resampling appearances of potentially confounded objects. These two components function as cores of the proposed approach.

We build a comprehensive testing environment for single-source domain generalization for cross-domain medical image segmentation. It covers cross-modality, cross-sequence (MRI) and cross-center settings with various anatomical structures. We hope this testing environment to facilitate future works on domain robustness for medical image segmentation.

II Related works

Considerable efforts have been made to alleviate domain shift for deep networks. Unsupervised domain adaptation (UDA) transfers a model trained on a source domain to a target domain using unlabeled target-domain data . Existing techniques are mainly based on distribution alignment , or self-training . Multi-source domain generalization (MDG) usually learns domain-invariant features in an one-off manner from multiple source domains. Recent techniques include meta-learning, style transfer , transfer learning, dynamic networks and so on. In medical imaging, recent methods based on meta-learning have achieved promising results. However, for both UDA and MDG, target domain data or multi-source data is often unavailable due to privacy and/or cost concerns in medical settings.

Single-source domain generalization requires training data from one domain only. Recent works such as propose to remove features that cause the largest loss gradients, which are believed to be domain-dependent. Liu et al. propose to unify statistics of image features, which control image styles , among images from different domains. A major stream of works exploits data augmentation: training the network on deliberately perturbed samples to improve network robustness to real-world perturbations . We review these stream of methods later with more details. Some most recent works such as employ multiple techniques: data augmentation , adversarial training and contrastive learning .

II-B Data augmentation for domain robustness

Theoretical analysis suggests that data augmentation improves generalization by enlarging the span of the data and by regularizing decision boundaries. In practice, Cutout strengthens robustness against feature missing caused by domain shift, by partially occluding training images. Mixup regularizes decision boundaries by interpolating among training samples. RandConv drives the network to learn shape information, which is domain-invariant, by randomly altering image textures using linear filtering. Our method is closely related to RandConv . However, we show in our experiments that the linear filtering in RandConv is oversimplified for accounting for domain gaps that occur in real-world settings. Adversarial data augmentation generates image perturbations that easily flip predictions of classifiers .

In medical imaging, Zhang et al. employ a stack of photometric and geometric transformations to training images to improve domain robustness. Billot et al. propose a contrast-agnostic brain MRI segmentation strategy, which synthesizes training examples by sampling from pre-built generative models of brain images. However, this method necessities well-defined generative models from segmentation labels to images, which is usually unavailable in most of medical imaging applications. Furthermore, these generative models are often oversimplified. AdvBias is specially designed for medical image segmentation. It employs an adversarial augmentation technique based on a multiplicative bias field model. It outperforms a series of competitive methods on cross-center MRI segmentation .

II-C Leveraging causality for robust deep learning

As discussed in Sec. I, learning causalities and mitigating confoundings usually require causal intervention . In causal intervention, the variable-of-interest in a causal relationship is fixed, while other variables are fairly incorporated. A model then learns causalities from these intervened samples. For example, a model that recognizes a camel might be mistakenly focusing on the background: desert , since most of pictures of camels are taken in deserts. In this case, intervention can be done by incorporating different backgrounds: training with pictures of camels that are taken in diverse backgrounds like grassland and city. By this mean the model would learns that it is the camel rather than the desert that cause a camel label. Causal relationships are usually modeled using structural causal models (SCM). The fixing operation is usually noted as do(⋅)do(\cdot) . The distribution p(Y∣do(X=camel))p(Y|do(X=\texttt{camel})), is called interventional distribution. Compared with conditional distribution p(Y∣X=camel)p(Y|X=\texttt{camel}) that reflects correlation in the observed data, p(Y∣do(X=camel))p(Y|do(X=\texttt{camel})) reflects causation .

Causal ideas have been used for discovering image features that are semantically essential and robust . Invariant risk minimization learns causal image representations by enforcing these representations to be Bayesian optimal in all environments. Mahajan et al. improve domain robustness using contrastive losses. Atzmon et al. propose a causal mechanism to generalize a model to novel samples with unseen combinations of attributions.

Our idea of using data-augmentation-based intervention is inspired by . proves that post-hoc data augmentation theoretically commutes with ”physical” intervention. Mitrovic et al. derive a practical loss function for causality-based domain generalization. Different from , we focus on the unanswered practical problem of designing an augmentation model tailored to the real-world problem: cross-domain medical image segmentation. Our work is also related to causal weakly-supervised segmentation by Zhang et al. , as both works study the adverse effect of confoundings among objects on image segmentation. While Zhang et al. focus on the intra-domain scenario, we focus on the effect of confoundings under domain shift.

III Method

Our causality-inspired data augmentation approach aims to improve network robustness against domain shift, in particular, shifts caused by the differences in acquisition processes . Based our analysis in Sec. I, we propose to expose the network to training examples that incorporate simulated intensity/texture shifts and shifted correlations among objects.

Specifically, our approach is a synergy of a global intensity non-linear augmentation technique (GIN) and an interventional pseudo-correlation augmentation technique (IPA). As shown in Fig. 1-B, GIN transfers training images to have diverse appearances while keeping the shapes of anatomical structures unchanged, discouraging the network from biasing towards appearances. IPA resamples possible appearances of potentially spuriously correlated objects (due to confounding) in the background and the object-of-interests, in independent and diverse manners. This is implemented as spatially-variable blending between two GIN-augmented images. The entire approach functions as additional steps in a standard data augmentation pipeline.

In the following sections, we first introduce the general problem formulation of the proposed data augmentation approach. We introduce GIN in Sec. III-B, with a detailed reasoning behind its design-of-choices. IPA is introduced in Sec. III-C, where we firstly reveal that it is the image acquisition process that naturally confounds objects in the background and the object-of-interests. We then describe how IPA removes confoundings. Finally, we summarize the overall training process.

A causal view of image generation and segmentation: We first introduce the problem formulation of our data-augmentation-based single-source domain generalization approach. Inspired by recent works , we model the data generation process and the (ideally, domain-invariant) segmentation process using the causal model shown in Fig. 1-A. Specifically, we make the following assumptions:

1) A→X←CA\rightarrow X\leftarrow C : An image XX is generated from two independent variables (factors): acquisition AA and content CC. CC represents the shapes of underlying anatomical structures of the patient, while AA represents the acquisition process. The factor AA maps different types of tissues of the patient in the scanner into different pixel values in the image.

2) C→S→YC\rightarrow S\rightarrow Y: There exists an ideal domain-invariant representation SS, determined by CC. SS is in the form of feature maps of the deep layers of the network and it contains the shape information of the object-of-interests. The ground-truth segmentation mask YY can be derived from SS.

3) X→fϕ(X)X\rightarrow f_{\phi}(X): The segmentation network fϕ(⋅)f_{\phi}(\cdot) takes XX as an input and predicts YY, by implicitly estimating SS.

AA and CC are independent: changing AA does not affect CC, SS and YY. Of note, our discussion is constrained to acquisition shift. CC is assumed to be unchanged across the source domain and the target domain in our experiments.

Causal intervention for domain robustness: According to , our argument that SS to be invariant to shifts of AA can be formally written as follows:

Here p(Y∣S,do(A=ai))p(Y|S,do(A=a_{i})) denotes the distribution raised by letting images to be generated from a specific acquisition process A=aiA=a_{i} , for example MRI. Using a symmetric notation, we let aja_{j} to be another acquisition process, for example CT. Eq. 1 suggests that ideally, this distribution should remain the same regardless of the acquisition processes.

In practice, as shown in Fig.1-A, we use a segmentation network fϕ(⋅)f_{\phi}(\cdot) parameterized by ϕ\phi to predict YY. To make fϕ(⋅)f_{\phi}(\cdot) domain invariant, we implicitly estimate SS in the last layers of fϕ(⋅)f_{\phi}(\cdot). However, the condition in Eq. 1 cannot be directly used to train fϕ(⋅)f_{\phi}(\cdot), as “physical” interventions on AA (scanning patients under all possible acquisition processes) is impractical. Fortunately, Ilse et al. have demonstrated that data augmentations can be used as “virtual” causal interventions. Therefore, we assume that for each aia_{i}, there exists a photometric transformation function Ti(⋅)\mathcal{T}_{i}(\cdot) being able to transform the image XX to be like from aia_{i} Arguably this transformation can be interpreted as generating counterfactual examples : given an observed image from a certain acquisition process, asking how this image would look like, had it been generated from another imaging process. We do not label our data augmentation as counterfactual since the aim of our approach is not to synthesize the appearance of a specific domain.. We therefore have:

Combining Eq. 1 and 2 leads to a practical domain invariance condition: minimizing the difference between distributions raised by different photometric transformations Ti(⋅)\mathcal{T}_{i}(\cdot) and Tj(⋅)\mathcal{T}_{j}(\cdot) . By combining this domain invariance condition and the image segmentation loss, we can now derive our loss function (inspired by ). For each iteration, we have

Here (x,y)∼p(X,Y)(\mathbf{x},\mathbf{y})\sim p(X,Y) denotes an image-label pair in the training dataset, Lseg(⋅,⋅)\mathcal{L}_{seg}(\cdot,\cdot) is a segmentation loss, e.g. cross-entropy; Ti(⋅)\mathcal{T}_{i}(\cdot) and Tj(⋅)\mathcal{T}_{j}(\cdot) are two different photometric transformations that simulates the effect of aia_{i} or aja_{j} respectively, randomly sampled from a family of photometric transformations at each iteration; D(⋅∥⋅)\mathcal{D}(\cdot\|\cdot) is the Kullback–Leibler divergence: measuring differences between two distributions; λdiv\lambda_{div} is a weighting coefficient. Similar divergence terms have also been used for semi-supervised learning , although they are not derived from a causal perspective.

As implied by Eq. 3, the photometric transformations {T(⋅)}\{\mathcal{T}(\cdot)\} serve as the core of the domain invariance condition. Since the target-domain data is unavailable, we resort to build {T(⋅)}\{\mathcal{T}(\cdot)\} in a bottom-up manner, based on our analysis on ingredients of acquisition shift, as discussed in Sec. I. We simulate {T(⋅)}\{\mathcal{T}(\cdot)\} as a combination of GIN and IPA.

III-B Global intensity non-linear augmentation

where ∥⋅∥F\|\cdot\|_{F} is the Frobenius norm.

By this mean, at each iteration, different intensities and textures are given to training images. As stable decision rules can hardly be built on randomly changing intensities/textures, the network would resort to invariant information like shapes.

Advantages: We highlight major advantages of the above configurations: Firstly, GIN is based on generic assumptions on intensity/texture transformations. It therefore avoids being over-specific to a certain target domain(s). In addition, GIN is computationally efficient: it is in the form of shallow networks and therefore easily exploits GPUs for acceleration. Also, it is differentiable, and therefore can be be integrated into adversarial augmentation frameworks to improved data efficiency.

III-C Interventional pseudo-correlation augmentation

Confounded objects in the background affect segmentation: Recall in Sec. I, spurious correlations (due to confoundings) between objects in the background and the object-of-interests in the source domain might be taken by a segmentation network as domain-specific clues for making predictionsIn our preliminary experiment, we verified the existence of decision rules that are based on the background, in medical image segmentation: We distorted the backgrounds of images by randomly swapping patches of the background (therefore the global image statistics would remain unchanged.). We then tested a segmentation network with these background-distorted images, and have observed substantial performance downgrade compared with results on the original undistorted images. This phenomenon has also been verified in general computer vision and these spurious correlations are sometimes termed as context bias .. These decision rules may break in the target domain, leading to performance downgrade. This is because confounded objects in the background that benefit segmentation in source domain, might not exist in the target domain. Alternatively, they might not co-shift in the same way as the object-of-interests.

From the perspective of network architectures, objects in the background often affect predictions of the object-of-interests by the following ways: 1) Background features can affect global feature statistics at normalization layers , since feature statistics are usually calculated across all spatial locations. 2) The large receptive fields often make the pixels of the object-of-interests and those of the neighboring background to be inevitably perceived and processed together .

The acquisition process naturally confounds objects: To mitigate the shifted-correlation effect, it is worthwhile to point out that it is the acquisition factor AA that leads to confounding. It naturally creates spurious correlations between certain objects in the background and the object-of-interestFor the ease of illustration, in the following analysis, we focus on the scenario where only one object-of-interest is available. Our conclusion naturally holds for multi-class segmentation as well, and has been validated by our experiments on abdominal and cardiac segmentations.. To demonstrate this, in an image XX, we consider the patch of object-of-interest XfX_{f} and the patch of a potentially correlated unlabeled object XbX_{b} in the background. We zoom-in the causal relations in Fig. 1-A using XfX_{f} and XbX_{b}, and redraw that in Fig. 3-A. We can see:

1) Xf→fϕ(X)←XbX_{f}\rightarrow f_{\phi}(X)\leftarrow X_{b}: Although XfX_{f} already contains sufficient information for delineating YY, in practice XbX_{b} also affects the network features and the output, as both XfX_{f} and XbX_{b} are processed by fϕ(⋅)f_{\phi}(\cdot).

2) fϕ(X)←Xb←A→Xff_{\phi}(X)\leftarrow X_{b}\leftarrow A\rightarrow X_{f}: More importantly, the confounding effect of AA that correlates XbX_{b} and XfX_{f} is revealed in the path Xb←A→XfX_{b}\leftarrow A\rightarrow X_{f}. This corresponds to the fact that given a certain acquisition process, the same imaging physical mechanism that maps different tissues to different pixel values, applies to both the object-of-interest and the objects in the background. Without such a path, appearances of XfX_{f} and XbX_{b} would vary independently in the training dataset. Stable correlations regarding their appearances could unlikely be established and learned.

Of note, as we assume the content factor CC remain unchanged across domains, we ignore the confoundings caused by CC, and omit CC in Fig. 3.

Removing confounding by intervention: We propose to mitigate the shifted correlation effect during the training stage, by removing the confounding Xb←A→XfX_{b}\leftarrow A\rightarrow X_{f} using the intervention do(Xf=xf)do(X_{f}=\mathbf{x}_{f}) . This operation resamples the appearances of correlated objects XbX_{b}’s, independent of XfX_{f}. This intervention in effect removes A→XfA\rightarrow X_{f}. Formally, we are learning the interventional distribution p(Y∣do(Xf=xf))p(Y|do(X_{f}=\mathbf{x}_{f})) based on the intervened causal diagram Fig. 3-B:

Here xf,xb∈x; (x,y)∼p(X,Y); a∼p(A)\mathbf{x}_{f},\mathbf{x}_{b}\in\mathbf{x};\ (\mathbf{x},\mathbf{y})\sim p(X,Y);\ a\sim p(A) and p(A)p(A) is a prior of possible acquisition processes. Eq. 5 translates to independently sampling possible appearances of XbX_{b}.

Unfortunately, to compute Eq. 5 we are faced with three practical issues: 1) we do not know which object is correlated with the object-of-interest and there is no ground-truth map of it; 2) there might be more than one objects in the background that correlates with the object-of-interest, and their effects might be entangled; 3) directly fixing xf\mathbf{x}_{f} using ground-truth masks y\mathbf{y} would make xf\mathbf{x}_{f} unnaturally stand out from the background, providing shortcuts for the network to recognize xf\mathbf{x}_{f}. Also, this intervention cannot be realized by GIN alone: GIN’s transformation functions are spatially-invariantIf two objects share the same appearance, their appearances would remain the same after being transformed by GIN..

Spatially-variable blending: As a practical solution, we employ interventional pseudo-correlation augmentation (IPA), which approximates the intervention do(Xf=xf)do(X_{f}=\mathbf{x}_{f}). IPA is built on appearance transformations of GIN. We use pseudo-correlation maps as surrogates of label maps of xb\mathbf{x}_{b}’s, for allocating transformation functions to different pixels of the image: Pixels that correspond to different values in the pseudo-correlation map would be given different transformation functions. Pseudo-correlation maps are generated using the unsupervised algorithm . To account for different potential spurious correlations, we use different randomly-sampled maps at each iteration. To avoid the shortcuts caused by fixing xf\mathbf{x}_{f} using y\mathbf{y}, we apply pseudo-correlation maps to both xb\mathbf{x}_{b}’s and xf\mathbf{x}_{f} (i.e. to the entire image).

III-D Training objective

The overall training is end-to-end using the loss function described in Eq. 3. For the ease of implementation we let the output of fϕ(⋅)f_{\phi}(\cdot) to be in the form of raw logits, and let p(y∣fϕ(⋅))p(\mathbf{y}|f_{\phi}(\cdot)) to be the probabilities obtained by passing the output of fϕ(⋅)f_{\phi}(\cdot) to a softmax function. For Lseg\mathcal{L}_{seg}, we employ a sum of multi-class cross-entropy loss and soft Dice loss. We set the weighting coefficient λdiv\lambda_{div} in Eq. 3 to be 10.0, same as in . After training, the segmentation network fϕ(⋅)f_{\phi}(\cdot) is ready to be directly applied to unseen testing domains. The overall algorithm flow is summarized in Algorithm 1.

IV Experiments

The proposed approach is evaluated in three cross-domain settings: 1) cross-modality abdominal segmentation from CT to T2-SPIR MRI (Abdominal CT-MR), 2) cross-sequence cardiac segmentation from bSSFP MRI to LGE MRI (Cardiac bSSFP-LGE), and 3) prostate segmentation on MRI across six centers (Cross-center Prostate). Details of the datasets and the source-target splits are summarized in Table I. All datasets are originally in 3-D and have been reformatted to 2-D, then resized to 192×\times192, and padded along the channel dimension to fit into the network. For the abdominal CT dataset, we applied a window of in Housefield values. For all MRI images, we clipped the top 0.5% of the histograms. We normalized all the 3-D scans to have zero mean and unit variance. For fairness of comparisons, for all the methods evaluated (including ERM), conventional geometric augmentations: affine transformations and elastic transformations; and intensity augmentations: brightness, contrast, gamma transformations and additive Gaussian noises were applied by default.

We employed the commonly-used Dice score (0-100) as the evaluation metric for measuring the overlap between the prediction and the ground truth. For abdominal and prostate segmentations, for the source domain, we used a 70%-10%-20% split for training, validation and testing sets; for the target domain(s), we used all the images for testing, same as in . For cross-center prostate segmentation, each time we took one domain as the source and the rest five domains as targets, and we computed Dice scores averaged by target domains. This 1-versus-5 experiment is repeated for each of all six domains. For cardiac segmentation, we employed the same data split as in the cross-sequence segmentation challenge .

IV-B Network architecture and training configurations

We configured the segmentation network fϕ(⋅)f_{\phi}(\cdot) as a U-Net , the most commonly used network architecture for medical image segmentation, with an EfficientNet-b2 backbone . For our proposed method, we trained the segmentation network using an Adam optimizer with an initial learning rate of 3×10−43\times 10^{-4} with learning rate decay. We evaluated our method at the 2k-th epoch where the learning rate decays to zero.

IV-C Quantitative and qualitative results

We compared our method with the empirical risk minimization (ERM) baseline and several recent single-source domain generalization methods. Among them Cutout enforces the model to be robust to corruptions by deliberately removing patches from training images. RSC defines features that lead to the largest gradients as non-robust features and removes them in training. MixStyle synthesize novel domains by mixing instance-level feature statistics . AdvBias is designed for medical images. It augments images using adversarial perturbations. Closely related to our work is RandConv , which employs a random linear intensity transformation model to synthesize novel domains.

Table II summarizes performances on three cross-domain segmentation scenarios, where a network is trained on the source domain and evaluated on the target domain(s). The proposed approach consistently outperforms peer methods. In particular, the performance gains of our method compared with the closely-related RandConv suggest that our approach simulates domain shifts in a more effective manner, leading to stronger robustness upon unseen domains. We also provide the upper bounds: i.e. training and testing in the target domain, in Table II. Qualitative examples are shown in Fig. 5.

To visualize the feature spaces, in Fig. 6 we show t-SNE of the target domain features collected at the last hidden layer of abdominal segmentation networks. As can be seen, for our proposed method, for the same class, target domain features stay close to those of the source domain; while features of different classes are separated.

IV-D Ablation studies

To investigate the effect of configurations of GIN on generalization performance, we conducted ablation studies on two key design-of-choices: the number of convolutional layers and the number of channels in hidden layers. Intuitively, only one or two layers may be insufficient for simulating non-linear transformations across domains in real world, while a too-large number of layers may lead to unrealistically aggressive augmentations that deviate from reality. The effect of numbers of channels in hidden layers is difficult to conjecture, due to the non-linearity of GIN.

We show quantitative results in Table III by varying the number of layers, and the number of channels in hidden layers from the default setting (4 layers and 2 channels). These experiments were conducted under the abdominal segmentation scenario, with IPA turned off for the ease of analysis. The left column of Table III agrees with our intuition regarding the number of convolutional layers.

IV-D2 Configurations of IPA

To examine the effect of interventional pseudo-correlation augmentation, we performed ablation study by removing IPA from the proposed approach. The results in the first two rows of Table IV validate the benefit of mitigating the shifted-correlation effect using IPA, especially for the cardiac and the prostate settings, where pp-values <0.001<0.001 under Wilcoxon signed-rank tests.

As a further exploration, we also examined an alternative design of pseudo-correlation maps: superpixels that are randomly sampled at each iteration , as depicted in Fig. 7. We present its results in the last row of Table IV. We conjecture the sub-optimal performance of the superpixel-based maps is due to the fact that superpixels often coincides with real ground-truth masks y\mathbf{y}’s, making xf\mathbf{x}_{f}’s unnaturally stand out and thus become shortcuts for the network.

V Discussion and conclusion

Domain robustness has been a challenge for deep learning based medical image computing for a long time. In this work, we propose a causality-inspired data augmentation approach for single-source domain generalization. From a methodological perspective, while previous multi-source domain generalization (MDG) and unsupervised domain adapation (UDA) methods are top-down solutions that learn a priori knowledge of out-of-domain data (assumed to be available), our data augmentation is a bottom-up approach based on the causal mechanism of acquisition shifts. Although challenging, pursing causalities and designing bottom-up methods encourage further theoretical investigations on domain shift, which in turn facilitates more principled techniques for robust learning. From a practical perspective, unlike UDA or MDG, our method does not require target-domain data or multi-source data to be available during training. Also, compared with UDA, our method is easier to deploy in real world: it does not require fine-tuning on the target domain (which more or less relies on expertise). Compared to peer single-source generalization techniques, our approach demonstrates consistently superior performances in our experiments.

In the current approach several limitations remain: First, although consistent performance gains are shown in all three scenarios, some crucial hyper-parameters like the number of layers in GIN and the configurations of IPA still require empirical choices. A more elegant augmentation technique that requires less empirical choices is desirable. In addition, as domain-specific information is suppressed during training, our method experiences slight performance downgrades in source-domain testing sets on abdominal CT (∼\sim -2.0/100) and prostate MRI’s (∼\sim -1.0/100), compared to the ERM counterparts. Potential solutions might reside in network architecture side, for example, a dynamic architecture which adaptively balances domain-specific and domain-invariant features.

Starting from the current methodology, several potential extensions arise: For example, as our method can efficiently produce huge amount of domain-shifted images, it is natural to combine the proposed method with multi-source domain generalization techniques . In addition, as our work focuses on image appearance, it is interesting to design methods targeting at domain shifts in terms of anatomical shapes.

References