SimCVD: Simple Contrastive Voxel-Wise Representation Distillation for Semi-Supervised Medical Image Segmentation
Chenyu You, Yuan Zhou, Ruihan Zhao, Lawrence Staib, James S. Duncan
Introduction
Medical image segmentation is a popular task in both machine learning and medical imaging communities . Compared to traditional segmentation approaches, deep neural network based segmentation methods have achieved much stronger performance in recent years with huge advances in representation learning . However, previous state-of-the-art approaches are mostly trained with a large amount of labeled data, which pose significant practical challenges in many medical segmentation tasks where there is a scarcity of labeled data due to the heavy burden of annotating images.
In recent years, a wide variety of semi-supervised methods have been designed to tackle these issues, which learn from limited labeled data along with a large amount of unlabeled data, achieving significant improvements in accuracy and greatly reducing the labeling cost. The common paradigms include adversarial learning, knowledge distillation, and self-supervised learning. Contrastive learning, a sub-area of self-supervised learning, has recently been noted as a promising direction since it has shown great promise in learning useful representations with limited human supervision . This is often best understood as pulling together semantically similar (positive) samples and pushing apart non-similar (negative) samples in a shared latent space. The representations uncovered by these contrastive objectives are capable of boosting the performance of any vision system especially in scenarios where the amount of annotated data available for the downstream tasks is extremely low, which is well suited for medical image analysis.
Despite advances in semi-supervised learning benchmarks, previous methods still face several major challenges: (1) Suboptimal performance: although prior works have achieved promising segmentation accuracy in the setting of limited annotations, semi-supervised models are usually not robust due to some information loss, compared with fully-supervised counterparts; (2) Geometric information loss: previous segmentation networks are poor at characterizing geometry, i.e., leveraging the intrinsic geometric structure of the images, such as the object boundary. As a consequence, it is often hard to accurately recognize object contours; and (3) Generalization ability: considering the limited amount of training data, training deep models is usually deficient due to over-fitting and co-adapting .
In this work, we address the question: can we advance state-of-the-art voxel-wise representation learning in a more extreme few-annotation-setting for medical image segmentation? To this end, we present SimCVD, a simple contrastive voxel-wise representation distillation framework, which can be utilized to produce superior voxel-wise representations from unlabeled data for improving network performance. Our proposed SimCVD, built upon the mean-teacher framework , can address the above-mentioned challenges as follows. First, SimCVD predicts the output geometric representations with only two different dropout masks (Figure 1). In other words, we pass two views of the geometric representations to the mean-teacher model, obtain two representations as “positive pairs”, by applying two independent dropout masks, and learn effective representations by efficiently associating positives and disassociating negatives in the shared latent space. Though this unsupervised learning strategy is simple, we find this approach is strikingly effective compared to other common data augmentation techniques (e.g., inpainting and local shuffle pixel). More importantly, as we will show, it achieves comparable performance to previous fully-supervised approaches. Through a series of thorough analyses, we find that dropout can be viewed as minimal data augmentation for performance improvement, and it can effectively regularize the training of deep neural networks, avoid representation collapse and enhance model generalization.
Second, we attribute the cause of the geometric information loss to the need for geometric shape constraints. We address this challenge by performing multi-task learning that jointly predicts a segmentation map along with a signed distance map (SDM) . The SDM calculates the signed distance function of the object, i.e., the distance of a voxel from the boundary of the object, with the sign determined by whether the voxel is within the object. Thus, it can be viewed as a global shape constraint on the labeled data. Considering that the SDM can provide a more flexible geometric measure of the object boundary, we move beyond the supervised learning scheme and exploit the regularity in geometric shapes among different object classes through distilling “boundary-aware” knowledge via a contrastive objective among the unlabeled data. This enables the model to learn boundary-aware features more effectively by encouraging the networks to produce segmentation maps with similar distance map distributions on the entire dataset.
Third, it is challenging to train the segmentation model on small training sets since deep neural networks trained on a limited amount of data are prone to over-fitting. To this end, we propose to use knowledge distillation (KD), which has been shown to be effective in segmentation and classification tasks . The key idea of KD is that a teacher model is first trained, and then used to guide the training of the student model for improving generalization ability. In the medical domain, most existing KD methods simply consider the segmentation problem as a pixel/voxel-level classification problem. In contrast, considering that medical image semantic segmentation is a structured prediction problem, we present a novel structured knowledge pair-wise distillation, which further use the structural knowledge from the mean-teacher model, while avoiding co-adapting and over-fitting.
Our contributions are summarized as follows. First, we propose a novel contrastive distillation model termed SimCVD featured by (i) boundary-aware representations that incorporate rich information of the object shape, (ii) a distillation objective which contrasts different distance map distributions jointly in the shared latent space, and (iii) a pair-wise distillation objective to further distill pair-wise structural knowledge. Second, we demonstrate that, in the setting of very limited annotation, simply using dropout can deliver more robust end-to-end segmentation performance compared to heavily relying on a large amount of labeled data. Third, we conduct experiments on two popular benchmark datasets to evaluate SimCVD. The results demonstrate that SimCVD significantly outperforms other state-of-the-art semi-supervised approaches, while achieving competitive performance compared to fully-supervised counterparts.
Related Work
Semi-Supervised Medical Image Segmentation In recent years, substantial efforts have been devoted to incorporating unlabeled data to improve network performance due to limited annotations. Yu et al. investigated an uncertainty map based on the mean-teacher framework to guide the student network to capture better features. Li et al. proposed to use signed distance fields for boundary prediction to improve the performance. Also, Luo et al. proposed a dual-task-consistency (DTC) model for semi-supervised medical image segmentation by jointly predicting the pixel-wise segmentation maps and the global-level level set representations on the unlabeled data. Our method aims at a more practical and challenging scenario: we train our model in a more extreme few-annotation setting that relies only on a small number of annotations, while achieving superior segmentation accuracy.
Contrastive Learning Self-supervised learning (SSL) has provided robust benefits to vision tasks by learning effective visual representations from unlabeled data in an unsupervised setting. It is based on a commonly-held belief that superior performance gains can be achieved through improved representation learning. Recently, contrastive learning, a type of self-supervised learning, has received a lot of interests . The key idea of contrastive learning is to learn powerful representations that optimize similarity constraints to discriminate similar pairs (positive) and dissimilar pairs (negative) within a dataset. The primary stream of subsequent work focuses on the choice of dissimilar pairs, which is critical to the quality of learned representations. The loss function used to quantify the contrast is chosen from several options, such as InfoNCE , Triplet , and so on. Recent studies introduced memory bank or momentum contrast to use more negative samples for contrast computation. In the context of medical imaging, Chaitanya et al. extended a contrastive learning framework to extract global and local cues in a stage-wise way, which requires human intervention and extensive training time. In contrast to Chaitanya et al., our unified work focuses on explicit modeling of the intrinsic geometric structure of the semantic objects in an end-to-end manner, and hence is able to recognize object boundaries more effectively and efficiently.
Knowledge Distillation The idea of knowledge distillation is to minimize the KL-divergence between the output distributions of the teacher model and the student model, and thus avoid over-fitting. KD has been applied to a variety of tasks , including image classification and semantic segmentation . Recent works found that the student model can outperform the teacher model when they share the same network architecture. Zhang et al. proposed to collaboratively train multiple student models with co-distillation, which improves performance of those individual models. At the same time, in the context of medical imaging, among those existing state-of-the-art KD methods, the self-ensemble mean-teacher framework is widely explored for image segmentation. Different from the existing methods that separately exploit class probabilities for each voxel, we consider knowledge distillation as a structured prediction problem by matching the relational similarity among all pairs of voxels from the encoded feature maps of the mean-teacher model. We have found that our approach significantly improves learning better voxel-wise representations.
Method
In this section, we introduce SimCVD, a semi-supervised segmentation network, which is built from scratch by effectively leveraging scarce labeled data and ample unlabeled data for improving end-to-end voxel-wise representation learning (See Figure 1). We first overview our proposed SimCVD and then describe the task formulation of SimCVD. Finally, we detail each component of SimCVD in the following subsections.
We aim to construct an end-to-end voxel-wise contrastive distillation algorithm to learn boundary-aware representations in the setting of extremely few annotations for volumetric medical imaging segmentation. Although the accuracy of supervised models is usually higher than that of semi-supervised models, the former requires much more labeled data than the latter. In many clinical situations, we only have few annotated data but a large amount of unlabeled data. This situation necessitates a semi-supervised segmentation algorithm that can utilize the unlabeled data to improve the segmentation performance.
To this end, we propose a novel contrastive distillation framework to advance state-of-the-art voxel-wise representation learning. In particular, our base multi-task segmentation network tackles two tasks simultaneously: classification and regression. Specifically, the segmentation network takes the input volume batch and jointly predicts the probability maps (classification) and the SDMs of the object (regression). To obtain better representations, we propose to perform structured distillation in the latent feature space, followed by contrasting the boundary-aware features in the prediction space, to learn more effective boundary-aware representations from 3D unlabeled data by regularizing the embedding space and exploring the geometric and spatial context of training voxels. At test time, we remove the mean teacher and two projection heads, and only the student network is deployed for the medical segmentation tasks.
2 Task Formulation
Our proposed SimCVD framework consists of a mean-teacher network, , and a student network, . Inspired by recent work , the optimization of these two networks can be achieved with an exponential moving average (EMA) which uses a weighted combination of the parameters of the student network and the parameters of the teacher network to update the latter. This strategy has been widely shown to improve training stability and the model’s final performance. Motivated by this idea, our training strategy is divided into two steps. At each iteration, we first optimize the student network by stochastic gradient descent. Then we update the teacher weights using an exponential moving average of the student weights .
The inputs to the two networks are perturbed versions of the same image. That is, given a volume input , we first add different perturbations (i.e., affine transformation and random crop) to generate two different images and . We then feed and with these two corresponding augmented images to obtain two confidence score (probability) maps and .
Before we present our proposed SimCVD in detail, we first describe our base architecture below.
3 Base Architecture
4 Boundary-aware Contrastive Distillation
Many prior methods distill knowledge merely in the shared prediction space by delivering the student network that matches the accuracy of the teacher network. However, this strategy is not robust for the following reasons: (1) the learned voxel-wise representations from the mean-teacher model are usually not robust due to the lack of geometric information; (2) the segmentation model still suffers from generalization issues; and (3) the network performance needs to be further improved. Therefore, we propose to perform boundary-aware contrastive distillation to train our model for better segmentation accuracy.
Our method differs from previous state-of-the-art methods in three aspects: (1) SimCVD imposes the global consistency in object boundary contours to capture more effective geometric information; (2) previous methods follow the standard setting in considering the relations of local patches, while SimCVD aims to exploit correlations among all pairs of voxels to improve robustness; and (3) due to the computational cost, SimCVD does not use a large memory bank. SimCVD trains the contrastive objective as an auxiliary loss during the volume batch updates. To specify our voxel-wise contrastive distillation algorithm on unlabeled sets, we define two discrimination terms: boundary-aware contrastive loss and pair-wise distillation loss.
where is a temperature hyperparameter. The indices and in the denominator are randomly sampled from a mini-batch of images such that 2D slices are sampled in total. and denote the 3D image index and slice index, respectively. The ’s in the denominator that are not are called negative samples. Inspired by the recent success , our boundary-aware contrastive loss is defined as:
where denotes a collection of the positive 2D slice pairs. Note that the index in is over all the unlabeled data, hence these data affect the training of and .
where measures the cosine of the angle between two ’s as their similarity. Note again that this loss also involves all the unlabeled data.
Overall Training Objective SimCVD is a general semi-supervised framework for combining contrastive distillation with geometric constraints. In our experiments, we train SimCVD with two objective functions — a supervised objective and an unsupervised objective. For the labeled data, we define the supervised loss in Section 3.3. For the unlabeled data, the unsupervised training objective consists of the boundary-aware contrastive loss, pair-wise distillation loss, and consistency loss in Section 3.4. The overall loss function is:
where are hyperparameters that balance each term.
Experimental Setup
We evaluated our approach on two popular benchmark datasets: the Left Atrium (LA) MR dataset from the Atrial Segmentation Challengehttp://atriaseg2018.cardiacatlas.org/, and the NIH pancreas CT dataset . For the Left Atrium dataset, it comprises 100 3D gadolinium-enhanced MR imaging scans (GE-MRIs) with expert annotations, with an isotropic resolution of . Following the experimental setting in , we use 80 scans for training, and 20 scans for evaluation. We employ the same pre-processing methods by cropping all the scans at the heart region and normalizing the intensities to zero mean and unit variance. All the training sub-volumes are augmented by random cropping to . For the pancreas dataset, it contains 82 contrast-enhanced abdominal CT scans. Following the experimental settings in , we randomly select 62 scans for training, and 20 scans for evaluation. In the pre-processing, we first truncate the intensities of the CT images into the window [, ] HU , and then resample all the data into a fixed isotropic resolution of . Finally, we crop all the scans centering at the pancreas region, and normalize the intensities to zero mean and unit variance. All the training sub-volumes are augmented by random cropping to . In this study, we compare all the methods on LA and the pancreas dataset with respect to 20% labeled ratio. To emphasize the effectiveness of SimCVD, we further validate all the methods with respect to 10% labeled ratio on LA dataset.
2 Implementation Details
In this study, all evaluated methods are implemented in PyTorch, and trained for iterations on an NVIDIA 1080Ti GPU with a batch size of . For data augmentation, we use standard data augmentation techniques (i.e., random rotation, flipping, and cropping). We set the hyper-parameters , , , , as , , , , , respectively. For the projection head, we set in the AlphaDropout layer, and output size for AdaptiveAvgPool2d. We use SGD optimizer with a momentum of and a weight decay of to optimize network parameters. The initial learning rate is set as and divided by every iterations. For EMA updates, we follow the experimental setting in , where the EMA decay rate is set to . We use the time-dependent Gaussian warming-up function to ramp up parameters, where and denote the current and the maximum training step, respectively. For fairness, we do not adopt any post-processing step.
In the testing stage, we adopt four metrics to evaluate the segmentation performance: Dice coefficient (Dice), Jaccard Index (Jaccard), 95% Hausdorff Distance (95HD), and Average Symmetric Surface Distance (ASD). Following , we adopt a sliding window strategy, which uses a stride with for the LA and for the pancreas.
Results
We compare SimCVD with published results from previous state-of-the-art semi-supervised segmentation methods, including V-Net , MT , DAN , CPS , Entropy Mini , UA-MT , ICT , SASSNet , DCT , and Chaitanya et al. on the LA dataset in two labeled ratio settings (i.e., 10% and 20%).
The quantitative results on the LA dataset are shown in Table 1. SimCVD substantially improved the segmentation accuracy in both 10% and 20% labeled cases. The results are visualized in Fig 2. Specifically, in the setting of 20% labeled ratio, our proposed SimCVD raises the previous best average results from 89.94% to 90.85% and from 81.82% to 83.80% in terms of Dice and Jaccard, even achieving comparable performance to the fully supervised baseline. Using the 10% labeled ratio, SimCVD further advances the state-of-the-art results from 87.49% to 89.03% in Dice. The gains in Jaccard, ASD, and 95HD are also substantial, achieving 80.34%, 2.59, and 8.34, respectively. This suggests that: (1) taking voxel samples with a contrastive objective yields better voxel embeddings; (2) incorporating pair-wise spatial labeling consistency can boost the performance by accessing more structural knowledge; and (3) utilizing a geometric constraint (i.e., SDM) is capable of helping identify more accurate boundaries. Leveraging all these aspects, we can observe consistent performance gains.
2 Experiments: Pancreas
To further evaluate the effectiveness of SimCVD, we compare our model on the pancreas CT dataset. Experimental results on the pancreas CT dataset are summarized in Table 2. We observe that our model consistently outperforms all previous methods, achieving up to 6.72% absolute improvements in Dice. As shown in Figs. 2 and 3, our method is capable of predicting high-quality object segmentation, considering the fact that the improvement in such a setting is difficult. This demonstrates: (1) the necessity of comprehensively considering both boundary-aware contrast and pair-wise distillation; and (2) the efficacy of global shape information. Compared to the previous strong models, our approach achieves large improvements on all the datasets, demonstrating its effectiveness.
Ablation Study
In this section, we conduct extensive studies to better understand SimCVD. We justify the inner working of SimCVD from two perspectives: (1) boundary-aware contrastive distillation (Section 6.1), and (2) the projection head (Section 6.2). In these studies, we evaluate our proposed method on the LA dataset with 10% labeled ratio (8 labeled and 72 unlabeled).
Ablation on Model Component In the model formulation, our motivation is to advance state-of-the-art voxel-wise representations by capturing the geometric and semantic information in 3D space. Rather than transferring knowledge across confidence score maps directly, our SimCVD distills “boundary-aware” knowledge from the teacher network. To validate the idea of boundary-aware contrastive distillation, we compare SimCVD to an ablative baseline (i.e., SimCVD w/o SDM). Table 3 (a) compares each component of SimCVD in the 10% labeled setting. First, we observe that removing the SDMs in training hurts the segmentation performance by , , , and absolute differences in terms of Dice, Jaccard, ASD, and 95HD. This confirms our intuition that the learned boundary-aware representations provide a good prior for improving segmentation accuracy. We also find that using adaptive max pooling strategy (i.e., SimCVD w/ adaptive max pooling) largely degrades the segmentation performance. Our segmentation results demonstrate that SimCVD is an effective approach, outperforming the best previous method with , , , and absolute differences in terms of Dice, Jaccard, ASD, and 95HD. We hypothesize that it is because “w/ adaptive max pooling” leads to information loss during training.
2 Analysis on Projection Head
To further understand how different aspects of our projection head contribute to the superior model performance, we conduct extensive experiments and discuss our findings below.
How to Interpret Dropout? Our experimental results have shown that SimCVD is an effective approach. In the following, we aim to answer two questions. First, how can we interpret SimCVD’s dropout training strategy? Can we view dropout as a form of data augmentation? Second, is it capable of exploiting additional informative cues in practice?
First, we examine whether removing dropout during training can achieve comparable performance. Table 4 shows the ablation result of our dropout on LA. As shown in Table 4, we observe that using dropout achieves a much better result on LA dataset. Compared to the setting , we find that “no dropout” () leads to a dramatic performance degradation by , , , absolute differences in terms of Dice, Jaccard, ASD, and 95HD, respectively. While in the case of , it also significantly hurts the network performance. On the other hand, we observe slight improvements on the other settings, compared to “no dropout”, but eventually underperform SimCVD. This clearly demonstrates the superiority of our dropout strategy to learn better representations with respect to different pairs of augmented images. We speculate that adding dropout can be interpreted as a minimal form of data augmentation, in which the positive pair takes two views of the same images, and their representations make a clear difference in dropout masks.
Effect of Augmentation Techniques To further examine our hypothesis, we compare common data augmentation techniques (i.e., local shuffle pixel, non-linear transformation, in-painting, out-painting) in Table 5. As is shown, the quantitative results reveal interesting behavior of different data augmentation: adding more data augmentation does not further contribute to the good model performance. We note that, somewhat surprisingly, it hurts the final prediction performance, and none of them outperforms the basic dropout mask. This suggests that by including these data augmentation techniques, it is possible to introduce additional noise during training, which leads to the representation collapse.
Effect of Pooling Size In Table 3, we demonstrate the network improvements from using adaptive mean pooling instead of adaptive max pooling. We investigate the effects of different pooling sizes in Table 4. Empirically, we observe that using a larger pooling size clearly improves performance consistently. However, we find that the results can not be improved further by increasing the pooling size to . In our implementation, we set the pooling size as .
Conclusion
In this work, we propose SimCVD, a simple contrastive distillation learning framework, which largely advances state-of-the-art voxel-wise representation learning on medical segmentation tasks. Specifically, we present an unsupervised training strategy, which takes two views of an input volume and predicts their signed distance maps of their object boundaries in a contrastive objective, with only two different dropout masks. We further conduct extensive analyses to understand the state-of-the-art performance of our approach, and demonstrate the importance of learning distinct boundary-aware representations and using dropout as the minimal data augmentation technique. We also propose to perform structural distillation by distilling pair-wise similarities, which achieves good performance improvements. Our experimental results show that SimCVD obtained new state-of-the-art results on two benchmarks in an extreme few-annotation setting.
We believe that our unsupervised training framework provides a new perspective on data augmentation along with unlabeled 3D medical data. We also plan to extend our method to solve multi-class medical image segmentation tasks.