Multi-Scale Coarse-to-Fine Segmentation for Screening Pancreatic Ductal Adenocarcinoma

Zhuotun Zhu, Yingda Xia, Lingxi Xie, Elliot K. Fishman, Alan L. Yuille

Introduction

Pancreatic cancer is one of the most dangerous killers to human lives, causing more than 330,000330\rm{,}000 deaths globally in 2014 . Pancreatic ductal adenocarcinoma (PDAC) is the most common type of pancreatic cancer, accounting for about 85%85\% of cancer cases. In early stages, this disease often has few symptoms and is very difficult to discover. By the time of diagnosis, the cancer has often spread to other parts of the body, leading to a very poor prognosis (e.g., a five-year survival rate of 5%5\% ). But, for cases diagnosed early, the survival rate rises to about 20%20\% . Hence, it is very important to study the possibility of detecting PDAC in common examinations, e.g., the abdominal CT scan.

The early diagnosis of pancreatic cancer requires much expertise in reading the scanned images and making decisions, but the increasing number of cases makes it impossible for a limited number of experienced radiologists to check all CT scans manually. Therefore, an artificial intelligence system for this purpose is in need. In particular, the radiologists in our team are interested in a system working on abdominal CT scans, which filters out a large fraction of normal cases, but preserves almost all abnormal cases for further investigation. To the best of our knowledge, there is no existing work on this task.

With the development of deep learning , it is possible to construct a system which learns from professional knowledge in data annotation, and apply it to helping doctors in various clinical purposes. The pancreas is one of the most challenging organs in CT segmentation . The difficulty mainly lies in its irregular shape and low contrast around the boundary. Powered by the recent progress in deep learning for 2D and 3D image segmentation, researchers designed various approaches towards accurate pancreas segmentation. In the pathological cases, the morphology of the pancreas can be largely impacted by the difference in the pancreatic cancer stage .

Our work is aimed to detect PDAC from a mixture of normal and abnormal CT scans. This is not a simple classification since radiologists also want to know the location of PDAC, we suggest a solution named segmentation-for-classification, which trains segmentation models and uses their outputs for classification. To deal with tumors of various sizes (Fig. 1), we adopt a segmentation network with multiple input scales, i.e., 64364^{3}, 32332^{3} and 16316^{3} volumes. But, voting that small input regions lead to a high false alarm rate, we adopt a coarse-to-fine testing strategy, which uses the 64364^{3} network for a coarse scan, and then the 323&16332^{3}\&16^{3} networks inside the bounding box to detect small tumors that are possibly ignored in the previous stage. A non-parameterized post-processing algorithm is designed to remove outliers.

Our contributions are three folds: 11) we voxelwisely annotate an abdominal CT dataset with 439439 cases in total, in which 136136 cases are diagnosed with PDAC while the remaining 303303 cases are normal, which is currently the largest PDAC dataset to the best of our knowledge; 22) we adopt a multi-scale segmentation-for-classification framework to conduct an interpretable abnormality detection, which provides radiologists with suspicious regions for further diagnosis; 33) our framework achieves a sensitivity of 94.1%94.1\% at a specificity of 98.5%98.5\%, which shows a promising direction to make a potential significant clinical impact.

The Segmentation-for-Classification Approach

Although some previous work suggested to classify CT or MRI volumes directly using 3D networks , we argue that a better solution is to perform tumor segmentation at the same time of classification. This makes the classification results interpretable by segmentation cues, by which radiologists can take a further investigation of the suspicious abnormal regions. In addition, this integrates voxel-wise annotations into the classification model as deep supervision, so that the entire network is better trained . Therefore, we propose a two-stage framework named segmentation-for-classification, in which a segmentation stage first extracts voxel-wise cues from the input CT scan, and a classification stage follows to summarize this information into the final prediction. Our multi-scale segmentation strategy is different from , which applied another network of the same scale in the fine stage. Tumor detection requires multiple scales.

Mathematically, let each training data be augmented by a segmentation mask Mn⋆\mathbf{M}_{n}^{\star} of the same dimensionality as X\mathbf{X}, so that mn,i⋆∈{0,1,2}{m_{n,i}^{\star}}\in{\left\{0,1,2\right\}} indicates the category of the ii-th voxel, i.e., in the tumor (mn,i=2{m_{n,i}}={2}), outside the tumor but inside the pancreas (mn,i=1{m_{n,i}}={1}), or outside the pancreas (mn,i=0{m_{n,i}}={0}). Note that the tumor voxel set is a subset of the pancreas voxel set. The segmentation module is a high-dimensional function M=s ⁣(X){\mathbf{M}}={\mathbf{s}\!\left(\mathbf{X}\right)}, which is implemented by a deep encoder-decoder network. The classification module is a binary function y=c ⁣(M){y}={c\!\left(\mathbf{M}\right)}. The overall framework is thus written as:

2 Training: Multi-Scale Deeply-Supervised Segmentation

We start with describing the segmentation stage. The tumor region in a pancreas, as shown in Fig. 1, can vary in scale, appearance and geometric properties. In particular, the largest tumor in our dataset occupies over one million voxels, but the smallest one has only thousands. This motivates us to train multi-scale networks to deal with such a large variation in scale.

In practice, we train three networks, taking input volumes of 64364^{3}, 32332^{3} and 16316^{3} voxels, respectively. Each segmentation network follows an encoder-decoder flowchart shown in Fig. 2. It has a series of convolutional layers to learn 3D patterns from training data. Down-sampling and up-sampling are implemented by max pooling and deconvolutional layers, respectively. Following , we introduce deep supervision in the training process, which is implemented by adding several auxiliary losses to intermediate layers, which delivers better performance for the normal and cystic pancreas segmentation in . Deep supervision is considered as a way of incorporating multi-stage visual cues, which constrains intermediate layers and improves the stability of training deep networks. Multi-scale segmentation is complementary to deep supervision, which aims at capturing visual patterns of various scales. As can be seen in experiments, multi-scale segmentation can take advantage of different scales, i.e., a large network produces a high specificity, and a small network gives a high sensitivity.

The training process starts with sampling patches of a specified size. Since the pancreas and the tumor only occupy a small fraction of the entire volume, a random sampling strategy may lead to that only few patches contain pancreas or tumor voxels, and thus the segmentation models are biased towards the background class. To deal with the issue, we sample lots of foreground patches for training the 32332^{3} and 16316^{3} networks. We first compute the region-of-interest (ROI) by padding a 3232-voxel margin around the minimal 3D bounding box covering the entire pancreas. Within it, we categorize the randomly sampled patches into three types (i.e., background, tumor and pancreas) according to the fraction of pancreas and tumor voxels, and make the numbers of training patches of these types o be approximately the same. Data augmentation is performed by randomly flipping patches and rotating by 90∘90^{\circ}, 180∘180^{\circ} and 270∘270^{\circ} over three axes.

We use the same configuration for training these networks. The base learning rate is 0.010.01 and decayed polynomially (the power is 0.90.9) in a total of 80,00080\rm{,}000 iterations (the mini-batch size is 1616, 3232 and 128128 for 64364^{3}, 32332^{3} and 16316^{3}, respectively). The weight decay and momentum are set to be 0.00050.0005 and 0.90.9, respectively.

3 Testing: Coarse-to-Fine Segmentation with Post-Processing

The first goal is to perform the pancreas and tumor segmentation. We first slide a 64364^{3} window in the entire CT volume. The spatial stride is 2020 along three axes, which is chosen to have the average testing time for each case within 1111 minutes on a TITAN Xp GPU. Based on the coarse segmentation, we compute the ROI, i.e., the smallest box covering all pancreas and tumor voxels padded by 3232, and crop the CT image accordingly. Then, we scan the ROI with sliding windows of 32332^{3} and 16316^{3} voxels, and the strides are set to be 1010 and 55, respectively. We do not run the two small networks on the entire volume because it can easily hallucinate tumors in the background regions. In addition, shrinking the scanning region for the 32332^{3} and 16316^{3} networks leads to a significant speedup in the testing process. The predictions of three networks are averaged into final segmentation.

Then, based on the segmentation mask, we classify each volume as normal or abnormal. Advised by the radiologists who desire the classification result to be explainable, we do not formulate the classifier c ⁣(⋅)c\!\left(\cdot\right) as another deep network, but use a simple, non-parametrized approach to filter out the outliers. We construct a graph on all voxels predicted as normal pancreas or tumor. Each voxel is a node, and there exists an edge between the adjacent voxels (each voxel is adjacent to 66 neighbors). We compute all connected component in the graph. A component is preserved if it is larger than 20%20\% of the maximal connected component, otherwise it is removed, i.e., all voxels within this component are predicted as background. To obtain our final goal, a volume is predicted as PDAC if at least KK voxels are predicted as tumor. In practice, we empirically set K=50{K}={50}.

Experiments

We collected a dataset with 303303 normal cases from potential renal donors, as well as 136136 biopsy-proven PDAC cases. Four experts in abdominal anatomy annotated the pancreas and tumor voxels on these data using the Varian Velocity software, and each case was checked by an experienced board-certified Abdominal Radiologist. For a radiologist, an average normal case took 2020 minutes, and an average abnormal case 4040 minutes to segment. Since the abnormal cases are much harder to obtain and annotate than the normal cases, we adopt a 44-fold cross-validation on our 136136 PDAC scans to have testing results on every abnormal case while we use a hard split of training and testing on our 303303 normal cases. All in all, each training set contains 103103 normal and 102102 abnormal cases where the normal-to-abnormal ratio is close to 11, and each testing set contains 3434 abnormal and 200200 normal cases. The average size of CT scans is 512<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>512<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>667512<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>512<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>667.

2 Segmentation Results

We first summarize the segmentation results in Table 1, which makes the normal v.s. abnormal classification to be interpretable by segmentation cues. The 64364^{3} network achieves reasonable pancreas and tumor segmentation accuracies. The segmentation result of normal pancreas is as high as 86.9%86.9\%, which means that the normal pancreases are easier to segment, as there are often unpredicted changes in shape and geometry in the abnormal cases. As a side comment, the lowest DSC of an abnormal pancreas is 38.4%38.4\%, lower than the number (44.0%44.0\%) of a normal pancreas. In tumor segmentation, we observe a lower accuracy and a higher standard deviation (57.3±28.1%57.3\pm 28.1\%). Except for the 1010 missing cases, we find 2020 more cases with a tumor DSC lower than 30%30\%. All these evidences imply the challenging of finding tumors considering their various size, shape and locations. Note that a recent work on the pancreatic cyst segmentation achieves a DSC of 63.4±27.7%63.4\pm{27.7}\% , which is not as hard as the tumor segmentation.

Going to smaller scales, fewer tumors are missed, though segmentation accuracies become lower. This is the tradeoff between sensitivity and specificity: a network with a smaller input region has the ability to detect tiny regions, but without seeing contexts, it can be easily confused by false positives. Thus, combining multi-scale predictions achieves a balance between sensitivity and specificity. Fig. 3 shows two examples that benefit from multi-scale segmentation.

We replace our backbone with 3D UNet and VNet at the 64364^{3} scale setting and report their results in Table 2 for comparison. We can find that the three backbones perform roughly similar in terms of the segmentation results. However, our backbone achieves the best results for the sensitivity and specificity.

3 Classification Results

Finally, we summarize classification results in Table 1, which is the crucial goal of making earlier diagnosis possible for doctors. Radiologists care more about a high sensitivity since they don’t want to miss a patient who has an abnormal pancreas, which inspires us to adopt a multi-scale strategy to improve the sensitivity while keeping a reasonable specificity. The model with multi-scale information achieves the best overall performance, i.e., a sensitivity of 94.1%94.1\% at a specificity of 98.5%98.5\%. These high scores imply that tumor segmentation provide strong cues for PDAC screening. We show all three false alarms in Fig. 4. The radiologists of our team confirmed that 22 out of these 33 false positives have focal fatty infiltration in the pancreas corresponding to the detected “tumors”. Focal fatty infiltration is difficult for radiologists to distinguish from tumor in current clinical practice. In this case, the predicted “false alarm” was not normal in view of our radiologists.

By augmenting our segmentation for classification framework with cues from number of predicted tumor voxels since the more voxels predicted as PDAC the more likely this case is abnormal, we can output a confident score for each case, indicating the possibility that this case suffers PDAC. More specifically, a confidence score is computed by a weighted sum of the volume size and the segmentation probability of predicted tumor voxels. By sorting all testing cases according to their confident scores, we obtain a ROC curve of sensitivity and specificity. From the ROC curve, we can make different emphasis to change the tradeoff between sensitivity and specificity, e.g., we can achieve a sensitivity of 98.5%98.5\% at a specificity of 95.6%95.6\%, or a specificity of 99.5%99.5\% at a sensitivity of 94.1%94.1\%.

Conclusion

In this paper, we study an important and challenging task, i.e., detecting pancreases diagnosed with PDAC in abdominal CT scans. This topic is crucial in saving lives from pancreatic cancer yet few studied before, possibly due to the lack of data. We propose a segmentation-for-classification framework which trains a segmentation network and performs interpretable abnormality classification by simply checking the existence of tumor voxels in each testing volume. There are two key points to improve classification accuracy, known as multi-scale network training and coarse-to-fine testing. To offer a best trade-off between sensitivity and specificity on our own collected dataset containing 303303 normal and 136136 PDAC cases, we achieve a sensitivity of 94.1%94.1\% at a specificity of 98.5%98.5\%. The strong numbers show the promising direction to make a significant impact in clinics for early detection of pancreatic cancer, which would save lives.

Acknowledgements This work was supported by the Lustgarten Foundation for Pancreatic Cancer Research.

References