DeepOrgan: Multi-level Deep Convolutional Networks for Automated Pancreas Segmentation
Holger R. Roth, Le Lu, Amal Farag, Hoo-Chang Shin, Jiamin Liu, Evrim Turkbey, Ronald M. Summers
Introduction
Segmentation of the pancreas can be a prerequisite for computer aided diagnosis (CADx) systems that provide quantitative organ volume analysis, e.g. for diabetic patients. Accurate segmentation could also necessary for computer aided detection (CADe) methods to detect pancreatic cancer. Automatic segmentation of numerous organs in computed tomography (CT) scans is well studied with good performance for organs such as liver, heart or kidneys, where Dice Similarity Coefficients (DSC) of 90% are typically achieved Wang et al. (2014), Chu et al. (2013), Wolz et al. (2013), Ling et al. (2008). However, achieving high accuracies in automatic pancreas segmentation is still a challenging task. The pancreas’ shape, size and location in the abdomen can vary drastically between patients. Visceral fat around the pancreas can cause large variations in contrast along its boundaries in CT (see Fig. 3). Previous methods report only 46.6% to 69.1% DSCs Wang et al. (2014), Chu et al. (2013), Wolz et al. (2013), Farag et al. (2014). Recently, the availability of large annotated datasets and the accessibility of affordable parallel computing resources via GPUs have made it feasible to train deep convolutional networks (ConvNets) for image classification. Great advances in natural image classification have been achieved Krizhevsky et al. (2012). However, deep ConvNets for semantic image segmentation have not been well studied Mostajabi et al. (2014). Studies that applied ConvNets to medical imaging applications also show good promise on detection tasks Cireşan et al. (2013), Roth et al. (2014). In this paper, we extend and exploit ConvNets for a challenging organ segmentation problem.
Methods
2 Convolutional neural network (ConvNet) setup
We use ConvNets with an architecture for binary image classification. Five layers of convolutional filters compute and aggregate image features. Other layers of the ConvNets perform max-pooling operations or consist of fully-connected neural networks. Our ConvNet ends with a final two-way layer with softmax probability for ‘pancreas’ and ‘non-pancreas’ classification (see Fig. 1). The fully-connected layers are constrained using “DropOut” in order to avoid over-fitting by acting as a regularizer in training Srivastava et al. (2014). GPU acceleration allows efficient training (we use cuda-convnet2https://code.google.com/p/cuda-convnet2).
5 Data augmentation
6 Cross-scale and 3D probability aggregation
Results & Discussion
Manual tracings of the pancreas for 82 contrast-enhanced abdominal CT volumes were provided by an experienced radiologist. Our experiments are conducted using 4-fold cross-validation in a random hard-split of 82 patients for training and testing folds with 21, 21, 20, and 20 patients for each testing fold. We report both training and testing segmentation accuracy results. Most previous work Wang et al. (2014), Chu et al. (2013), Wolz et al. (2013) uses leave-one-patient-out cross-validation protocols which are computationally expensive (e.g., hours to process one case using a powerful workstation Wang et al. (2014)) and may not scale up efficiently towards larger patient populations. More patients (i.e. 20) per testing fold make the results more representative for larger population groups.
0.2 Evaluation:
To the best of our knowledge, this work reports the highest average DSC with 71.8% in testing. Note that a direct comparison to previous methods is not possible due to lack of publicly available benchmark datasets. We will share our data and code implementation for future comparisonshttp://www.cc.nih.gov/about/SeniorStaff/ronald_summers.htmlhttp://www.holgerroth.com/. Previous state-of-the-art results are at 68% to 69% Wang et al. (2014), Chu et al. (2013), Wolz et al. (2013), Farag et al. (2014). In particular, DSC drops from 68% (150 patients) to 58% (50 patients) under the leave-one-out protocol Wolz et al. (2013). Our results are based on a 4-fold cross-validation. The performance degrades gracefully from training (83.66.3%) to testing (71.810.7%) which demonstrates the good generality of learned deep ConvNets on unseen data. This difference is expected to diminish with more annotated datasets. Our methods also perform with better stability (i.e., comparing 10.7% versus 18.6% Wang et al. (2014), 15.3% Chu et al. (2013) in the standard deviation of DSCs). Our maximum test performance is 86.9% DSC with 10%, 30%, 50%, 70%, 80%, and 90% of cases being above 81.4%, 77.6%, 74.2%, 69.4%, 65.2% and 58.9%, respectively. Only 2 outlier cases lie below 40% DSC (mainly caused by over-segmentation into other organs). The remaining 80 testing cases are all above 50%. The minimal DSC value of these outliers is 25.0% for . However Wang et al. (2014), Chu et al. (2013), Wolz et al. (2013), Farag et al. (2014) all report gross segmentation failure cases with DSC even below 10%. Lastly, the variation of enforcing within a structured prediction CRF model achieves only 68.2% 4.1%. This is probably due to the already high quality of and in comparison.
We present a bottom-up, coarse-to-fine approach for pancreas segmentation in abdominal CT scans. Multi-level deep ConvNets are employed on both image patches and regions. We achieve the highest reported DSCs of 71.810.7% in testing and 83.66.3% in training, at the computational cost of a few minutes, not hours as in Wang et al. (2014), Chu et al. (2013), Wolz et al. (2013). The proposed approach can be incorporated into multi-organ segmentation frameworks by specifying more tissue types since ConvNet naturally supports multi-class classifications Krizhevsky et al. (2012). Our deep learning based organ segmentation approach could be generalizable to other segmentation problems with large variations and pathologies, e.g. tumors.
This work was supported by the Intramural Research Program of the NIH Clinical Center. The final publication will be available at Springer.