Deep Learning and Conditional Random Fields-based Depth Estimation and Topographical Reconstruction from Conventional Endoscopy

Faisal Mahmood, Nicholas J. Durr

I Introduction

COLORECTAL cancer (CRC) is the third most commonly diagnosed cancer in the United States . Colonoscopy screening can significantly reduce colorectal cancer mortality by detecting and removing premalignant lesions. However, this approach has well-known limitations and recent studies have suggested that gastroenterologists can easily miss more than 20% of clinically relevant polyps . Approximately 60% of colorectal cancer cases detected after optical colonoscopy are associated with missed lesions . In addition to the well-characterized problem of missed polyps, non-polypoid lesions are even more difficult to screen for and are increasingly recognized as harboring significant malignant potential .

One of the most effective ways to increase lesion detection rates is by using chromoendoscopy. Chromoendoscopy increases both polypoid and non-polypoid lesion contrast by iteratively spraying and rinsing a topical dye through the colon, effectively encoding surface topography as color contrast . However, chromoendoscopy is not used in routine screening because it doubles procedure time and requires specialized training.

Computational measurement of colon topography has the potential to improve lesion detection rates while meeting practical colonoscopy workflow and clinical constraints. Surface features could be used to amplify lesion contrast, assist in geometric lesion classification , augment conventional images , or improve computer-aided lesion detection algorithms. Current state-of-the-art lesion detection and classification methods rely on color and texture of the lesions for feature extraction . However, surface topography of the colon could prove vital for automatic lesion detection and classification. Lastly, 3D mapping of the colon surface may be useful for advanced colonoscopy quality metrics, such as fractional coverage of the examination .

Despite the recent advances in computer vision (CV) and image processing, colonoscopy remains a particularly challenging environment for depth estimation and 3D reconstruction. Colonoscopes have a monocular camera with close light sources, a wide field-of-view, and both the endoscope and the colon are in frequent motion. The unpredictable movement, limited working area, small endoscope size, non-uniform colon texture patterns and deformable nature of the colon render conventional CV techniques like shape from stereo (SfSt) and shape from texture (SfT) inadequate for robust depth estimation. More advanced approaches attempt to reconstruct colon surfaces with restrictive assumptions , and there is currently no model-based approach that robustly and accurately estimates depth from a colonoscopy video. Photometric stereo endoscopy captures some 3D structure of the mucosa but is inherently qualitative due to the unknown working distances from each object point to the endoscope.

Learning-based methods have been used for monocular depth prediction for conventional computer vision applications , particularly for autonomous navigation . A dictionary learning-based approach has recently been employed for depth estimation for colonoscopy images in . However, the virtual colonoscopy data used for training does not simulate the optical properties of an actual endoscope and there was no validation of the technique presented on real endoscopy images. Promising preliminary results have also been shown for 3D monocular reconstruction for functional endoscopic sinus surgery , but the network used was shallow with a single layer and eight nodes trained from only 36 images. A small number of images were used because it is difficult to get ground truth training data. Moreover, the texture in the data is patient specific and cannot be used to estimate depth from other patients, requiring a new training set to be acquired at the beginning of each new procedure. Recent work on monocular 3D reconstruction for assisted navigation in bronchoscopy uses deep learning for monocular depth estimation . However, they validate their method only on phantom data and require training from patient specific CT data every time depth estimates are required from a new patient.

The concept of combining CNNs with a graphical model for structured learning problems has demonstrated considerable promising results . Pixel-wise labeling, that assigns a continuous or discrete label to each pixel of the image, has generally been tackled with feature engineering. Recently CNNs have been extensively used with great success for a variety of CV tasks . However, such approaches lack spatial consistency such as smooth transitions. Spatial consistency has traditionally been captured by probabilistic graphical models such as CRFs , and can augment CNNs to improve depth estimation accuracy .

I-B Contributions and Significance

Traditional approaches of learning-based depth estimation are trained on data that include patient-specific texture, color, and shape, making them difficult to generalize without acquiring a large amount of ground truth data. Low-level texture details are patient-specific and not diagnostic, such as vascular patterns. High-level texture, on the other hand, contains clinically-relevant features that can be generalized across patients. Although details from texture and color are important, these approaches fail to exploit what may be the strongest cue of depth in the small working distances encountered in endoscopy, the inverse square fall-off in light intensity with propagation distance. Most CT colonoscopy (CTC) or virtual colonoscopy software packages such as Slicer 3D and Viatronix do not utilize an accurate model for the optical properties of an endoscope.

Depth values for a specific object in view of the endoscope are inherently continuous, thus depth estimation from monocular endoscope images can easily be formulated as a Conditional Random Fields (CRF) learning problem. In this work we develop and train a CNN-CRF network on a large dataset of realistic endoscopy data with ground truth depth. Our specific contributions can be summarized as follows:

We developed an accurate optical model of an endoscope that includes the inverse-square law of intensity fall-off, to generate synthetic and virtual images of the colon with ground truth depth.

We use this data to train a network that consists of the unary and pairwise parts of a joint CNN-CRF framework.

We validate our results on test data from the digital synthetic colon, a silicone colon phantom, and endoscopy images collected from a porcine colon registered with CT to give accurate ground truth depth.

To the best of our knowledge, this is the first deep learning network trained on synthetically-generated endoscopy images. In practice, large datasets of labeled or annotated medical images are not generally available due to privacy issues, scarcity of experts available for annotation, underrepresentation of rare conditions which leads to highly correlated features of the normal condition and non standardized datasets. This problem has recently been tackled with transfer learning, which shows promising results on conventional computer vision networks fine tuned for medical images . However, transfer learning can lead to artifacts specifically for regression problems . We show that the significant performance benefits of training with large datasets can be realized by utilizing synthetically-generated medical data with an accurate forward model for the imaging system and an anatomically-realistic model of the organ. We further show that this network accurately estimates depth in real endoscopy images after transfer to a synthetic-like domain.

II Methods

A large dataset of endoscopy images with corresponding ground truth depth maps is required for training a CNN to estimate depth from a monocular scene. This data is challenging to generate because depth sensors are impractical to couple to a small endoscope and must receive regulatory approval to be used in humans. Moreover, the high level texture of the colon is patient-specific and cannot be used to efficiently learn depth. To circumvent these obstacles, we generate several datasets of images from synthetic, phantom, porcine, and human models, which have, increasing realism, decreasing quality of ground truth depth, and decreasing dataset sizes.

Synthetic Colon - Virtual Endoscopy Data.

To train our network, we generated over 100,000 texture free endoscopy images, each with an associated ground truth depth from a digital synthetic colon. This synthetic colon phantom was generated using Blender and the data were recorded using a virtual endoscope with parameters selected to mimic the range found in common colonoscopes. The virtual colon had anatomically realistic diameter, bending angles and polyps (Fig. 1). The rendered images have a resolution of 720×576px720\times 576px and a varying viewing angle between 110−140o110-140^{o}. A Mitchell-Netravalie filter was used to prevent aliasing. Two virtual light sources were placed on either side of the camera on the virtual endoscope and each was configured to provide inverse square fall-off of illumination intensity. The depth of the scene was recorded by calculating the distance from the camera to each point on the synthetic colon being imaged. Fig. 1 shows the synthetic colon and representative endoscopy data with aligned depth maps generated using this procedure. We varied the position of the virtual endoscope to generate a diverse set of endoscopy data and in order to accurately model the effects of endoscope motion and light illuminating similar surfaces from different angles (Fig 2). The virtual endoscope was randomly translated along the horizontal axis within the bounds of the synthetic colon and randomly rotated between 0−18000-180^{0} to generate a diverse set of data. This dataset was used for pre-training the CNN-CRF network.

Colon Phantom CT - Virtual Endoscopy Data.

Although the synthetic colon data may be sufficient for learning the inverse square law, our synthetic colon model lacks high spatial frequency detail. To incorporate sensitivity to these features, we also generated training data from virtual endoscopy images from a CT dataset of a silicone colon phantom (The Chamberlain Group Colonoscopy Trainer #2003). The CT reconstruction was performed using filtered back-projection and was filtered using a Ram-Lak filter with linear interpolation in Slicer 3D . The data was then imaged using the virtual endoscope described previously. These endoscopy images with higher frequency details help the network learn the properties of inverse intensity fall-off on ridges and polyps in the colon. This reconstructed CT data was filtered to reduce the effects of fine texture in the scene which may be specific to the cadaver the phantom was molded from. Over 100,000 images were collected from this setup. Fig. 3 shows a portion of the reconstructed colon phantom, rendered virtual endoscopy images, and corresponding ground truth depth. This dataset was used to fine-tune our CNN-CRF depth estimation network.

To validate the accuracy of our trained model on real tissue, we dissected a porcine colon and fixed it to a half-pipe scaffold with a diameter of 5.1cm5.1cm and a 9090 degree bend to simulate the anatomy of the human transverse and descending colon (Fig. 7). Metallic pins were used as fiducial markers for localization and size estimation. The tissue was then imaged using a benchtop cone beam CT scanner with 720 projections at half-degree increments by rotating the scaffold on a stepper motor stage. The CT projections were reconstructed using filtered back-projection and a Ram-Lak filter . The resulting 3D model was imaged using the virtual endoscope in Blender. The scaffold was then covered with black foil and was imaged with an optical endoscope (Misumi MO-V5006L) with a 104o104^{o} 2.3mm/f102.3mm/f10 wide angle lens with (Misumi L23010IR-M5.5-53). The CT and optical endoscopy results were registered to get ground truth depth maps.

Optimization-based multimodal registration was performed on the virtual endoscopy image collected from a CT reconstructed porcine colon and an optical endoscopy image of the same view. A one-plus-one evolutionary optimizer was used to optimize a Mattes mutual information metric for similarity. The growth factor was set to 1.52, initial radius was set to 0.30, the radius adjustment was set to 0.014 and the maximum number of random spatial samples used to compute the metric was set to 500. The optimization was run for 250 iterations at 3 pyramid levels.

Human Endoscopy Data. Human endoscopy data available from the NIH and other datasets available from the MICCAI endoscopy challenge were used to qualitatively evaluate the performance of our network. Since real endoscopy data can have specular reflections we used graph-based in-painting to partially remove specular reflections.

Table I compares various datasets used in this study and CTC or virtual colonoscopy data. CTC has been used in other depth estimation studies such as , however poor cleaning is a major limiting factor . Moreover, recent studies on nonpolypoid lesion detection using CTC restrict to polyps larger than 66 mm because of the limited practical resolution of in-vivo CTC and the requirement of post-processing to remove artifacts from incomplete preparation . CTC also has a relatively high miss rate for non-polypoid lesions . For these reasons CTC data is not used for training.

II-B Deep Learning with Conditional Random Fields

Inspired by success of using similar models for analyzing conventional images in , we implemented an algorithm to estimate depth using continuous CRF and CNN. A continuous CRF is able to exploit the continuous depth values within specific regions of an endoscopic scene. Moreover, unlike several previous methods, this method does not require assumptions since the log-likelihood optimization problem can be directly solved because the partition function can be analytically calculated. Fig. 4 shows a top level flow diagram of the setup. The following sections describe the unary and pairwise parts and the training in detail. Then we discuss how the network trained on virtual endoscopy data is adapted to real data.

As with general CRFs the conditional probability distribution of the raw data can be defined as:

The energy function, E(\text{\boldmathy},\text{\boldmathx}) can be defined in terms of the unary potentials ψ\psi and pairwise potentials ϕ\phi over nodes N\mathcal{N} and edges S\mathcal{S} of xx,

where, the unary part, ψ\psi, regresses the depth from each superpixel and the pairwise part, ϕ\phi, enforces smoothness between similar neighboring superpixels. γ\gamma and β\beta are the two learning parameters associated with the unary and pairwise terms respectively.

Unary Potential. The unary part is designed to regresses superpixel-wise depth for an input endoscopy image. Similar to the unary potential can be defined as follows,

where hih_{i} is the regressed depth of superpixel ii and γ\gamma represents CNN parameters. The architecture of the training network is described systematically in Fig. 5. The CNN used for the unary part makes use of recent developments in fully convolutional networks (FCNs). Unlike standard CNNs which are composed of convolution followed by fully connected layers and produce non-spatial outputs, FCNs can take images of any size and produce spatial convolutional maps. FCNs have have been extensively used for complicated problems specifically for semantic segmentation . We initialize the first five layers from Alex-Net . Two additional 512512 channel convolutional layers with a filter size of 3×33\times 3 are added to the network (as shown in α\alpha, Fig. 5). α\alpha is capable of taking an input image of any size and giving 512512 channel convolutional maps. A typical problem with all fully convolutional architectures is that the feature maps produced can be significantly smaller than the actual size of the images. We mitigate this problem using a convolution map up-sampling step. For ease of implementation we use nearest neighbor up-sampling. Moreover, we incorporate a super-pixel pooling layer similar to to acquire super-pixel features from convolutional maps.

Learning h(\text{\boldmath\gamma}) and β\beta. The overall energy function defined in Eq. 3 can now be populated with unary and pairwise terms and can be written as,

For simplicity and explicit vector calculations the term Ai,j=∑k=1KβkSi,jkA_{i,j}=\sum_{k=1}^{K}\beta_{k}S_{i,j}^{k} can be defined as the affinity matrix, and Di,i=∑jAi,jD_{i,i}=\sum_{j}A_{i,j} as a n×nn\times n diagonal matrix. \text{\boldmathL}=\text{\boldmathD}-\text{\boldmathA} defines the graph Laplacian for further simplicity we notate \text{\boldmath\xi}=\text{\boldmathI}+\text{\boldmathL} where II is a n×nn\times n identity matrix. Using these notations Eq. 6 can be simplified as,

Assigning \text{\boldmath\zeta}=2\text{\boldmathh}^{\top}-\text{\boldmath\xi}\text{\boldmathy}^{\top}, the probability density function in Eq. 2 can now be simplified to the following form,

Given Pr(\text{\boldmathy}|\text{\boldmathx}) we can now calculate the negative log-likelihood which simplifies to,

The negative log-likelihood of the training data is minimized during the training process and the optimization problem can be represented as the following objective function,

Optimization Solution. The optimization problem is solved by standard stochastic gradient decent-based back-propagation. For the unary part the the partial derivatives of -\log Pr(\text{\boldmathy}|\text{\boldmathx}) are calculated with respect to γ\gamma. In Eq. 9 only the terms with hh represent a term with γ\gamma so all other terms are excluded as a result of the partial derivative,

For the pairwise part the partial derivatives are calculated with respect to β\beta

Depth Estimation. To estimate the depth of a new endoscopic image, the MAP problem in Eq. 1 must be solved. Here we show that a closed form solution of the problem exists based on the definitions presented above.

To solve the above maximization problem, the partial derivative of the maximization term has to be calculated with respect to yy. Thus, all terms without yy can be ignored, simplifying the problem to,

This clearly shows that the problem has a close form solution and can be solved.

II-C Adversarial Training for Domain Adaptation

Since our network was trained on synthetic data, where low-level patient specific texture details are absent, we include a domain adaption step to test the network on real images. For new input images that contain this texture, we developed a network that transforms them to a synthetic like representation. This bridges the gap between real and synthetic domains. We use adversarial training between a discriminator network and a transformer network. This setup is based on recent advances in generative adversarial networks and adversarial training . The transformer network takes batches of synthetic images for unsupervised training and learns to remove patient-specific texture from the input images. The discriminator, which is embedded in the transformer’s loss function, classifies the output as real or synthetic. Once the training reaches Nash equilibrium, the transformer is able to fool the discriminator every time and can perfectly transform a real image to its synthetic counterpart. To prevent the synthetic-like representation of a real image from deviating significantly from the original image we use a self-regularization term to preserve patient independent features. If θt\theta_{t} and θd\theta_{d} represent the learning parameters for the transformer and discriminator respectively, \mathbfitDθd\mathbfit{D}_{\theta_{d}} represents the trained discriminator and \mathcal{T}_{\theta_{t}}(\text{\boldmathx}) represent the output of the transformer then the overall transformer loss term can be defined as,

III Experiments

We implemented the training networks using VLFeat Mat-ConvNet using MATLAB 2017a and CUDA 8.0. The training data was prepared by oversegmenting each virtual endoscopy image into superpixels using SLIC and corresponding ground truth depth was assigned to each superpixel. The data was randomized to prevent the network from learning too many similar features quickly. The network was pre-trained on synthetic colon data and fine tuned on colon phantom data. 55%55\% of the data was used for training and 40%40\% for validation and 5%5\% for testing. Training was done using K80 GPUs on the Maryland Advanced Computing Cluster (MARCC). Momentum was set at 0.9 as suggested in and both weight decay parameters (λ1,λ2)(\lambda_{1},\lambda_{2}) were set to 0.0007. The learning rate was initialized at 0.00001 and decreased by 20% every 20 epoches. These parameters were tuned to achieve best results. A total of 300 epochs were run and the epochs with least log⁡10\log 10 error were selected to avoid the selection of an over-fitted model.

III-B Quantitative Evaluation Metrics

We evaluated the three datasets mentioned in section II based on metrics used by other monocular depth estimation work , mostly for conventional vision. These metrics are:

Relative Error: 1N∑y∣ygt−yest∣ygt\frac{1}{N}\sum_{y}\frac{\lvert y_{gt}-y_{est}\rvert}{y_{gt}}

Root Mean Square Error:1N∑y(ygt−yest)2\sqrt{\frac{1}{N}\sum_{y}(y_{gt}-y_{est})^{2}}

Average log10log10 Error: 1N∑y∣log⁡10ygt−log⁡10yest∣\frac{1}{N}\sum_{y}\lvert\log_{10}y_{gt}-\log_{10}y_{est}\rvert

III-C Comparative Analysis

It is not possible to compare our results directly with existing endoscopy depth estimation work because of the diversity of datasets and evaluation methods used. However, we do make a comparison with an FCN regression model that does not employ CRFs. Simple FCNs have recently been used for a variety of CV tasks including work for endoscopy . This comparison allows us to judge the benefit of using a graphical model. The CRF loss layer in the network is replaced with least squares regression. However, we do not claim this to be a direct comparison with or other works because the data used is drastically different and their setup does not use of super-pixels.

III-D Results with Synthetic Colon and Phantom Virtual Endoscopy Data

The trained network was tested on images from the synthetic colon and silicone colon phantom that were not used for training (Fig. 6). Using 10,000 randomly-selected test images we observe that the accuracy of the network improves by every metric with the CNN-CRF method for images which are very similar to the training data (Table II, III).

III-E Results with Porcine Colon Real Endoscopy Data

As mentioned earlier optical and virtual endoscopy views from a CT reconstruction were registered to get ground truth depth maps for optical endoscopy images (Fig. 7). For a fair comparison, the depth map from endoscopy was filtered through the same pipeline of filters used for the CT reconstruction. Only registered regions of the two depth maps were compared. Representative images from this process are shown in Fig. 8 and a the algorithm performance on 1,4601,460 porcine colon images is summarized in Table IV.

III-F Qualitative Results with Real Human Endoscopy Data

We tested the trained network on real data from colonoscopy images available from the NIH and the MICCAI endoscopy challenge databases . The results in Fig. 9, show that the network can regress coarse depth maps that match intuitive cues from the image. These depth maps were then used to reconstruct the topography of the surface of the colon. Fig. 10 compares our method with 3D monocular reconstruction from the tubular assumption-based approach presented by Hong et al. .

IV Conclusions

This paper presents a novel architecture for monocular endoscopy depth estimation and topographical reconstruction that uses the advantages of a joint CNN and CRF-based framework. Unlike previous, approaches this method does not require geometric assumptions. The network was trained using 200,000 images from synthetically-generated data and CT-reconstructions imaged using a virtual endoscope. To the best of our knowledge, this is the first work which trains a network from a large set of synthetically generated and rendered medical images. This is a particularly relevant approach to 3D endoscopy applications because, despite the clinical need, there are no practical alternatives to acquiring large datasets of real endoscopy images with corresponding ground truth. We validate our network on real colon tissue and endoscopy by generating a test dataset using a porcine colon and mounting it on a scaffold followed by CT and registered optical endoscopy. Our work adds to the active area of 3D endoscopy research and has the potential to improve CAD algorithms for detection, segmentation and classifications of lesions.

The limitations of the current method include artifacts due to specular reflections, cases where inverse of intensity might not be the major cue and instances where the pairwise similarities can give rise to artifacts. Moreover, there were several sources of testing errors beyond the inherent accuracy of the network. The reconstruction, refinement, and filtering of the raw CT data all contribute to inaccuracies in the depth map used for ground truth. The CT data also includes streaking artifacts due to non-uniform x-ray absorption, specifically around the metallic pin fiducial markers. More error sources include inconsistency of the stepper motor which rotated the scaffold and errors related to registering the virtual endoscopy CT view with the optical endoscopy image.

Our future work will focus on generalizing the concept of synthetic data generation for medical images and utilizing depth estimation as an additional cue for other endoscopy applications. The proposed network can also be used as an initialization for future deep learning-oriented endoscopy applications.

V Acknowledgments

The authors thank Dr. J. Webster Stayman and Mr. Steven Tilley II (Department of Biomedical Engineering, Johns Hopkins University, Baltimore, MD) for collecting cone-beam CT data on the porcine colon, and Dr. Jeffrey Siewerdsen for his insightful feedback on the manuscript. The authors also thank the staff at Maryland Advanced Computing Cluster (MARCC) for their efficient technical support and training.

References