Representing Topological Self-Similarity Using Fractal Feature Maps for Accurate Segmentation of Tubular Structures
Jiaxing Huang, Yanfeng Zhou, Yaoru Luo, Guole Liu, Heng Guo, Ge Yang
Introduction
Accurate segmentation of tubular structures is of significant importance across a wide variety of areas. In the area of biological research, for example, the accurate segmentation of tubular structures such as the endoplasmic reticulum (ER, Fig. 1) is critical to the study of related human disease mechanisms . In the area of clinical research, the accurate segmentation of blood vessels (Fig. 1) is essential to the early diagnosis of diseases such as retinopathy and stroke . Similarly, in the area of remote sensing, the accurate extraction of roads from aerial imagery is essential for navigation and route planning . However, accurate segmentation of tubular structures from images remains challenging due to factors such as their complex morphology and geometry, low image signal-to-noise ratio, and poor image contrast.
A wide variety of techniques have been developed for the segmentation of tubular structures. Classical methods rely on manually crafted features such as intensity, texture, and shape. For example, previous studies have utilized deformable shape models to fit tubular structures, leveraging their geometric properties. However, these techniques often cannot handle challenges posed by factors such as poor contrast, high noise and complex background. Deep learning methods have revolutionized image segmentation and achieved substantial improvements in segmentation performance. Recent studies on using deep learning models for segmentation of tubular structures have focused on optimization of loss functions and refinement of model structures . However, these studies primarily aim to achieve high segmentation performance utilizing a limited input of images without providing additional information to their segmentation models. Tubular structures exhibit distinct topology and geometry that are vital for segmentation. In this study, we explore using characteristics of their structure to assist deep learning models in segmentation.
One notable characteristic of tubular structures is their topological self-similarity, namely large and complex tubular structures exhibit similar topological patterns at different scales. For example, Fig. 1 shows that if “one junction with multiple edges” is considered a primary structural component, tubular structure entities such as the ER network and the retinal blood vessel network exhibit similar topology at both global and local scales. To quantify the topological self-similarity of tubular structures, we utilize the fractal theory, which characterizes self-similarity of intricate structures at different scales . Central to this theory is the fractal dimension (FD), a key parameter used previously to describe the textural attributes of images.
In addition to topology and geometry, edges and skeletons are important characteristics in defining tubular structures. In the segmentation of these structures, boundary accuracy and skeletal continuity are crucial for downstream tasks. For example, a commonly encountered problem in the segmentation of interconnected tubular structures (Fig. 1) is the breakages in segmentation results due to factors such as low-contrast or blurring.
To enhance the segmentation quality of tubular structures, we exploit the topological self-similarity as well as edge and skeletal characteristics of tubular structures. The main research contributions of this study are as follows:
1) We have developed a strategy to incorporate fractal features into deep learning networks. Specifically, we extend the fractal dimension from the image-level to the pixel-level and generate the fractal feature map (FFM), which characterizes topological self-similarity and textural complexity of each region within an image or its associated label. The FFM computed from the image is denoted as , while the FFM derived from the label is indicated as . Utilization of as an additional model input and as an additional loss function weight substantially enhances segmentation performance.
2) We develop the multi-decoder network (MD-Net) by extending the U-Net architecture with an edge decoder and a skeleton decoder. These decoders enable the model to simultaneously predict the boundaries and skeletons of image objects, in addition to the primary segmentation masks. By incorporating related constraints within the loss function, our model focuses not only on achieving accurate target segmentation but also allocates increased attention to boundary delineation and skeleton preservation, thereby enhancing the overall prediction quality.
3) We have demonstrated the versatility and robustness of FFMs. Incorporation of into the vanilla U-Net and HR-Net enhances segmentation performance, indicating that FFM can be used as a plug-in for different models.
Related Work
Classical methods have been proposed to improve the performance of segmenting tubular structures by taking into account their geometric characteristics. Firstly, various methods utilize active contours and compute geodesics or minimal distance curves to approximate the contours of tubular structures, thereby effectively delineating their boundaries . Secondly, tree structure-based methods utilize intrinsic shape priors to assist segmentation of different tubular structures. These methods employ a bottom-up method to identify tubular objects and a top-down grouping strategy to recognize tree structures, generating corresponding shape priors . Lastly, various centerline-based methods such as the one developed in use multiscale detection strategies to accurately identify centerlines so that distance transform can be used to provide valuable information for segmentation of tubular structures.
Deep learning methods have also been proposed to integrate topological and geometrical prior knowledge of tubular structures to enhance the performance of their segmentation. The integration is primarily achieved in three ways.
1) Convolutional kernel design. The popular deformable convolution and dilated convolution aim to overcome the limitations of geometric transformations in CNNs and have demonstrated exceptional performance in complex segmentation tasks. Additionally, DSC-Net utilizes dynamic snake convolution to accurately capture the distinct features of tubular structures . By adaptively focusing on slender and winding local structures, DSC-Net achieves improved performance in capturing the intricacy of tubular structures.
2) Model architecture design. Various architecture designs have been proposed to learn the topological and geometrical features of tubular structures. In , a global transformer and dual local attention network are employed to simultaneously capture global and local features to effectively learn the complex geometric properties of tubular structures. Dong et al. propose an enhanced Deformable U-Net that exploits flexible deformable convolutional layers to better generate clear boundaries for 3D cardiac cine MRI. High-frequency components that have strong capabilities to perceive thin structures are fused in to enhance the performance of segmenting thin structures.
3) Loss function design. Various loss functions have been explored for the segmentation of tubular structures. In , a similarity measure centerlineDice (clDice) based on the intersection of segmentation masks and their respective skeletons is introduced. A loss function based on clDice is proposed to enable the networks to generate segmentations with more accurate connectivity information and topology preservation. Araujo et al. design a loss function based on the morphological closing operator that allows models to produce more topologically coherent masks and consistent vascular trees. Wong et al. apply a novel Persistence Diagram Loss that quantifies topological correctness of segmentation over fine-grained structures.
In this study, instead of relying on intricate designs of model architectures or loss functions, we incorporate FFMs as a model input and a loss function weight to enhance the perception of topological self-similarity and textures of images.
2 Fractal Theory and Applications
Despite their topological and geometrical complexity, tubular structures often exhibit topological similarities across different spatial scales. This observation suggests that their complex spatial patterns can be effectively described using simple texture features. Fractal geometry provides a means to describe the irregular or fragmented shapes of natural features and other intricate objects that traditional Euclidean geometry struggles to analyze . Specifically, fractal features offer the capacity to describe and characterize the topological and geometrical complexity as well as textural composition of tubular image objects. Thus, fractal geometry has been applied to image classification and segmentation tasks.
For classification, Roberto et al. and Lin et al. utilize the fractional Brownian motion (FBM) model to extract fractal features of images. These features are then fed into a support vector machine or a convolutional neural network (CNN) classifier to differentiate between different objects. For segmentation, the FBM model is utilized in to extract fractal features from images. Such features are combined with classical methods such as thresholding and region growth techniques for image object segmentation.
Although existing methods have demonstrated good performance by leveraging fractal features, there is still a gap in exploring the integration of fractal features with deep learning for segmentation tasks. The inherent self-similarity observed in tubular structures aligns well with the fractal theory. In this study, we address this gap by incorporating FFMs into the segmentation model, aiming to provide a new and reliable source of information to enhance segmentation performance.
Method
In fractal geometry, the fractal dimension (FD) provides a quantitative measure of an image’s degree of self-similarity and roughness. FD can be estimated via the property of self-similarity . Given an image , it is self-similar if comprises distinct copies of itself scaled down by a factor of . Consequently, for an image, the FD is defined as:
Although the definition of FD based on self-similarity is simple and concise, its direct estimation becomes impractical when dealing with irregular images. To overcome this challenge and estimate the FDs of images, the box-counting method is employed.
Box-counting Method: Consider a grayscale image with dimensions , where denotes the maximum gray level (typically ). We can model the image as a three-dimensional space with indicating the two-dimensional position and a third coordinate denoting the gray value. Then the 3-D space is subdivided into smaller cubic regions, or “boxes”, each with dimensions . Here, is a given scale used be a multiple of the sidelength of a pixel in and can be a multiple of the gray level unit in z-direction. Given and , the value of calculated using the following formula:
Given a grid located at position , suppose that the minimum gray value is contained within the box (), and the maximum gray value within the box (). The minimum number of boxes required to encompass all gray levels within the grid at is computed as:
Considering all grids, the number of boxes that can cover all the patches is expressed as:
where . We can obtain a series of using differing values of . Finally, the fractal dimension can be estimated from the least-squares linear fitting of versus , as illustrated in Equation (1). The flow of box-counting method is summarized in the supplementary material.
Fractal Feature Map: Although fractal dimension and fractal features have been previously applied to classification tasks , the inherent differences between classification and segmentation tasks preclude the direct application of image-level fractal dimension to segmentation endeavors. We extend the calculation of FD from the image-level to the pixel-level to generate the FFM of an image. As depicted in Fig. 2, the process begins with the utilization of a sliding window technique. Within a window, for example, the FD of this region is computed using the box-counting method. Subsequently, the window is shifted along both the horizontal and vertical directions with a step size of 1, resulting in the calculation of the FFM for the entire image. The algorithm for generating the FFM is summarized in Algorithm 1.
The FFMs of the images (denoted as ) are incorporated into the segmentation model as additional input channels (Image, ), enhancing the focus on fractal structure and self-similarity after the normalization process. To mitigate the impact of image noise on the calculation of FD, we adopt the box-counting method proposed in as they substitute the gray value with the standard deviation to make the method more robust.
2 Multi-Decoder Network
In segmenting tubular structures, accurate boundary detection and preservation of global topological connectivity are crucial requirements. To meet these requirements, we propose a new model that we refer to as Multi-Decoder Network (MD-Net).
In addition to predicting the segmentation mask, we introduce an edge decoder and a skeleton decoder to generate the edge and skeleton of the tubular structure, respectively, as depicted in Fig. 3. When (Image, ) are fed into MD-Net, the Encoder employs convolution layers to extract a series of low-level to high-level image features. These features are simultaneously transmitted to the three Decoders by skip connections for prediction. Specifically, in the Encoder, each convolution layer involves the repeated application of two 3x3 convolutions, followed by a rectified linear unit (ReLU), and a 2x2 max pooling operation with a stride of 2 for down-sampling. As for the Decoders in MD-Net, each convolution layer includes an up-sampling of the feature map, followed by a 2x2 convolution that reduces the number of feature channels by half. Subsequently, the feature is concatenated with the corresponding feature copied from the Encoder. This concatenated feature is then processed using two 3x3 convolutions, each followed by a ReLU activation function.
During training, the ground truths for edges and skeletons are derived from the annotated masks by employing the findContours function from the OpenCV library and the skeletonization algorithms in the scikit-image library. We have visualized the boundaries and skeletons in the supplementary material to corroborate the accuracy of the extraction techniques. During the inference stage, the final structural segmentation is the output of the object decoder. Edge and skeleton predictions are only used in model training.
Loss Function: Given an image with pixels, the segmentation ground truth is denoted as , where foreground and background pixels are labeled as 1 and 0, respectively. The prediction of the model is represented as \hat{y}\ \epsilon\, indicating the probabilities of individual pixels being classified as foreground. To compute the loss, we utilize the differentiable Soft IOU Loss , denoted as in Equation (5), where is a smoothing coefficient.
The BCE loss function is applied separately to the edge and skeleton segmentation tasks, denoted as .
For MD-Net, which has three segmentation tasks, the loss function is composed as in Equation (8). The weights and are typically set to 1.0, 0.5 and 0.5, respectively.
3 Fractal Feature Constrained Loss
The value of FD serves as an indicator of the texture complexity present in a given image region. Hence, it is advantageous to incorporate pixel-level FFM as weights of the loss function. This allows for a higher weight to be assigned to regions that exhibit greater complexity, as accurately segmenting such regions poses a greater challenge. Consequently, we calculate the FFM for each label associated with the image, denoted as , and employ as pixel-level weights within the loss function to guide the convergence of the model. Considering the specificity of the edge decoder and skeleton decoder, we have thus restricted the application of exclusively to the object decoder. The Constrained Loss of MD-Net is as follows:
Experiments
We evaluate the FFM and MD-Net using five publicly available tubular datasets including the endoplasmic reticulum (ER) and mitochondrial (MITO) organelle network datasets, the Retinal OCT-Angiography vessel Segmentation (ROSE) and Structured Analysis of the Retina (STARE) retinal vessel datasets, and the Massachusetts Roads (ROAD) remote sensing dataset. By selecting these datasets, our aim is to comprehensively evaluate the performance of our methods across different types of structures and domains, providing a robust assessment of their capabilities. To further examine the influence of FFM, we included the non-tubular dataset NUCLEUS .
All images in ER, MITO, STARE, and NUCLEUS are randomly sampled patches with size of for training and sliding window with size of for validation and testing. A crop size of is used randomly during training and regularly during testing for the ROAD dataset as in . Horizontal and vertical flipping as well as // rotation are used for data augmentation. In addition, data augmentation for ROSE is conducted by rotation of an angle from to as described in . Detailed partitions of the six datasets and experimental setup are shown in the supplementary material.
2 Performance Comparison
We select several segmentation networks for comparison, including the vanilla U-Net , U-Net++ , nnU-Net , and HR-Net . Additionally, we compare our methods with SOTA models for diverse datasets. For the task of retinal vessel segmentation, we choose three-stage , OCTA-Net , GT-DLA , and AF-Net , as these networks are tailored for this task. Furthermore, we compare the performance of MD-Net in the ROAD dataset with DCU-Net , Dconn-Net , and the DSC-Net , a tubular structure segmentation network with specially designed loss . To assess the generalization of FFM, we apply to both U-Net and HR-Net. All models are trained on the same dataset with the same hyperparameter settings.
3 Evaluation Metrics
The models are assessed using three types of metrics: volumetric, topology, and distance.
1. Volumetric scores, including intersection-over-union (IoU), accuracy (ACC), centerlineDice (clDice) , and AUC. These metrics quantify the overlap and agreement between the prediction and ground truth.
2. Topology scores. To evaluate the topological correctness of the segmentations, we calculate the Betti Error for the sum of Betti Numbers and . In the evaluation of ROSE and STARE, Error represents the only.
3. Distance score: Hausdorff Distance (HD) quantifies the accuracy of the boundary delineation and provides information about the spatial closeness of the predicted and ground truth boundaries.
Since the nucleus is oval in shape and not tubular, no Betti Error and clDice are used in segmentation performance evaluation and no skeletal decoder is used in the MD-Net. In this study, the evaluation metrics IoU, ACC, AUC, and clDice, are expressed as percentages (%). The Hausdorff Distance (HD) is measured in pixels (px). All performance metrics are calculated for each image and averaged.
4 Configurations
All models are implemented using PyTorch (version 1.12.0) and executed on 4 NVIDIA 3090 GPU cards. During the training phase, we employed the SGD optimizer with an initial learning rate of 0.05. To dynamically adjust the learning rate as the training progressed, we incorporated warm-up and exponential decay techniques. For all datasets, a fixed batch size of 32 is employed, ensuring consistency in the training process. Additionally, to prevent overfitting, a regularization weight of 0.0005 is applied. The entire training process spanned 50 epochs, providing sufficient iterations for model convergence and learning.
5 Quantitative Evaluation
Based on the results presented in Tab. 1, several conclusions can be drawn.
Performance of MD-Net: Our proposed model, MD-Net, demonstrates superior performance in terms of segmentation accuracy and topological continuity compared to the other models. This superiority is observed across five tubular datasets as well as one non-tubular dataset. We have also performed t-test to check whether improvements of our methods over competing methods in performance are statistically significant. Results are provided in supplementary.
In tubular datasets, MD-Net outperforms other methods in terms of segmentation results. Specifically, in the ER, MITO, ROSE, STARE, and ROAD datasets, MD-Net achieves improvements in IoU of , , , , and , respectively, compared to existing SOTA methods. Furthermore, when considering topological continuity and boundary extraction, MD-Net demonstrates the best performance with the minimum error and Hausdorff Distance. These results highlight the ability of MD-Net, equipped with an edge decoder and a skeleton decoder, to effectively capture edge and skeleton features of thin tubular structures, resulting in more accurate and continuous topology segmentation outcomes.
Even in the non-tubular dataset NUCLEUS, MD-Net achieves the best segmentation results with an IoU of , ACC of , and HD of . This demonstrates the effectiveness of MD-Net and FFM in datasets with simpler structures.
Generalization of FFM: is incorporated into U-Net and HR-Net as an additional input channel without modifying the training parameters. The results in Tab. 1 indicate that the inclusion of leads to improved segmentation performance for both U-Net and HR-Net.
For U-Net, the segmentation performance (IoU) shows improvements of , , , , , and on the ER, MITO, ROSE, STARE, ROAD, and NUCLEUS datasets with the assistance of , respectively. Similarly, for HR-Net, the IoU improves by , , , , , and on the same datasets, respectively. On average, brings an improvement of and in IoU for U-Net and HR-Net, respectively, in the segmentation task of tubular datasets.
In conclusion, FFM demonstrates notable generalization capabilities, leading to enhanced performance across various models. Particularly in datasets with intricate structures, FFM consistently outperforms the original models in segmentation tasks. This improvement can be attributed to the inherent characteristics of FD involved in FFM.
and : Fractal Dimension is a significant metric used to quantify the complexity of image textures. In the context of model training, we leverage the as weights for the loss function. This method enables us to guide the model’s focus towards regions with higher complexity, as indicated by a higher FD. Through extensive experimentation, we have observed that incorporating FFM-based weights into the loss function yields improvements in both the training performance and efficiency of the model, as depicted in Tab. 1. These findings highlight the tangible benefits of including as a weight parameter in optimizing the training process.
6 Qualitative Evaluation
To facilitate a more comprehensive comparison of the experimental outcomes across various models, we employed visualizations to present the segmentation results, as depicted in Fig. 4. The visual analysis provides a clear and intuitive representation of the model’s performance. The observed results validate that the integration of FFM leads to enhanced segmentation outcomes in terms of both edge accuracy and topological continuity. More visualization results can be found in the supplementary material.
Ablation Study
To ascertain that the observed performance improvement of the model is attributed to the FFM and not the increase of input channels, we replace the with the image, thereby transforming the input (image, ) to (image, image). Additionally, we computed the image’s Hurst feature (HF, classical fractal analysis), mean feature (MF, directional analysis), and contrast feature (CF, directional analysis) as introduced by . These features replaced the as input and were trained and evaluated on U-Net and MD-Net. Experimental results demonstrated the advantages of the as shown in Tab. 2. This finding indicates that the inclusion of the FFM plays an effective role in achieving better performance in the segmentation task.
2 Robustness of Fractal Feature Maps
The computation of the FFM is influenced by the selection of window size and step size. We used different window and step sizes to generate different and tested their performance. The results in Tab. 3 demonstrate the robustness of FFM. When used for training the U-Net and MD-Net, the calculated using different window sizes consistently yield results with a variation range of on the ER and STARE datasets. Similarly, applying different steps during the calculation of FFMs does not lead to significant changes in performance. It should be noted that the efficiency of computing can be substantially improved by increasing the step size.
The numerical results in Tab. 2 and Tab. 3 are calculated by volumetric score IoU. Results with more metrics can be found in the supplementary material.
3 Decoders of MD-Net
To evaluate the effectiveness of the edge decoder and the skeleton decoder, a comparative ablation analysis was conducted between MD-Net and U-Net on the ER and STARE datasets. Subsequently, we proceeded to remove either the edge decoder or the skeleton decoder from MD-Net and trained the modified models. The results shown in Tab. 4 confirm that both the edge decoder and the skeleton decoder contribute to improving the segmentation performance. It is noteworthy that the best performance is achieved when all components are utilized simultaneously.
4 Limitations of Fractal Feature Map
The time complexity of FFMs’ computation is associated with the size of image (), window size , and step size . Its time complexity is represented as: . Calculation of FFM indeed incurs computational overhead but is optimized by using multithreading and larger . Currently, the generation of FFMs for a image takes about 150 milliseconds on a Xeon 8336C CPU, given a step size of 1 and a window size of 5. Detailed information on training and inference time can be found in the supplementary material.
Conclusions
In this study, we propose a method that uses fractal feature maps (FFMs) along with a multi-decoder network (MD-Net) for the semantic segmentation of tubular structures. FFMs are used to capture the texture and self-similarity of image regions or label regions at the pixel-level through the fractal dimension. are used as an additional input to enhance the segmentation model’s perception of tubular structures and are used as a weight for the loss function to guide the model training. MD-Net improves the quality of segmentation by simultaneously predicting edges and skeletons through the incorporation of an edge decoder and a skeleton decoder. The proposed method is evaluated on five datasets of tubular structures and one dataset of non-tubular structures. The results demonstrate the superiority of MD-Net with FFM in the segmentation of tubular structures and stable performance in the segmentation of non-tubular structures. Furthermore, the performance improvement observed when incorporating the FFM module into vanilla U-Net and HR-Net underscores the broad applicability and potential of incorporating fractal information into deep learning models.
Acknowledgements
This work was supported in part by the National Natural Science Foundation of China (grants 92354307, 91954201, 31971289, 32101216), the Strategic Priority Research Program of the Chinese Academy of Sciences (grant XDB37040402) and the Fundamental Research Funds for the Central Universities (grant E3E45201X2).