Image Fusion Transformer
Vibashan VS, Jeya Maria Jose Valanarasu, Poojan Oza, Vishal M. Patel
Introduction
Image fusion has proved to be critical in many real world applications, e.g., military , computer vision , remote sensing , and medical imaging . It refers to combining different images of the same scene to integrate complementary information and generate a single fused image. For example, the images captured using a visible sensor are rich in fine details like colour, contrast, and texture. However, visible sensors fail to distinguish between objects and background under poor lighting conditions. The images obtained using thermal sensors capture salient features that distinguish objects from the background during daytime and nighttime. However, thermal sensors lack texture and colour space information about the object. This is because visible sensors work in 300-530 m wavelength while thermal sensors work in 8-14 m wavelength . It would be very useful to have a single image that contains complementary information from visible and thermal sensors. Image fusion specifically tackles this by fusing the complementary attributes from different sources to generate a detailed scene representation .
Traditional methods for image fusion include sparse representation (SR) based methods ; multi-scale transformation based methods ; saliency-based methods and low-rank representation (LRR) based methods . Even though these methods achieve competitive performance, there exists several shortcomings: 1) They have a poor generalization ability as they rely on handcrafted feature extraction, 2) Dictionary learning in SR and LRR is time-consuming , and 3) Different sets of source images require different fusion strategies.
Recent image fusion works have explored CNN-based fusion techniques, , , which outperform traditional ones by overcoming the aforementioned shortcomings. Though existing CNN-based fusion techniques improve generalization ability by learning local features, they fail to extract long-range dependencies in the images. This results in the loss of some essential global context that might be useful for an exemplary fused image. Therefore we argue that integrating local features with long-range dependencies can add global contextual information, which in turn helps to improve fusing performance further. With this motivation, we propose an Image Fusion Transformer (IFT) with a novel Spatio-Transformer (ST) fusion strategy that effectively learns both local features and long-range information at multiple scales to fuse the complementary information from given images (see Fig. 1). The main contributions of this work can be summarized as follows:
We propose a novel fusion method, called Image Fusion Transformer (IFT), that utilizes both local information and models long-range dependencies to overcome the lack of global contextual understanding that exists in recent image fusion works.
The proposed method utilizes a novel Spatio-Transformer (ST) fusion strategy, where a spatial CNN branch and a transformer branch are employed to utilize both local and global features to fuse the given images better.
The proposed method is evaluated on multiple fusion benchmark datasets, where we achieve competitive results compared to the existing fusion methods.
Related works
Traditional image fusion methods employed discrete cosine transform (DCT) , sparse representation (SR), , principal component analysis (PCA) , etc. to extract useful features. However, these feature extraction methods lack generalizability. Moreover, images captured from different sources require different fusion strategies. As a result, traditional fusion strategies are designed in a source-specific manner. Compared to traditional methods, deep learning based methods have shown promising improvement in computer vision tasks such as classification , segmentation and detection . Motivated by this, to overcome these issues with the traditional methods for image fusion, the deep learning-based approaches were explored.
Li proposed a technique where at first, the visible and thermal images are decomposed. Later, they perform fusion using the decomposed images by average feature fusion and deep learning-based feature fusion. Another work proposed by Li uses a fully convolution-based model to fuse the visible and thermal images. Here, features from source images are extracted using a DenseNet encoder and fused using a CNN-based fusion layer. A CNN-based decoder is then used to get the fused image. Li further extended their previous work to an end-to-end fusion strategy for multi-scale deep features minimizing the proposed detail and feature loss. Xu proposed a unified unsupervised end-to-end framework that tackles the fusion problem by integrating it with continual learning. However, all of these methods focus on learning spatial local features between source images and do not consider the long-range dependencies present within the source images. In this work we explore extracting long-range features in addition to the local features to enhance the fusion quality further.
2 Transformers
Transformer model architecture was first proposed by Vaswani and has been proven to be extremely important in Natural Language Processing (NLP) literature over the years. The success of transformer-based models can be attributed to their ability to capture better long-range information compared to recurrent neural networks and CNNs. Motivated by their success, Dosovitskiy proposed a Vision Transformer (ViT) for image classification. This has sparked a significant interest in developing transformer-based methods for vision problems like object detection and segmentation . Hence, in this work, we also exploit a transformer-based architecture to obtain improved image fusion performance by enabling the model to encode long-range dependencies from the images.
Proposed method
The proposed Image Fusion Transformer (IFT) is a fusion network that takes in input source images and generates an enhanced fused image. IFT consists of three parts: encoder network, Spatio-Transformer (ST) fusion network and a nested decoder network as illustrated in Fig. 2. The encoder network consists of four encoder blocks, where each encoder block contains a convolution layer with kernel size followed by ReLU and max-pooling operation. For a given source input, we extract deep features at multiple-scales from each convolution block of the encoder network. These extracted features from both images are then fused at multiple scales using the ST fusion network. The ST fusion network consists of a spatial branch and a transformer branch. The spatial branch consists of conv layers and a bottleneck layer to capture local features. The transformer branch consists of an axial attention-based transformer block to capture long-range dependencies (or global context). Finally, we obtain the fused image by training the nested decoder network with the fused features as an input. The decoder network is based on the RFN-Nest architecture.
2 Self-attention and axial-attention
where , and are query, key and value at any arbitrary location and and are computed as , and , respectively. From Eq. 1, we can infer that self-attention computes long-range affinities throughout the entire feature map unlike CNN. However, this self-attention mechanism is computationally expensive due to its quadratic complexity.
Hence, we employ the axial attention mechanism which is computationally more efficient. Specifically, in axial attention, self-attention is first performed over the feature map height axis and then over the width axis, thus reducing computational complexity. Moreover, Wang , proposed a learnable positional embedding to axial attention query, key and value to make the affinities sensitive to the positional information. These positional embeddings are parameters that are learnt jointly during training. Therefore, for a given input , the self-attention along the height axis can be computed as:
3 Spatio-Transformer (ST) fusion strategy
The proposed ST fusion block consists of two branches: the spatial and transformer branches. In the spatial branch, we use a conv block and a bottleneck layer to capture local features. In the transformer branch, we use axial attention to learn global-contextual features by modelling long-range dependencies through the self-attention mechanism. We add these two features to obtain a fused feature map containing enhanced local and global-context information. Moreover, we applied our ST fusion strategy at multiple scales and then forwarded it to the decoder network to obtain the final fused image. ST fusion block is illustrated in Fig. 3.
4 Loss function
The proposed method is trained to preserve fine structural details and retain the salient foreground and background details. The overall training objective to train IFT, denoted as , can be given as:
where is the structural similarity loss, which is computed as follows
where and are the fused and input source image, respectively. Also, measures structural similarity. If tends to 1, then the fused image retains most of the structural details from the source images and vice versa. The feature similarity loss is calculated as follows
where is the number scales at which deep features are extracted; denote fused image, input source 1 image and input source 2 image, respectively. Also, are trade-off parameters to balance the loss magnitude. is the fused feature map while and correspond to the encoded feature maps of the input source 1 and input source 2 images, respectively. This loss constrains the fused deep features to preserve salient structures, thus enhancing the fused feature space to learn more salient features and preserve fine details. Here, is a hyperparameter.
Experiments and results
Implementation details. For visible and infrared fusion, we train our model on 80000 pairs of visible and infrared images in the KAIST dataset. We test on 21 pairs of visible and infrared images in the TNO Human Factors dataset during testing. Further, we follow RFN-Nest experimental setup by resizing the images to and setting hyperparameters equal to 6, 3, 100, 700. For all experiments, we set the learning rate, epoch, and batch size equal to , 4 and 2, respectively.
For the experiment with MRI and PET images, the network is trained on 9981 cropped patches with image pairs obtained from the Harvard MRI and PET datasets. The trained model is evaluated on 20 pairs of MRI and PET images sampled from the Harvard MRI and PET image fusion dataset. During training, we resize the images to and convert the PET images to IHS scale to fuse the channel with an MRI image. For all experiments, we set the learning rate, epoch, and batch size equal to , 4 and 2, respectively.
Infrared and visible image fusion. From Table 1, we can observe that the proposed method outperforms the existing methods in En and MI metrics. Our method can capture both local and long-range dependencies generating sharper content and preserve most of the visual information compared to other methods. Further, our method produces competitive performance in SCD and MS-SSIM metrics than other methods solely focusing on local image fusion. Qualitative fusion results are illustrated in the top row of Fig. 4. The red box highlights the human and the yellow box highlights the reconstruction of fine features. In the red box, we can observe that capturing long-range dependencies results in assigning the same intensity all over the human for IFT compared to other CNN-based methods. In addition, we can observe in the yellow box that our model can reconstruct fine details as it captures both long-range and local information.
MRI and PET image fusion From Table 2, we can infer that our method outperforms all the existing methods in Entropy and CC metrics by preserving local and long-range information. In SD and MG metrics, it produces competitive performance compared to the existing techniques. Qualitative fusion results are illustrated in the bottom row of Fig. 4 and the red box highlights the intensity variation of PET colors in the fused image. From Fig. 4 we can observe that Structure-aware method lacks color intensity variations in the fused image; whereas, DDcGAN and IFT exhibit better intensity variations and brighter colors. Moreover, IFT color variation is more similar to PET than DDcGAN, thanks to IFT’s ability to encode long-range dependencies. This is inferred from the high performance in Correlation Coefficient (CC) metric from Table 2.
Ablation study. An ablation study is conducted on the ST Fusion branch and the results are reported in Table 3. Spatial-based image fusion performs fusion using only local features, whereas transformer-based image fusion performs fusion operation utilizing long-range dependencies. However, it is crucial to capture both local and long-range features for understanding overall representation, which results in better image fusion. From Table 3, it is evident that our proposed ST fusion network outperforms only spatial or transformer-based image fusion in all metrics by capturing both local and long-range dependencies. Hence, this supports our argument that image fusion is improved by integrating long-range dependencies with local features.
Conclusion
In this work, we proposed an Image Fusion Transformer (IFT) network where we developed a novel Spatio-Transformer (ST) fusion strategy that attends to both local and long-range dependencies. In the ST fusion strategy, a CNN branch and a transformer branch are introduced to fuse local and global features. The proposed method is evaluated on multiple fusion benchmark datasets where we achieve better results compared to the existing fusion methods. Moreover, we perform an ablation study to show the effectiveness of extracting local and long-range information while doing fusion.