Data augmentation for deep learning based accelerated MRI reconstruction with limited data
Zalan Fabian, Reinhard Heckel, Mahdi Soltanolkotabi
Introduction
In magnetic resonance imaging (MRI), an extremely popular medical imaging technique, it is common to reduce the acquisition time by subsampling the measurements, because this reduces cost and increases accessibility of MRI to patients. Due to the subsampling, there are fewer equations than unknowns, and therefore the signal is not uniquely identifiable from the measurements. To overcome this challenge there has been a flurry of activity over the last decade aimed at utilizing prior knowledge about the signal, in a research area referred to as compressed sensing .
Compressed sensing methods reduce the required number of measurements by utilizing prior knowledge about the images during the reconstruction process, traditionally via a convex regularization that enforces sparsity in an appropriate transformation of the image. More recently, deep learning techniques have been used to enforce much more nuanced forms of prior knowledge (see Ongie et al. and references therein for an overview). The most successful of these approaches aim to directly learn the inverse mapping from the measurements to the image by training on a large set of training data consisting of signal/measurement pairs. This approach often enables faster reconstruction of images, but more importantly, deep learning techniques yield significantly higher quality reconstructions. Thus, deep learning techniques enable reconstructing a high-quality image from fewer measurements which further reduces image acquisition times. For instance, in an accelerated MRI competition known as fastMRI Challenge , all the top contenders used deep learning reconstruction techniques.
Contrary to classical compressive sensing approaches, however, deep learning techniques typically rely on large sets of training data consisting of images along with the corresponding measurement. This is also true about the use of deep learning techniques in other areas such as computer vision and Natural Language Processing (NLP) where superb empirical success has been observed. While large datasets have been harvested and carefully curated in areas such as vision and NLP, this is not feasible in many scientific applications including MRI. It is difficult and expensive to collect the necessary datasets for a variety of reasons, including patient confidentiality requirements, cost and time of data acquisition, lack of medical data compatibility standards, and the rarity of certain diseases.
A common strategy to reduce reliance on training data in classification tasks is data augmentation. Data augmentation techniques are used in classification tasks to significantly increase the performance on standard benchmarks such as ImageNet and CIFAR-10. For a comprehensive survey of image data augmentation in deep learning see . More specific to medical imaging, data augmentation techniques have been successfully applied to registration, classification and segmentation of medical images. Recently, several studies have demonstrated that data augmentation can significantly reduce the data needed for GAN training for high quality image generation. In a classification setting, data augmentation consists of adding additional synthetic data obtained by performing invariant alterations to the data (e.g. flips, translations, or rotations) which do not affect the response (i.e., the label).
In image reconstruction tasks, however, data augmentation techniques are less common and much more difficult to design because the response (the measurement) is affected by the data augmentation. For example, measurements of a rotated image are not the same as measurements from the original image. In the context of accelerated MRI reconstruction, augmentation techniques such as randomly generated undersampling masks and simple random flipping have been applied, and authors in Schlemper et al. note the importance of rigid transforms in avoiding overfitting. However, an effective pipeline of augmentations for training data reduction and thorough experimental studies thereof are lacking.
The goal of this paper is to explore the benefits of data augmentation techniques for accelerated MRI with limited training data. By carefully taking into account the physics behind the MRI acquisition process we design a data augmentation pipeline, which we call MRAugment Code: https://github.com/MathFLDS/MRAugment, that can successfully reduce the amount of training data required. Our contributions are as follows:
We perform an extensive empirical study of data augmentation in accelerated MRI reconstruction. To the best of our knowledge, this work is the first in-depth experimental investigation focusing on the benefits of data augmentation in the context of training data reduction for accelerated MRI.
We propose a data augmentation technique tailored to the physics of the MR reconstruction problem. It is not obvious how to perform data augmentation in the context of accelerated MRI or in inverse problems in general, because by changing an image to enlarge the training set, we do not automatically get a corresponding measurement, contrary to classification problems, where the label is retained.
We demonstrate the effectiveness of MRAugment on a variety of datasets. On small datasets we report significant improvements in reconstruction performance on the full dataset when MRAugment is applied. Moreover, on small datasets we are able to surpass full dataset baselines by using only a small fraction of the available training data by leveraging our proposed data augmentation technique.
We perform an extensive study of MRAugment on a large benchmark accelerated MRI data set, specifically on the fastMRI dataset. For 8-fold acceleration and multi-coil measurements (multi-coil measurements are the standard acquisition mode for clinical practice) we achieve performance on par with the state of the art with only of the training data. Similarly, again for 8-fold acceleration and single-coil experiments (an acquisition mode popular for experimentation) MRAugment can achieve the performance of reconstruction methods trained on the entire dataset while using only 33% of training data.
We reveal additional benefits of data augmentation on model robustness in a variety of settings. We observe that MRAugment has the potential to improve generalization to unseen MRI scanners, field strengths and anatomies. Furthermore, due to the regularizing effect of data augmentation, MRAugment prevents overfitting to training data and therefore may help eliminate hallucinated features on reconstructions, an unwanted side-effect of deep learning based reconstruction.
Background and Problem Formulation
MRI is a medical imaging technique that exploits strong magnetic fields to form images of the anatomy. MRI is a prominent imaging modality in diagnostic medicine and biomedical research because it does not expose patients to ionizing radiation, contrary to competing technologies such as computed and positron emission tomography.
However, performing an MR scan is time intensive, which is problematic for the following reasons. First, patients are exposed to long acquisition times in a confined space with high noise levels. Second, long acquisition times induce reconstruction artifacts caused by patient movement, which sometimes requires patient sedation in particular in pediatric MRI . Reducing the acquisition time can therefore increase both the accuracy of diagnosis and patient comfort. Furthermore, decreasing the acquisition time needed allows more patients to receive a scan using the same machine. This can significantly reduce patient cost, since each MRI machine comes with a high cost to maintain and operate.
Since the invention of MR in the s there has been tremendous research focusing on reducing their acquisition time. The two main ideas are to i) perform multiple acquisitions simultaneously and to ii) subsample the measurements, known as accelerated acquisition or compressed sensing . Most modern scanners combine both techniques, and therefore we consider such a setup.
where is the number of coils. Obtaining fully-sampled k-space data is time-consuming, and therefore in accelerated MRI we decrease the number of measurements by undersampling in the Fourier-domain. This undersampling can be represented by a binary mask that sets all frequency components not sampled to zero:
We can write the overall forward map concisely as
2 Traditional accelerated MRI reconstruction methods
Traditional compressed sensing recovery methods for accelerated MRI are based on assuming that the image is sparse in some dictionary, for example the wavelet transform. Recovery is then posed typically as a convex optimization problem:
3 Deep learning based MRI reconstruction methods
In recent years, several deep learning algorithms have been proposed and convolutional neural networks established new state of the art in MRI reconstruction significantly surpassing the classical baselines. Encoder-decoder networks such as the U-Net and its variants were successfully used in various medical image reconstruction and segmentation problems . These models consist of two sub-networks: the encoder repeatedly filters and downsamples the input image with learned convolutional filters resulting in a concise feature vector. This low-dimensional representation is then fed to the decoder consisting of subsequent upsampling and learned filtering operations. Another approach that can be considered a generalization of iterative compressed sensing reconstructions consists of unrolling the data flow graph of popular algorithms such as ADMM or gradient descent iterations and mapping them to a cascade of sub-networks. Several variations of this unrolled method have been proposed recently for MR reconstruction, such as i-RIM , Adaptive-CS-Net , Pyramid Convolutional RNN and E2E VarNet .
Another line of work, inspired by the deep image prior focuses on using the inductive bias of convolutional networks to perform reconstruction without any training data . Those methods do perform significantly better than classical un-trained networks, but do not perform as well as neural networks trained on large sets of training data.
MRAugment: a data augmentation pipeline for MRI
In this section we propose our data augmentation technique, MRAugment, for MRI reconstruction. We emphasize that data augmentation in this setup and for inverse problems in general is substantially different from DA for classification problems. For classification tasks, the label of the augmented image is trivially the same as that of the original image, whereas for inverse problems we have to generate both an augmented target image and the corresponding measurements. This is non-trivial as it is critical to match the noise statistics of the augmented measurements with those in the dataset.
We are given training data in the form of fully-sampled MRI measurements in the Fourier domain, and our goal is to generate new training examples consisting of a subsampled k-space measurement along with a target image. MRAugment is model-agnostic in that the generated augmented training example can be used with any machine learning model and therefore can be seamlessly integrated with existing reconstruction algorithms for accelerated MRI, and potentially beyond MRI.
In the following subsections we first argue why we generate individual coil images with the augmentation function, then discuss the design of the augmentation function itself.
As mentioned before, we are augmenting complex-valued, noisy images. This noise enters in the measurement process when we obtain the fully-sampled measurement of an image as , and is well approximated by i.i.d complex Gaussian noise, independent in the real and imaginary parts of each pixel . Therefore, we can write where has the same distribution as due to being unitary. Since the noise distribution is characteristic to the instrumental setup (in this case the MR scanner and the acquisition protocol), assuming that the training and test images are produced by the same setup, it is important that the augmentation function preserves the noise distribution of training images as much as possible. Indeed, a large mismatch between training and test noise distribution leads to poor generalization .
Let us demonstrate why it is non-trivial to generate augmented measurements for MRI through a simple example. A natural but perhaps naive approach for data augmentation is to augment the real-valued target image instead of the complex valued . This would allow us to directly obtain real augmented images from a real target image just as in typical data augmentation. However, this approach leads to different noise distribution in the measurements compared to the test data due to the non-linear mapping from individual coil images to the real-valued target and works poorly. Experiments demonstrating the weakness of this naive approach of data augmentation can be found in Section 4.5.
In contrast, if we augment the individual coil images directly with a linear function , which is our main focus here, we obtain the augmented k-space data
where represents the augmented signal and the noise is still additive complex Gaussian. A key observation is that in case of transformations such as translation, horizontal and vertical flipping and rotation the noise distribution is exactly preserved. Moreover, for general linear transformations the noise is still Gaussian in the real and imaginary parts of each pixel.
To elaborate further, in the multi-coil case our augmentation pipeline applies transformations to the underlying object modulated by the different coil sensitivity maps. In particular, the fully sampled measurement of the th coil in the image domain takes the form
where is i.i.d Gaussian noise obtained via a unitary transform of the original measurement noise. Assuming linear augmentations, the augmented coil image from MRAugment can be written as
where the additive noise is still Gaussian. As seen in (3.2), MRAugment transforms images modulated by the coil sensitivities, therefore the sensitivitiy maps are also indirectly augmented. However, the models we experimented with had no issues learning the proper mapping from augmented measurements with transformed sensitivity maps as our experimental results show.
It is natural to ask if data augmentation would be possible by directly augmenting the object, before the coil sensitivities are applied. If the sensitivity maps are known or are estimated a priori, one may recover the object from the various coils as
where is the complex conjugate of and due to typical normalization . Then, we can apply the augmentation as
Finally, we obtain the augmented coil images as
Comparing (3.3) with (3.2), one may see that now the augmentation is directly applied to the ground truth signal bypassing the coil sensitivities. However, comparing this result in (3.3) with the original unaugmented coil images in (3.1) reveals that the additive noise in (3.3) has a very different distribution from the original i.i.d Gaussian, even worse noise on different augmented coil images are now correlated. Finally, the sensitivity maps are typically not known and need to be estimated before we can apply this augmentation technique, which can introduce additional inaccuracies in the augmentation pipeline.
This discussion motivates our choice to i) augment complex-valued images directly derived from the original k-space measurements, ii) consider simple transformations which preserve the noise distribution and iii) augment individual coil images as in (3.2). Next we overview the types of augmentations we propose in line with these key observations.
2 Transformations used for data augmentation
We apply the following two types of image transformations in our data augmentation pipeline: Pixel preserving augmentations, that do not require any form of interpolation and simply result in a permutation of pixels over the image. Such transformations are vertical and horizontal flipping, translation by integer number of pixels and rotation by multiples of . As we pointed out in Section 3.1, these transformations do not affect the noise distribution on the measurements and therefore are suitable for problems where training and test data are expected to have similar noise characteristics. General affine augmentations, that can be represented by an affine transformation matrix and in general require resampling the transformed image at the output pixel locations. Augmentations in this group are: translation by arbitrary (not necessarily integer) coordinates, arbitrary rotations, scaling and shearing. Scaling can be applied along any of the two spatial dimensions. We differentiate between isotropic scaling, in which the same scaling factor is applied in both directions (: zoom-in, : zoom-out) and anisotropic scaling in which different scaling factors are applied along different axes.
Figure 2 provides a visual overview of the types of augmentations applied in this paper. Numerous other forms of transformations may be used in this framework such as exposure and contrast adjustment, image filtering (blur, sharpening) or image corruption (cutout, additive noise). However, in addition to the noise considerations mentioned before that have to be taken into account, some of these transformations are difficult to define for complex-valued images and may have subtle effects on image statistics. For instance, brightness adjustment could be applied to the magnitude image, the real part only or both real and imaginary parts, with drastically different effects on the magnitude-phase relationship of the image. That said, we hope to incorporate additional augmentations in our pipeline in the future after a thorough study of how they affect the noise distribution.
3 Scheduling and application of data augmentations
A critical question is how to schedule over training in order to obtain the best model. Intuitively, in initial stages of training no augmentation is needed, since the model can learn from the available original training examples. As training progresses the network learns to fit the original data points and their utility decreases over time. We find schedules starting from and increasing over epochs to work best in practice. The ideal rate of increase depends on both the model size and amount of available training data.
Experiments
In this section we explore the effectiveness of MRAugment in the context of accelerated MRI reconstruction in various regimes of available training data sizes on various datasets. We start with providing a summary of our main findings, followed by a detailed description of the experiments. Additional reconstructions and more experimental details can be found in the supplementary material.
In the low-data regime (up to images), data augmentation very significantly boosts reconstruction performance. The improvement is large both in terms of raw SSIM and visual reconstruction quality. Using MRAugment, fine details are recovered that are completely missing from reconstructions without DA. This suggests that DA improves the value of reconstructions for medical diagnosis, since health experts typically look for small features of the anatomy. This regime is especially important in practice, since large public datasets are extremely rare.
In the moderate-data regime ( images) MRAugment still achieves significant improvement in reconstruction SSIM. We want to emphasize the significance of seemingly small differences in SSIM close to the state of the art and invite the reader to visit the fastMRI Challenge Leaderboard that demonstrates how close the best performing models are.
In the high-data regime (more than images) data augmentation has diminishing returns. It does not notably improve performance of the current state of the art, but it does not degrade performance either. Our experiments in the latter two regimes however strongly suggest that data augmentation combined with much larger models may lead to significant improvement over the state of the art, even in the high-data regime. However, without larger models it is expected that in a regime of abundant data, DA does not improve performance. For the models and problem considered here, this is around images. We hope to investigate the effectiveness of MRAugment combined with such larger models in our future work.
Additional benefits of data augmentation include improved robustness under shifts in test distribution, such as improved generalization to new MRI scanners and field strengths. Furthermore, we observe that data augmentation can help to eliminate hallucinations by preventing overfitting to training data.
We use the state-of-the-art End-to-End VarNet model , which is as of now one of the best performing neural network models for MRI reconstruction. We measure performance in terms of the structural similarity index measure (SSIM), which is a standard evaluation metric for medical image reconstruction. We study the performance of MRAugment as a function of the size of the training set. We construct different subsampled training sets by randomly sampling volumes of the original training dataset and adding all slices of the sampled volumes to the new subsampled dataset. For all experiments, we apply random masks by undersampling whole k-space lines in the phase encoding direction by a factor of and including of lowest frequency adjacent k-space lines in order to be consistent with baselines in . For both the baseline experiments and for MRAugment, we generate a new random mask for each slice on-the-fly while training by uniformly sampling k-space lines, but use the same fixed mask for each slice within the same volume on the validation set (different across volumes). This technique is standard for models trained on the fastMRI dataset and not specific to our data augmentation pipeline. For augmentation probability scheduling we use
where is the current epoch, denotes the total number of epochs, and unless specified otherwise. This schedule works resonably well on datasets of various size that we have studied and has not been fine-tuned to individual experiments. Ablation studies on the effect of the scheduling function is deferred to the supplementary.
2 Low-data regime
For the low-data regime we work with two different datasets, the Stanford 2D FSE dataset and the 3D FSE Knee dataset described below and demonstrate significant gains in reconstruction performance.
Stanford 2D FSE dataset. First, we perform experiments on the Stanford 2D FSE dataset, a public dataset of fully-sampled MRI volumes of various anatomies including lower extremity, pelvis and more. We use training-validation split, randomly sampled by volumes. We generate random splits in order to minimize variations in reconstruction metrics due to validation set selection and report the mean validation SSIM over runs along with the standard errors.
We plot a training curve of validation SSIM with and without data augmentation in Figure 3(a). The regularizing effect of data augmentation prevents overfitting to the training set and improves reconstruction SSIM on the validation dataset even in case of training longer than in the baseline experiments without data augmentation. Figure 3(b) compares mean validation SSIM when the model is trained in different data regimes from to of all training data. MRAugment leads to significant improvement in reconstruction SSIM and this improvement is consistent across different train-val splits and training set sizes. We achieve higher mean SSIM using only of the training data with MRAugment than training on the full dataset without DA. On the full dataset, we improve reconstruction SSIM from to , and MRAugment achieves even larger gains in the lower data regime. Figure 4 provides a visual comparison of a reconstructed slice emphasizing the benefit of data augmentation.
Stanford Fullysampled 3D FSE Knees dataset. The Stanford Fullysampled 3D FSE Knees dataset consists of fully-sampled k-space volumes of knees. We use the same methodology to generate training and validation splits and evaluate results as in case of the Stanford 2D FSE dataset.
This dataset has significantly less variation compared to the Stanford 2D FSE dataset. Consequentially, we observe strong overfitting early in training if no data augmentation is used (Figure 5(a)). However, applying data augmentation successfully prevents overfitting. Furthermore, in accordance with observations on the Stanford 2D FSE dataset, data augmentation significantly boosts reconstruction SSIM across different data regimes (Figure 5(b)).
3 High-data regime
Next, we perform an extensive study on the fastMRI dataset , the largest publicly available fully-sampled MRI dataset with competitive baseline models, that allows us to investigate the utility of MRAugment across a wide range of training data regimes. More specifically, we use the fastMRI knee dataset, for which the original training set consists of approximately MRI slices in volumes and we subsample to , , and of the original size. We measure performance on the original (fixed) validation set separate from the training set.
Single-coil experiments. For single-coil acquisition we are able to exactly match the performance of the model trained on the full dataset using only a third of the training data as depicted on the left in Fig. 7. Moreover, with only of the training data we achieve comparable SSIM to the model trained on the full dataset. The visual difference between reconstructions with and without data augmentation becomes striking in the low-data regime. As seen in the top row of Fig. 6, the model without DA was unable to reconstruct any of the fine details and the results appear blurry with strong artifacts. Applying MRAugment greatly improves reconstruction quality both in a quantitative and qualitative sense, visually approaching that obtained from training on the full dataset but using hundred times less data.
Multi-coil experiments. As depicted on the right in Fig. 7 for multi-coil acquisition we closely match the state of the art while significantly reducing training data. More specifically, we approach the state of the art SSIM within using of the training data and within with of training data. As seen in the bottom row of Fig. 6, when using only of the training data we successfully reconstruct fine details comparable to that obtained from training on the full dataset, while high frequencies are completely lost without DA.
Finally, we perform ablation studies on the fastMRI dataset and demonstrate that both pixel preserving and interpolating transformations individually improve reconstruction SSIM. Furthermore their effect is complementary: the best results are obtained by adding all transformations to the pipeline. Moreover, we investigate the effect of the data augmentation scheduling function and demonstrate that exponential scheduling results in better performance compared to a constant augmentation probability. We also evaluate a range of different augmentation schedules and show that both significantly lower or higher probabilities lead to poorer reconstruction SSIM. All ablation experiments can be found in the supplementary material.
4 Model robustness
In this section we investigate further potential benefits of data augmentation in scenarios where training examples from the target data distribution are not only scarce as studied before, but unavailable. Distribution shifts can have a detrimental effect on a variety of reconstruction methods . Furthermore, we show some initial experimental results how data augmentation may help avoiding hallucinated features appearing on reconstructions due to overfitting.
Unseen MR scanners. First, we explore how data augmentation impacts generalization to new MRI scanner models not available in training time. Different MRI scanners may use different field strenghts for acquisition, and higher field strength typically correlates with higher SNR. Approximately half of the volumes in the fastMRI knee dataset have been acquired by a scanner, whereas the rest by three different scanners. We perform the following experiments:
: We train and validate on volumes acquired using scanners. Volumes in the validation set have been imaged by a scanner not in the training set.
: We train on all volumes acquired by scanners and validate on the scanner.
: We train on all volumes acquired by the scanner and validate on all other scanners.
Table 1 summarizes our results. Data augmentation consistently improves reconstruction SSIM on unseen scanner models. Similarly to our main experiments, the improvement is especially significant in the low-data regime. We observe that DA provides the greatest benefit when training on scanners and testing on models. We hypothesize that data augmentation can hinder the model from overfitting to the higher noise level present on acquisitions during training thus resulting in better generalization on the lower noise volumes.
Unseen anatomies. We demonstrate how data augmentation may help improving generalization on new anatomies not included in the training set, even in the high-data regime. We train a VarNet model on the complete fastMRI knee train dataset using the hyperparameters recommended in Sriram et al. , and evaluated the network on the fastMRI brain validation dataset throughout training. We repeated the experiment with the same hyperparameters, but with MRAugment turned on. The results can be seen in Fig. 9. The regularizing effect of data augmentation impedes the network to overfit to the training dataset, thus the resulting model is more robust to shifts in test distribution. This results in higher reconstruction quality in terms of SSIM on unseen brain data when using MRAugment.
Hallucinations. An unfortunate side-effect of deep learning based reconstructions may be the presence of hallucinated details. This is especially problematic in providing accurate medical diagnosis and lessens the trust of medical practicioners in deep learning. We observe that data augmentation has the potential benefit of increased robustness against hallucinations by preventing overfitting to training data, as seen in Fig. 8.
5 Naive data augmentation
We would like to emphasize the importance of applying DA in a way that takes into account the measurement noise distribution. When applied incorrectly, DA leads to significantly worse performance than not using any augmentation.
We train a model using ’naive’ data augmentation without considering the measurement noise distribution as described in Section 3.1, by augmenting real-valued target images. We use the same exponential augmentation probability scheduling for MRAugment and the naive approach. As Fig. 10(a) demonstrates, reconstruction quality degrades over training using the naive approach. This is due to the fact that as augmentation probability increases, the network sees less and less original, unaugmented images, whereas the poorly augmented images are detrimental for generalization due to the mismatch in train and validation noise distribution. On the other hand, MRAugment clearly helps and validation performance steadily improves over epochs. Fig. 10(b) provides a visual comparison of reconstructions using naive DA and our data augmentation method tailored to the problem. Naive DA reconstruction exhibits high-frequency artifacts and low image quality caused by the mismatch in noise distribution. These drastic differences underline the vital importance of taking a careful, physics-informed approach to data augmentation for MR reconstruction.
Conclusion
In this paper, we develop a physics-based data augmentation pipeline for accelerated MR imaging. We find that MRAugment yields significant gains in the low-data regime which can be beneficial in applications where only little training data is available or where the training data changes quickly. We also demonstrate that models trained with data augmentation are more robust against overfitting and shifts in the test distribution. We believe that this work opens up many interesting venues for further research with respect to data augmentation for inverse problems in the low data regime. First, learning efficient data augmentation from the training data has been investigated in prior literature and would be a natural extension of our method. Second, finding the optimal augmentation strength throughout training is challenging, therefore an adaptive scheme that automatically schedules the augmentation probability would potentially further improve upon our results. Such a method is proposed in Karras et al. for discriminator augmentation in GAN training, where the augmentation strength is adjusted based on an discriminator overfitting heuristic with great success. Finally, combining our technique with a generative framework such as AmbientGAN that generates high quality samples of a target distribution from noisy partial measurements, could be potentially used to synthesize fully-sampled MRI data from few noisy k-space measurements.
Acknowledgments
M. Soltanolkotabi is supported by the Packard Fellowship in Science and Engineering, a Sloan Research Fellowship in Mathematics, an NSF-CAREER under award #1846369, DARPA Learning with Less Labels (LwLL) and FastNICS programs, and NSF-CIF awards #1813877 and #2008443. R. Heckel is supported by the IAS at TUM, the DFG (German Research Foundation), and by the NSF under award IIS-1816986.
References
Appendix outline
The following appendix provides additional experimental details, enlarged images of reconstructed slices and extra discussions not included in the main paper. The organization of the supplementary material is as follows:
FastMRI dataset. Appendix A provides additional details on the fastMRI dataset and the experimental setup used in our experiments. We plot reconstruction metric results in PSNR in addition to the SSIM comparison in the main paper (Fig. 7) and demonstrate gains comparable to that measured in SSIM. Furthermore, we plot randomly picked reconstructions from the validation set in order to provide a comprehensive view of reconstruction quality. Finally, we apply MRAugment with a model different from the one used in our main experiments on the fastMRI dataset to demonstrate the wider applicability of DA for deep learning based MR reconstruction.
Stanford datasets. In Appendices B and C we provide more details on the Stanford datasets and the experimental details. Moreover, further reconstructed slices are depicted complementing the ones in the main paper.
Robustness. In Appendix D we give more details on the robustness experiments from Section 4.4, describing the MR scanner models used in the experiments.
Ablation studies. In Appendix E we perform ablation studies on the fastMRI dataset to investigate the utility of various augmentations and the effect of augmentation scheduling on the final reconstruction.
Finally, our code is published at https://github.com/MathFLDS/MRAugment. We refer to this code for additional detail regarding the implementation. We note that MRAugment pipeline can be seamlessly integrated with any existing MR reconstruction code, and can be applied to the fastMRI code base by only a couple of lines of additional code. We hope that the utility and ease of use of MRAugment will prove useful for a wider range of practitioners.
Appendix A Experiments on the fastMRI dataset
The fastMRI dataset is a large open dataset of knee and brain MRI volumes. The train and validation splits contain fully-sampled k-space volumes and corresponding target reconstructions for both (simulated) single-coil and multi-coil acquisition. The knee MRI dataset we are focusing on in this paper includes train volumes ( slices) and validation volumes ( slices). The target reconstructions are fixed size center cropped images corresponding to the fully-sampled data of varying sizes. The undersampling ratio is either ( acceleration) or ( acceleration). Undersampling is performed along the phase encoding dimension in k-space, that is columns in k-space are sampled. A certain neighborhood of adjacent low-frequency lines are always included in the measurement. The size of this fully-sampled region is of all frequencies in case of acceleration and in case of acceleration.
Dataset sampling. We use the fastMRI single-coil and multi-coil knee dataset for our experiments. For creating the sub-sampled datasets, we uniformly sample volumes from the training set, and add all slices from the sampled volumes. Our validation results are reported on the whole validation dataset. Images in the dataset have varying dimensions. Due to GPU memory considerations we center-cropped the input images to pixels (which covers most of the images). We use random undersampling masks with acceleration and fully-sampled low-frequency band, undersampled in the phase encoding direction by masking whole kspace lines. We generate a new random mask for each slice on-the-fly while training, but use the same fixed mask for each slice within the same volume on the validation set (different across volumes).
Model. We train the default E2E-VarNet network from Sriram et al. with cascades (approx. parameters) for both the single-coil and multi-coil reconstruction problems. For single-coil data we remove the Sensitivity Map Estimation sub-network as sensitivity maps are not relevant in this problem.
Hyperparameters and training. We use an Adam optimizer with learning rate following Sriram et al. . We train the baseline model on the full training dataset for epochs. For the smaller, sub-sampled datasets we train for the same computational cost as the baseline, that is we train for epochs on th of the training data. Without data augmentation, we observe a saturation in validation SSIM during this time. With data augmentation we trained longer as we still observe improvement in validation performance after the standard number of epochs. We report the best SSIM on the validation set throughout training. We train on 4 GPUs for single-coil data and on 8 GPUs for multi-coil data. The batch size matches the number of GPUs used for training, since a GPU can only hold a single datapoint.
Data augmentation parameters. The transformations and their corresponding probability weights and ranges of values are depicted in Table 5. We adjust the weights so that groups of transformations such as rotation (arbitrary, by ), flipping (horizontal or vertical) or scaling (isotropic or anisotropic) have similar probabilities. For both the affine transformations and upsampling we use bicubic interpolation. Due to computational considerations we only use upsampling before transformations for the single-coil experiments.
A.2 Additional experimental results on the fastMRI dataset
Comparison of additional metrics. In order to provide more in-depth comparison for our main experiment, here we provide results on PSNR as an additional image quality metric, extending our results from Figure 7. We observe significant and consistent improvement in PSNR when applying MRAugment (Fig. 11) with similar trends to SSIM: the improvement is the most prominent in the low-data regime, but still significant in the moderate domain.
Additional reconstructions. In order to demonstrate that MRAugment works well across a wide range of MR slices, here we provide additional reconstructions with and without data augmentation. In multi-coil reconstructions the visual differences are more subtle, therefore we magnified regions with fine details for better comparison.
Figures 15 and 16 provide a comprehensive set of reconstructions across all subsampling ratios with and without data augmentation for the single-coil and multi-coil slices additional to the ones presented in Figure 6.
Even though the most visible improvement on reconstructions is observed when training data is especially low ( subsampling), Figure 12 provides a closer look at a slice where significant details are recovered by MRAugment using of training data.
Figures 17 and 18 provide more reconstructed slices randomly sampled from the validation dataset with and without DA.
Other models. Even though we demonstrated our DA pipeline on E2E-VarNet, the potential of our technique is not limited to a specific model. We performed preliminary experiments on i-RIM , another high performing model on single-coil MR reconstruction. We kept the hyperparameters proposed in Putzky and Welling for the single-coil problem with modifications as follows. Due to computational considerations, we decreased the number of invertible layers to with kernel strides in each layer, resulting in a model with parameters. In order to further reduce training time, we trained on volumes without fat suppression that take up of the full fastMRI knee dataset. We refer to this new reduced dataset as ’full’ in this section. Finally, we trained on center crops of input images for each experiment. We used ramp scheduling with augmentation probability linearly increasing from to . The acceleration factor and undersampling mask were the same as for other experiments. As depicted in Fig. 13 , our experiments show that applying data augmentation to only of the training data can match the performance of the model trained on the full dataset.
Appendix B Experiments on the Stanford 2D FSE dataset
The Stanford 2D FSE dataset is a public dataset of fully-sampled MRI volumes of various anatomies including lower extremity, pelvis and cardiac images. All measurements have been acquired by the same MRI scanner using multi-coil acquisition, however volume dimensions and the number of receiver coils vary between volumes. The total number of MRI slices is about of the fastMRI knee training dataset.
Dataset sampling. When random sampling, we randomly select volumes of the original dataset and add all slices of the sampled volumes. For volumes where multiple contrasts are available, we arbitrarily pick the first one and discard the others. We scale all measurements by to approximately match the range of fastMRI measurements. Unlike the fastMRI dataset, Stanford 2D FSE is not separated into training, validation and test sets. Therefore, we use training-validation split in our experiments, where we generate random splits in order to minimize variations in reconstruction metrics due to validation set selection and show the mean of validation SSIMs over all runs. When performing experiments on less training data, we keep of the full dataset as validation set and only subsample the train split. We use no center-cropping on the training images as volume dimensions vary strongly. We undersample the measurements by a factor of and generate masks the same way as in the fastMRI experiments detailed in Section A.
Model. We train the default E2E-VarNet network as used in the multi-coil fastMRI experiments detailed in Section A.
Hyperparameters and training. We use Adam optimizer with a learning rate of as in our other experiments. For the baseline experiments without data augmentation, we train the model for epochs, after which we see no significant improvement in reconstruction SSIM and the model overfits to the training dataset. With data augmentation we train for epochs, as validation SSIM increases well after epochs and we observe no overfitting. In all experiments, we report the mean of best validation SSIMs over independent runs.
Data augmentation parameters. In all data augmentation experiments on the Stanford 2D FSE dataset we use exponential schedulig with . The range of values for the various transformations is almost identical to that in Table 5. For more details we refer the reader to the attached source code.
Appendix C Experiments on the Stanford Fullysampled 3D FSE Knees dataset
The Stanford Fullysampled 3D FSE Knees dataset is a public MRI dataset of fully-sampled k-space volumes of knees, acquired by the same MRI scanner. Each volume consists of slices of images with a multi-coil acquisition using receiver coils. The full dataset consists of slices, or about of the fastMRI knee training dataset.
Experimental setup. We apply the same dataset sampling, metric reporting method, model and hyperparameters (includig data augmentation scheduling) as in the Stanford 2D FSE experiments in Section B. We scale all original measurements by .
Appendix D Robustness experiments
Validation on unseen MRI scanners. We explore how data augmentation impacts generalization to new MRI scanner models not available in training time. Different MRI scanners may use different field strenghts for acquisition, and higher field strength typically correlates with higher SNR. Volumes in the fastMRI knee dataset have been acquired by the following different scanners (followed by field strength): MAGNETOM Aera (1.5T), MAGNETOM Skyra (3T), Biograph mMR (3T) and MAGNETOM Prisma Fit (3T). The number of slices acquired by the different scanners are shown in Table 5. We perform the following experiments:
: We train and validate on volumes acquired by scanners with field strength. The training set only contains scans from MAGNETOM Skyra and MAGNETOM Prisma Fit, and we validate on Biograph mMR scans.
: We train on all volumes acquired by scanners and validate on the scanner (MAGNETOM Aera).
: We train on all volumes acquired by the scanner and validate on all other scanners.
We combine volumes corresponding to the same scanner model from the official train and validation sets to form the validation set, however for training we only use volumes of the given model from the training set. Table 1 summarizes our results.
Appendix E Ablation studies
Transformations. We performed ablation studies on of the fastMRI knee training dataset in order to better understand which augmentations are useful. We use the multi-coil experiment with all augmentations as baseline and tune augmentation probability for other experiments such that the probability that a certain slice is augmented by at least one augmentation is the same across all experiments. We depict results on the validation dataset in Table 5. Both pixel preserving and general (interpolating) affine transformations are useful and can significantly increase reconstruction quality. Furthermore, we observe that their effect is complementary: they are helpful separately, but we achieve peak reconstruction SSIM when all applied together. Finally, the utility of pixel preserving augmentations seems to be lower than that of general affine augmentations, however they come with a negligible additional computational cost.
Augmentation scheduling. Furthermore, we investigate the effect of varying the augmentation probability scheduling function. The results on the validation dataset are depicted in Table 5, where exponential, denotes the exponential scheduling function in
with and constant, means we use a fixed augmentation probability throughout training. We observe that scheduling starting from low augmentation probability and gradually increasing is better than a constant probability, as initially the network does not benefit much from data augmentation as it can still learn from the original samples. Furthermore, too low or too high augmentation probability both degrade performance. If the augmentation probability is too low, the network may overfit to training data as more regularization is needed. On the other hand, too much data augmentation hinders reconstruction performance as the network rarely sees images close to the original training distribution.