CuNeRF: Cube-Based Neural Radiance Field for Zero-Shot Medical Image Arbitrary-Scale Super Resolution

Zixuan Chen, Jian-Huang Lai, Lingxiao Yang, Xiaohua Xie

Introduction

Medical imaging techniques such as computed tomography (CT) and magnetic resonance imaging (MRI) are critical tools in assisting clinical diagnosis. However, the acquisition of high-quality medical slices is a resource-intensive process, which requires subjects to be exposed to considerable ionizing radiations for a long time, increasing the lifetime risk of cancer . To reduce the burden on subjects, a feasible approach is to reconstruct high-resolution (HR) medical volumes from low-resolution (LR) ones.

To tackle medical image super-resolution (MISR) challenges, early studies employed optimization methods and interpolation methods . Subsequently, a series of methods have adopted convolutional neural networks to learn the LR-HR mappings. Recently, medical image arbitrary-scale super-resolution (MIASSR) methods have received widespread attention in the MISR community, aiming to employ a single model to upsample medical volumes at arbitrary scales. Although these methods achieve acceptable HR results, they still have two major issues: (i) Existing MIASSR methods rely on the supervision from HR volumes, yet high-quality HR volumes are not always available; (ii) These methods may be susceptible to the distribution gap between training and test data, producing non-existent details. These drawbacks limit the application scenarios of existing MIASSR methods.

To address the above-mentioned limitations, we present a zero-shot MIASSR framework – Cube-based NeRF (CuNeRF), which aims to yield arbitrary upsampling images after training on a test LR volume itself (see Figure 1). Specifically, we draw inspiration from the neural radiance field (NeRF) to estimate the continuous volumetric representation from discrete samples (LR volumes) instead of fitting the mapping between LR and HR volumes. Since directly applying NeRF on medical volumes may result in grid-like artifacts (see Figure 2, detailed explanation is provided in Section 4.1), our CuNeRF tackles such aliasing issues via the proposed differentiable modules: cube-based sampling, isotropic volume rendering, and cube-based hierarchical rendering. As shown in Figure LABEL:fig:contribution, CuNeRF can build a continuous mapping between the coordinate and the corresponding intensity value in the training data, which is capable of generating medical slices at arbitrary scales and free viewpoints in a continuous domain. Comprehensive experiments on the MSD Brain Tumour (MRI) and KiTS19 (CT) datasets show that CuNeRF yields impressive performance in 3D and volumetric MISR at various upsampling scales, outperforming state-of-the-art methods.

To the best of our knowledge, CuNeRF is the first zero-shot MIASSR framework that can continuously upsample medical volumes at arbitrary scales.

We address the hole-forming issues via the proposed techniques: cube-based sampling, isotropic volume rendering, and cube-based hierarchical rendering.

Extensive experiments on CT and MRI modalities for 3D MISR and volumetric MISR show CuNeRF favorably surpasses state-of-the-art MIASSR methods.

Related Works

In this section, we first review implicit neural representation and then introduce some impressive progress in medical image super-resolution. Recent surveys provide a comprehensive review of super-resolution methods.

Learning implicit neural representations (INRs) from discrete samples to form a continuous function has been a long-standing research problem in computer vision for numerous tasks. A recent trend in this field is to map discrete representations to coordinate-based continuous neural representations through implicit functions formed by neural networks, such as multi-layer perceptron (MLP). Chen et al. proposed a method to learn the INR of 2D images using the local implicit image function. Subsequent work extended this to apply in the video domain. Currently, most 3D view-synthesis methods are based on the neural radiance fields (NeRF) framework. NeRF can model a volumetric radiance field to render novel views with impressive visual quality using standard volumetric rendering and alpha compositing techniques . However, NeRF has the drawback of requiring massive training views and lengthy optimization iterations to learn the correct 3D geometry. Several follow-up works have attempted to optimize NeRF’s training procedures, such as reducing the required training views , accelerating convergence and rendering speed . Other works aim to adapt NeRF to various domains, such as generative modeling , anti-aliasing , unbounded representation , and RGB-D scene synthesis . Recently, some researchers employ INR-based methods to reconstruct medical images from discrete-sampled data , more details can be seen in the recent survey .

2 Medical Image Super Resolution

Medical image super-resolution (MISR) is an important task in medical image processing, which aims to reconstruct high-resolution (HR) medical slices from corresponding low-resolution (LR) ones. Initially, some conventional methods like and widely-used interpolation methods like bicubic and tricubic interpolations were employed in the early research. Inspired by , recent studies have shifted their focus towards using deep learning-based super-resolution networks in the medical domain. Lim et al. employ deep learning-based super-resolution networks to upsample medical images. Some studies upsample each 2D LR medical slice to acquire the corresponding HR one, such as . On the other hand, Chen et al. and Wang et al. use 3D DenseNet-based networks to generate HR volumetric patches from LR ones. Yu et al. build a transformer-based MISR network to address volumetric MISR challenges. Recent studies have been focusing on medical image arbitrary-scale super-resolution (MIASSR) , which aims to upsample medical slices at arbitrary scales by a single model. Inspired by Meta-SR , Peng et al. deals with volumetric MISR on the zz-axis at integer scales. Wu et al. propose ArSSR, an INR-based method that can upsample MRI volumes at arbitrary scales in a continuous domain. Wang et al. propose a weakly-supervised framework that uses unpaired LR-HR medical volumes for optimization. However, these methods deeply rely on the HR medical volumes, which limits the application scenarios.

Preliminary: NeRF

Neural radiance field (NeRF) aims to build the continuous mapping from (x,d)(\mathbf{x},\mathbf{d}) to (c,σ)(\mathbf{c},\sigma), where x=(x,y,z)\mathbf{x}=(x,y,z) and d=(θ,ϕ)\mathbf{d}=(\theta,\phi) denote spatial location and viewing direction, while c\mathbf{c} and σ\sigma represent the content color and volume density, respectively. NeRF’s techniques can be summarized as follow:

Ray sampling. NeRF first constructs the ray r(t)=o+td\mathbf{r}(t)=\mathbf{o}+t\mathbf{d} that emits from the center of projection o\mathbf{o} and passes through the materials along the viewing direction d\mathbf{d}. Subsequently, NeRF samples NN points along the ray from near plane tnt_{n} to far plane tft_{f} predefined. For each sampling point r(tk)\mathbf{r}(t_{k}), NeRF employs a positional encoding function γ(⋅)\gamma(\cdot) to map the location xk\mathbf{x}_{k} and view direction d\mathbf{d} into higher dimensional space as:

where ρ\rho denotes an arbitrary vector and LL is a hyperparameter set to 1010 as default.

Volume rendering. The pixel color C(r)\mathbf{C}(\mathbf{r}) can be modeled as the integral of the corresponding ray r\mathbf{r} based on Beer-Lambert Laws as:

where c(⋅)\mathbf{c}(\cdot) and σ(⋅)\sigma(\cdot) denote the color and volume density functions. In practice, NeRF employs a multi-layer perceptron (MLP) FΘF_{\Theta} to estimate these two functions. For each sampling point r(tk)\mathbf{r}(t_{k}), MLP FΘF_{\Theta} predicts the corresponding color ck\mathbf{c}_{k} and volume density σk\sigma_{k} by:

Given the estimated results of the NN sampling points from tnt_{n} to tft_{f}, we can approximate the volume rendering integral using numerical quadrature as introduced by :

where C^(r)\hat{\mathbf{C}}(\mathbf{r}) is the predicted color of the pixel.

Hierarchical volume rendering. NeRF also refine the result by allocating samples proportionally to their expected volume distribution based on the coarse estimations. NeRF simultaneously optimizes two MLPs, i.e., the coarse one FΘcF^{c}_{\Theta} and the fine one FΘfF^{f}_{\Theta}. Specifically, NeRF first samples NcN_{c} points and obtain the coarse output C^c(r)\hat{\mathbf{C}}_{c}(\mathbf{r}) by Eq 4, which can be rewrited as C^c(r)=∑i=1Ncwici\hat{\mathbf{C}}_{c}(\mathbf{r})=\sum^{N_{c}}_{i=1}w_{i}\mathbf{c}_{i}. A piecewise-constant PDF related to the sampling points along the ray can be produced by w^=wi/∑j=1Ncwj\hat{w}=w_{i}/\sum^{N_{c}}_{j=1}w_{j}. NeRF then samples NfN_{f} points from this distribution by inverse transform sampling (ITS) and computes the fine outputs C^f(r)\hat{\mathbf{C}}_{f}(r) using all Nc+NfN_{c}+N_{f} sorted sampling points. Let R\mathcal{R} represent the batch, and these two MLPs can be optimized by the following rendering loss:

where g.t.g.t. denotes the ground truth of the rendering pixels.

Method

In this section, we first analyze the limitations of NeRF for rendering medical volumes and elaborate on our motivations. Subsequently, based on our findings, we propose cube-based NeRF (CuNeRF), a novel yet efficient method to deal with “zero-shot” medical image arbitrary-scale super-resolution (MIASSR), which extends NeRF’s application scenarios in the medical domain. Specifically, we first normalize the medical volumes into the range of $byvolumetricnormalization,andthentrainthemodelviaproposeddifferentiablemodules:cube−basedsampling,isotropicvolumerendering,andcube−basedhierarchicalrendering.Duringtraining,CuNeRFisbuildingacoordinate−intensitycontinuousfunctionwhoseinputisa3Dlocationby volumetric normalization, and then train the model via proposed differentiable modules: cube-based sampling, isotropic volume rendering, and cube-based hierarchical rendering. During training, CuNeRF is building a coordinate-intensity continuous function whose input is a 3D location\mathbf{x}==(x,y,z)andtheoutputisthecorrespondingpixelvalueand the output is the corresponding pixel valuec$. After optimization, CuNeRF can predict pixels at any spatial position within the range. As a result, CuNeRF is capable to render medical slices at free viewpoints and arbitrary scales by feeding the corresponding plane equations. Figure 4 depicts the overall framework of our CuNeRF, and the subsequent techniques are described in the following subsections.

As shown in Figure 2, NeRF’s sampling strategy may not be suitable for directly applying to medical volumes, which may produce grid-like artifacts in the results. To explain this limitation, we provide an example of NeRF’s modeling strategies applied to medical volumes in Figure 3 (a). Visually, NeRF is trained to model the volumetric space along the ray emitted by each training pixel. Since medical volumes only contain three orthogonal slices, which differs from multi-view photos collected by conventional cameras, and thus NeRF’s modeling techniques cannot cover the entire representation fields, leaving some “holes” (i.e., unmodeled space) within the regions between adjacent training pixels. Consequently, NeRF may produce sub-optimal results while rendering the contents within the holes. As shown in Figure 2, NeRF†NeRF† is trained on three-orthogonal views. produces grid-like artifacts in upsampling medical volumes, which demonstrates NeRF may struggle to render high-quality HR medical volumes from the corresponding LR ones.

To address the hole-forming issues caused by NeRF’s ray sampling, we introduce cube-based sampling, which samples cubes (3D volumetric space) instead of rays (1D space) to fill the hole regions between adjacent training pixels by the spatial overlaps, as demonstrated in Figure 3 (b). To adapt cube-based sampling, we further propose isotropic volume rendering and cube-based hierarchical rendering modules. These modules will be introduced in the following subsections.

2 Volumetric Normalization

where x^o\hat{\mathbf{x}}_{\mathbf{o}} is set to (0,0,0)(0,0,0) as the center xo=(H2,W2,L2)\mathbf{x}_{\mathbf{o}}=(\frac{H}{2},\frac{W}{2},\frac{L}{2}) of the medical volume. To adapt the positional encoding γ(⋅)\gamma(\cdot) introduced in Eq 1, each positional coordinate xt=(xt,yt,zt)\mathbf{x}_{t}=(x_{t},y_{t},z_{t}) within the medical volume is transformed into the field coordinate x^t=(x^t,y^t,z^t)\hat{\mathbf{x}}_{t}=(\hat{x}_{t},\hat{y}_{t},\hat{z}_{t}). The normalization function N(⋅)\mathcal{N}(\cdot) is formulated as:

where PP is a hyperparameter as the padding size.

3 Cube-based Sampling

Implicit neural representation methods aim to build the continuous representation of medical volumes. However, NeRF suffers from hole-forming issues, which may leave some unmodeled spaces in their representation fields, and thus synthesizes grid-like artifacts in upsampled results. To circumvent the holes forming in the representation fields, we propose a novel sampling strategy: cube-based sampling, which samples cubes (3D volumetric space) instead of rays (1D space). Specifically, for the spatial position x^t\hat{\mathbf{x}}_{t}, CuNeRF samples a set of points within the cube space B(x^t,l2)\mathcal{B}(\hat{\mathbf{x}}_{t},\frac{l}{2}). Each point x^i\hat{\mathbf{x}}_{i} is chosen under the uniform distribution U\mathcal{U} by:

where ll denotes the edge length of the cube. We employ the group of these NN sampling points to approximate the cube space. Due to the spatial overlaps between adjacent cubes, the representation fields can be well-covered by employing the proposed cube-based sampling in optimization. As a result, the representation fields can be densely modeled with the same sampling time as NeRF .

4 Isotropic Volume Rendering

As introduced in 3, the pixel color related to the ray r\mathbf{r} is computed by an integral in Eq 2. Intuitively, the pixel color C(x^t,l)\mathbf{C}(\hat{\mathbf{x}}_{t},l) related to the cube space B(x^t,l2)\mathcal{B}(\hat{\mathbf{x}}_{t},\frac{l}{2}) can be computed by the following triple integral as:

where (x^n,y^n,z^n)=(x^t−l2,y^t−l2,z^t−l2)(\hat{x}_{n},\hat{y}_{n},\hat{z}_{n})=(\hat{x}_{t}-\frac{l}{2},\hat{y}_{t}-\frac{l}{2},\hat{z}_{t}-\frac{l}{2}) denotes the initial location of the triple integral while c(⋅)\mathbf{c}(\cdot) and σ(⋅)\sigma(\cdot) represent the color and volume density functions. However, since NeRF samples NN points to approximate the volume rendering integral of the ray using numerical quadrature in Eq 4, it is required to sample N3N^{3} points to model the cube with the same density, leading to massive computational costs.

where r^=∥(l2,l2,l2)∥p\hat{r}=\|(\frac{l}{2},\frac{l}{2},\frac{l}{2})\|_{p} denotes the max distance of rr within the cube. The derivation detail of Eq 10 is shown in the supplementary materials. Given NN sampling points by the proposed cube-based sampling, CuNeRF first sorts these points by the distance rr. Subsequently, the integral of the cube is approximated via numerical quadrature rules:

where C^(x^t,l)\hat{\mathbf{C}}(\hat{\mathbf{x}}_{t},l) denotes the predicted color of x^t\hat{\mathbf{x}}_{t}.

5 Cube-based Hierarchical Rendering

To refine the results, CuNeRF allocates sampling points proportionally to their expected volume distribution within the cube. Similar to NeRF, CuNeRF also simultaneously optimizes the coarse and fine MLPs. As obtaining the coarse output C^c(x^t,l)\hat{\mathbf{C}}_{c}(\hat{\mathbf{x}}_{t},l), CuNeRF first samples NfN_{f} numbers of rr using ITS. Subsequently, for each rr, we use the hierarchical sampling function ζp(⋅)\zeta_{p}(\cdot) to select important points:

where λ=∥g.t.−C^f(x^t,l)∥12\lambda=\|g.t.-\hat{\mathbf{C}}_{f}(\hat{\mathbf{x}}_{t},l)\|^{\frac{1}{2}} is an adaptive regularization term to alleviate the overfitting brought by the “coarse” part.

6 Medical Slice Synthesis

After optimization, CuNeRF can predict the pixels at any spatial coordinates within the representation fields. Therefore, CuNeRF can represent medical slices with free viewpoints and arbitrary scales by feeding the corresponding plane coordinates. Detailed techniques are described in the following, and we show some examples in Section 5.2.

Free-Viewpoint Rendering. To render a medical slice with the given position x^\hat{\mathbf{x}} and viewpoint d\mathbf{d}, we first construct a base plane Po\mathcal{P}_{o} at x^o\hat{\mathbf{x}}_{o}. Subsequently, we employ the translation matrix MT\mathcal{M}_{T} to move slices from x^o\hat{\mathbf{x}}_{o} to x^\hat{\mathbf{x}}. Finally, since the viewpoint d\mathbf{d} can be represented as rotating ϕ\phi degrees around a certain axis n⊥n_{\perp}, we can obtain the rotation matrix MR\mathcal{M}_{R} via Rodrigues’ rotation formula . Thus, the target plane Pt\mathcal{P}_{t} can be calculated as:

The target medical slices can be obtained by feeding the points sampled within Pt\mathcal{P}_{t} into our CuNeRF.

Arbitrary-Scale Rendering. To render a medical slice with the given sampling scale δ\delta, we first follow the above process to obtain the target plane Pt\mathcal{P}_{t}. Then, we calculate the scale matrix MS\mathcal{M}_{S} based on δ\delta, and sample the points under the translation MSPt\mathcal{M}_{S}\mathcal{P}_{t}. By feeding the sampling points into our CuNeRF, we can obtain the desired medical slices.

Experiments

In this section, we conduct extensive experiments and in-depth analysis to demonstrate the superiority of our CuNeRF in representing high-quality medical images at arbitrary scales. For fair comparisons, the hyperparameters and model settings are consistent in all experiments.

Datasets. We comprehensively compare our CuNeRF and the existing advances in 2 different modalities: CT and MRI. More specifically, we select T1-weighted MRI volumes from the Medical Segmentation Decathlon (MSD) while we also take CT volumes from the 2019 Kidney Tumor Segmentation Challenge (KiTS19) datasets, respectively. All MRI volumes have the same dimension of 240<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>240<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>155240<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>240<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>155. The image size of each CT slice is 512<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>512512<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>512, while the number of CT slices is different.

In experiments, we resize all the CT volumes into 5123512^{3}. All the MRI and CT volumes are normalized into $$. The degradation strategy is the nearest-neighbor interpolation. For the compared supervised methods , we select 50 LR-HR MRI pairs to finetune the pre-trained model of , and also select 150 LR-HR CT pairs to train. The evaluation set consists of 80 medical volumes, including 40 CT and 40 MRI volumes. Note, following the ZSSR settings reported in , we train NeRF† and our CuNeRF on each LR test volume itself (see Figure 1), while the HR volumes are only used for evaluations.

Multi-Layer Perceptron Architecture. Figure 5 depicts MLP’s architecture, where the input is a 3D location x^=(x^,y^,z^)\hat{\mathbf{x}}=(\hat{x},\hat{y},\hat{z}) encoded by γ(⋅)\gamma(\cdot) and the output is a 2D union of pixel intensity c\mathbf{c} and volume density σ\sigma. The parameter size of the proposed MLP is about 0.58M

where φ∼U[0,π]\varphi\sim\mathcal{U}[0,\pi] and θ∼U[0,2π]\theta\sim\mathcal{U}[0,2\pi], respectively. Let p=2p=2 as default, we consider a spherical parameterized as (x^,y^,z^)=ζ2(r,φ,θ)(\hat{x},\hat{y},\hat{z})=\zeta_{2}(r,\varphi,\theta), where φ∈[0,π],θ∈[0,2π],r>0\varphi\in[0,\pi],\theta\in[0,2\pi],r>0. This change of variables from the Cartesian system gives us a differential term:

Therefore, the volume rendering function in Eq 10 can be simplified from Eq 9 as follow:

For training, we employ Adam as the optimizer with a weight decay of 10−610^{-6} and a batch size of 20482048. The maximum iteration is set to 250000250000, and the learning rate is annealed logarithmically from 2<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>10−32<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>10^{-3} to 2<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>10−52<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>10^{-5}. Similar to NeRF, CuNeRF first samples 6464 points for the coarse MLP FΘcF^{c}_{\Theta} and feeds 192192 points (the sorted union of 6464 coarse and 128128 fine points) into the fine MLP FΘfF^{f}_{\Theta}. The training time for each 512<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>512<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>512512<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>512<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>512 volume is about 0.8<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>∼</mo></mrow><annotationencoding="application/x−tex">∼</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.3669em;"></span><spanclass="mrel">∼</span></span></span></span></span>30.8<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>∼</mo></mrow><annotation encoding="application/x-tex">\sim</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.3669em;"></span><span class="mrel">∼</span></span></span></span></span>3 hours.

For testing, the number of sampling points is set to 1616 (88 for coarse MLP and 88 for fine MLP), which can reduce considerable computational costs. The results are obtained by feeding all the coordinates of the given plane equations into our model (seeing details in Section 4.6). The inference time for rendering a 256<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>256<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>256256<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>256<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>256 volume is about 30 secs. Note that we do not use any pre- and post-processing techniques to improve our results in the experiments.

Evaluation Metrics. We use two quantitative metrics: Peak Signal-to-Noise Ratio (PSNR) and Structured Similarity Index (SSIM) to measure the image quality of different methods. Note we report the average SSIM on axial, coronal, and sagittal planes for volumetric MISR.

2 Experimental Results

We compare the proposed CuNeRF with 5 state-of-the-art methods, including 2 supervised MIASSR methods: ArSSR and SAINT , 1 supervised MISR method: TVSRN , 1 conventional method: bicubic interpolation, and NeRF† . Given a upsampling scale δ\delta, we evaluate these methods under the following two settings: (i) 3D MISR. Upsampling the donwsampled volume from Hδ<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>Wδ<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>Lδ\frac{H}{\delta}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>\frac{W}{\delta}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>\frac{L}{\delta} to H<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>W<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>LH<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>W<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>L; (ii) volumetric MISR. Upsampling the donwsampled volume from H<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>W<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>LδH<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>W<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>\frac{L}{\delta} to H<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>W<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>×</mo></mrow><annotationencoding="application/x−tex">×</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6667em;vertical−align:−0.0833em;"></span><spanclass="mord">×</span></span></span></span></span>LH<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>W<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>×</mo></mrow><annotation encoding="application/x-tex">\times</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord">×</span></span></span></span></span>L. Note that NeRF† and CuNeRF are trained with the same settings and similar parameter size (±\pm0.02M).

Quantitative Comparison. We report 3D MISR and volumetric MISR based on the increasing upsampled scales in Table 1 and Table 2, respectively. As demonstrated, for the 3D MISR challenge on MRI volumes, CuNeRF surpasses all the competitors with a consistent preferable performance at various upsampling scales. For volumetric MISR challenge on CT volumes, CuNeRF achieves comparable performance to SAINT and TVSRN . Compared to fully-supervised MIASSR methods: ArSSR and SAINT , our CuNeRF is more robust at presenting large-scale medical slices and capable to deal with different modalities (CT and MRI), suggesting CuNeRF owns broader application scenarios. It is also worth noting that NeRF† achieves comparable performance for volumetric MISR but fails in 3D MISR. Since volumetric MISR only aims to acquire the pixels along the zz-axis, the experimental results of NeRF† confirm our motivations.

Visual Comparison. We visualize the rendering results of CuNeRF and other competitors on MRI (rows 1 and 2) and CT (rows 3 and 4) modalities in Figure 6. It can be observed that CuNeRF well represents the medical slices at various scales. Compared to the exhibited methods, CuNeRF is most similar to the ground truths, achieving better visual verisimilitude and reducing aliasing artifacts, especially in representing large-scale medical slices. Since NeRF† exhibits grid-like artifacts in rendering high-quality medical slices at larger-valued scales, the visualization results prove the effectiveness of CuNeRF, which extends NeRF’s capability to continuously represent medical images.

Free-Viewpoint & Arbitrary-Scale Rendering. As shown in Figure 7, CuNeRF can synthesize medical images at continuous-valued scales (a). Moreover, CuNeRF is capable to yield medical slices with a viewpoint rotating 360 degrees around an arbitrary coordinate axis n⊥n_{\perp}. Compared to existing methods, CuNeRF is capable to provide richer visual information for clinical diagnosis.

3 Ablation Study

In this subsection, we conduct comprehensive experiments to prove the correctness of CuNeRF’s design. We first carry out ablation studies to investigate the effectiveness of the proposed modules. Subsequently, we evaluate the CuNeRF’s performance under different settings.

CuNeRF’s ablation variants. We evaluate against several ablations of the proposed CuNeRF with each module: CuS, IVR and LA\mathcal{L}_{A} represent cube-based sampling, isotropic volume rendering, and adaptive rendering loss, respectively. The baseline model here is NeRF†. As reported in Table 3, the baseline model struggles to deal with 3D MISR issues (row 1), while adopting CuS instead of ray sampling can significantly improve the performance (row 2). Compared to NeRF’s volume rendering function, employing IVR (row 3) can further improve the slice synthesis quality, suggesting IVR can better estimate the volumetric distribution, reducing aliasing artifacts raised by undersampling. Since the coarse term of NeRF’s rendering loss may affect the optimization, LA\mathcal{L}_{A} (row 4) is able to alleviate this distraction.

Conclusion

In this paper, we present Cube-based NeRF (CuNeRF), a zero-shot framework for medical image arbitrary-scale super-resolution (MIASSR). Instead of learning the mapping between LR-HR pairs, CuNeRF learns the continuous volumetric representation from LR volumes, thus a well-trained model can yield medical images at arbitrary viewpoints and scales in a continuous domain. Extensive experiments demonstrate that CuNeRF outperforms state-of-the-art methods, yielding better visual effects and reducing artifacts at various upsampling factors.

Acknowledgement

This project is supported by the Natural Science Foundation of China (No. 62072482).

References