MambaVC: Learned Visual Compression with Selective State Spaces
Shiyu Qin, Jinpeng Wang, Yimin Zhou, Bin Chen, Tianci Luo, Baoyi An, Tao Dai, Shutao Xia, Yaowei Wang
Introduction
Visual compression is a long-standing problem in multimedia processing. In the past few decades, classical standards have dominated for a long time. With the advent of deep neural architectures like CNNs and Transformers , learned compression methods have emerged and shown ever-improving performance, gaining increasing interest over traditional ones.
The core of visual compression is the neural network design to eliminate redundant information and capture content distribution, where it naturally presents a dilemma between rate-distortion optimization and model efficiency. While CNN-based methods remain popular in many resource-limited scenarios thanks to the hardware-efficient convolution operators, their local receptive field limits global context modeling capacity and thus restricts compression performance. In contrast, Transformer-based methods excel in the global perception with attention mechanisms and thereby benefit redundancy reduction. However, their quadratic complexity in computation and memory raises efficiency concerns. Although some hybrid approaches like TCM combine CNNs and Transformers to balance compression efficacy and efficiency, it is not a sustainable direction for further development. Unlike prior work, we are committed to exploring promising solutions beyond engineering trade-offs toward this issue.
Recently, state space models (SSMs) , particularly the structured variants (S4) , have been extensively studied. Mamba stands out as a representative work, whose data-dependent selective mechanism enhances critical information extraction while eliminating irrelevant noise from the input. This hints that Mamba-based models can effectively gather global context and thus enjoy advantages for compression. Furthermore, Mamba integrates structured reparameterization tricks and utilizes a hardware-efficient parallel scanning algorithm, assuring faster training and inference on GPUs. These compelling features inspire us to investigate Mamba’s potential for visual compression.
In this paper, we introduce MambaVC, a simple, strong and efficient visual compression network with selective state spaces. Inspired by Liu et al. , we design a visual state space (VSS) block as the nonlinear activation function after each downsampling in the neural compression network, which integrates a specialized 2D selective scanning (2DSS) mechanism for spatial modeling. The 2DSS performs selective scanning along 4 pre-defined traverse paths in parallel, which helps to capture comprehensive global contexts and facilitates effective and efficient compression.
We conduct extensive experiments on image and video benchmark datasets. Without the bells and whistles, MambaVC achieves a superior rate-distortion trade-off with lower computational and memory overheads compared to CNN- and Transformer-based counterparts, some as demonstrated in Figure 1(a). More encouragingly, we show that MambaVC exhibits even stronger performance on high-resolution image compression, as shown in Figure 1(b). These favorable results are consistent with SSM’s efficient long-range modeling capacity, shedding light on its potential in many important yet challenging applications, such as compressing high-definition medical images and transmitting high-resolution satellite imagery. We also compare and analyze different designs from various aspects, including spatial redundancy, effective receptive field, and information loss in the compression process, to facilitate a comprehensive understanding of MambaVC’s efficacy.
In summary, our contributions are as follows:
We develop MambaVC, the first visual compression network with selective state spaces. The designed 2DSS improves global context modeling and helps effective and efficient compression.
Extensive experiments on benchmark datasets show superior performance and competitive efficiency of MambaVC on image and video compression. The strong results highlight a new promising direction of compression network design beyond CNNs and Transformers.
We showcase MambaVC’s particular effectiveness and scalability in high-resolution compression, prompting its potential in many important but challenging applications.
We compare and analyze different network designs thoroughly, showing the MambaVC’s advantages regarding various aspects to validate and understand its effectiveness.
Related Works
Learned Visual Compression In the past decade, learned visual compression has demonstrated remarkable potential and made a significant impression. The prevailing methods can be categorized into CNN-based and Transformer-based approaches. Early works, such as CNNs with generalized divisive normalization (GDN) layers , achieved good performance in image compression. Attention mechanisms and residual blocks were integrated into the VAE architecture later. However, the limited receptive field constrained the further development of these models. With the explosion of Vision Transformers , Transformer-based compression models have shown strong competitiveness. Yet, their substantial computational and storage demands are daunting. Recent efforts have attempted to combine the strengths of both approaches, but led to even increased computational complexity as shown in Figure 1(a). The trade-off between model performance and efficiency remains a pressing issue that needs to be addressed.
State Space Models SSMs are recently proposed models combined with deep learning to capture the dynamics and dependencies of long-sequence data. LSSL first leverages linear state space equations for modeling sequence data. Later, the structured state-space sequence model (S4) employs a linear state space for contextualization and shows strong performance on various sequence modeling tasks, especially with lengthy sequences. Building on it, numerous have been proposed, and Mamba stands out with its data dependency and parallel scanning. Many works have consequently extended Mamba from Natural Language Processing (NLP) to the vision domain such as image classification , multimodal Learning and others . However, the application of the Mamba for visual compression remains unexplored. In this work, we explore how to transfer the success of Mamba to build effective and efficient compression models.
Method
Modern SSMs approximate this continuous-time ODE through discretization. Concretely, they discretize the continuous parameters and by a timescale , using the zero-order hold trick:
Then the discretized version of eq. 1 is reformulated as follows:
Mamba further incorporates data-dependence to , and , enabling an input-aware selective mechanism for better state-space modeling. While the recurrent nature restricts the fully parallel capacity, Mamba ingeniously implements structural reparameterization tricks and the hardware-efficient parallel scanning algorithm to compensate for the overall efficiency.
2 The proposed MambaVC
We illustrate the architecture of MambaVC in Figure 2(a). Given an image , we first obtain the latent and hyper latent using the encoder and the hyper encoder , respectively:
At the decoder side, we first use a hyper decoder to obtain the initial mean and variance:
Then we divide the latent to slices and compute slice-wise information by:
Next, we use the decoder to reconstruct image from the quantized latent :
Finally, we optimize the following training objectives:
where is the Lagrangian multiplier to control the rate-distortion trade-off.
2.2 Visual State Space (VSS) Block
Analogously, the gating branch computes the weight vector by:
Finally, the two branches are combined to produce the output feature map:
where denotes the element-wise product.
2.3 2D Selective Scan (2DSS)
We then apply reversed operations to the contextual token sequences by the following folding patterns:
In the end, we merge the transformed feature maps to obtain the output feature map:
2.4 Extension to Video Compression
We also extend MambaVC to video compression to explore its potential. Here we choose the scale-space flow (SSF) , a renowned learned video compression model, as the base framework for extension. We upgrade the CNN-based transforms in 3 parts (i.e., I-frame compression, scale-space flow, and residual) of SSF with the developed VSS blocks. We call this extension by MambaVC-SSF. We will show and discuss the experimental results in Section 4.4.
Experiments
For image compression, we select the Flickr30k dataset in , consisting of 31,783 images. Each model is trained for 2M steps. For the first 1.2M steps, each batch consists of 8 randomly cropped 256256 images; for the next 0.8M steps, each batch includes 2 randomly selected 512512 upsampled images. The learning rate starts at 10-4 and drops to 10-5 at 1.8M steps, finally drops to 10-6 at 1.95M steps. We employe in rate-distortion loss.
For video compression, models are all trained on Vimeo-90k for 1M steps at a learning rate of 10-4 and an additional 0.6M steps at 10-5. In the first phase, each batch contains 8 randomly cropped 256256 images; in the second phase, each batch contains 8 randomly cropped 384256 images. We optimize video model for MSE distortion metric. In particular, we use . Inspired by , we process each video sequence in original and reversed order respectively during each optimization step.
1.2 Baselines
We conduct a comprehensive and thorough evaluation of MambaVC in two directions on Kodak , CLIC2020 , JPEG-AI and UHD with different image resolution. First, we compare it with state-of-the-art methods, including MLIC+ , Mixed , GLLMM , QResVAE , ELIC , STF , WACNN , Entroformer , Swin-ChARM , Invcompress and traditional coding method BPG444 and VTM-15.0 . Secondly, we validate the superiority of MambaVC over its convolutional and Transformer variants in terms of performance and efficiency. Specifically, we replace the VSS Block in MambaVC with swin transformer and GDN layer, respectively, naming them SwinVC and ConvVC. Detailed structures are shown in Appendix A.
Meanwhile, we evaluate variant SSF on MCL-JCV and UVG , comparing it with standard codecs AVC(x264), HEVC(x265) and the test model implementation of HEVC, called HEVC (HM). All methods fix the GOP size to 12.
2 Standard Image Compression
The rate-distortion performance on Kodak dataset is shown in Figure 3. For fairness, all shown learned methods are optimized for minimizing MSE. Both PSNR and MS-SSIM are tested to demonstrate the robustness of MambaVC. Compared to the previous best methods MLIC+ , our approach yields an average PSNR improvement of 0.1 dB, while requiring only half the computational complexity and 60% of the memory overhead as shown in Figure 1(a).
2.2 Comparison of Variants
The RD curves for all variants on Kodak are shown in Figure 11(a). To provide a clearer comparison of the performance among different variants, Figure 11(b) illustrates the percentage of rate savings relative to VTM-15.0 for achieving equivalent PSNR. Figure 11 demonstrates that MambaVC consistently outperforms SwinVC and ConvVC in various scenarios. SwinVC, as highlighted in previous work, surpasses ConvVC. Both MambaVC and SwinVC exhibit higher compression efficiency compared to VTM-15.0, whereas ConvVC falls short. As the rate increase, SwinVC’s performance advantage slightly diminishes, while MambaVC remains unaffected. In Table 1, we present the BD-rate of different variants compared to VTM-15.0 across four datasets. MambaVC achieves an average bitrate savings of 13.35%, while SwinVC achieves an average savings of 1.94%. In contrast, ConvVC consumes an average of 4.76% more bits. Notably, MambaVC is the only variant that surpasses VTM-15.0 on UHD , highlighting its potential for high-resolution images, which will be discussed in the next section. See Section C.1 for further details.
3 High-Resolution Image Compression
Recent work has demonstrated Mamba’s advantages in long-range modeling. To explore this potential in visual compression, we compare our MambaVC against SwinVC and ConvVC on images of varying resolutions in two ways. Specifically, we downsample high-resolution images from the UHD by different factors to create multiple sets of images with the same distribution but different sizes. As shown in Figure 1(b), MambaVC saves more bits as the resolution increases compared to the other variants. To mitigate the impact of specific dataset distributions, we test across four datasets with different resolutions. As indicated in Table LABEL:SwinVC_ConvVC_BD-rate, the performance advantage of MambaVC on the high-resolution UHD is significantly greater than on the lower-resolution Kodak . For datasets with similar sizes, like CLIC2020 and JPEG-AI , the performance advantage is relatively consistent. We also record the change in computational cost across different resolutions. As shown in Table LABEL:SwinVC_ConvVC_Computational_complexity, with increasing image sizes, the computational gap widened from an initial 0.23 TMACs and 0.1 TMACs to a final 12.96 TMACs and 5.46 TMACs, separately. These results indicate that MambaVC has a distinct advantage in compressing high-resolution images. This potential may influence the future development of specialized fields such as medical imaging and satellite imagery.
4 Video Compression with SSF Backbone
Following the configuration of Agustsson et al. , we evaluated our method on the MCL-JCV and UVG datasets. To ensure a more comprehensive comparison, we also construct the CNN- and Swin-Transformer-based counterparts with MambaVC-SSF, denoted as SwinVC-SSF and ConvVC-SSF, respectively. Detailed configurations for different models can be found in Sections 4.1.1 and MambaVC: Learned Visual Compression with Selective State Spaces. Figure 4 presents the RD curves of MambaVC-SSF with its different variants and traditional methods. The mamba-based model outperforms its convolutional and transformer counterparts. However, the performance improvement in video compression is not as pronounced as in image compression, possibly because merely changing the nonlinear transformation structure is insufficient to capture more redundancy. Additionally, all variants still fall short of HM in performance on the MCL-JCV dataset, indicating significant room for further improvement.
5 Computational and Memory Efficiencies
To explore the advantage of Mamba’s linear complexity in visual compression, we evaluate the memory overhead and computational complexity on the Kodak dateset . As results shown in Table 4, MambaVC exhibits the best performance across different variants. While MLIC+ incurs greater computational cost due to its adoption of a more advanced entropy model, it doesn’t achieve superior performance. On the other hand, method combining convolution and transformer, while falling short in both computational and storage aspects compared to SwinVC and ConvVC, further underscores the significance of MambaVC as a novel framework.
6 Analysis
Learned visual compression redundancy removal involves two key steps: nonlinear encoding transform and using a conditionally factorized Gaussian prior distribution to decorrelate the latent .
Specifically, the former converts the input signal from the image domain to the feature domain, while the latter uses a hyper network to learn the mean and variance of latent , assuming a Gaussian distribution, to further reduce correlation. As various correlations and redundancies are eliminated, less information needs to be entropy coded, thereby improving compression efficiency. To this end, we visualized the correlation between each spatial pixel in and its surrounding positions, which we refer to as latent correlation. Figure 5 indicates that MambaVC has lower correlations at all distances compared to SwinVC and ConvVC. Theoretically, decorrelated latent should follow a standard normal distribution (SND). To verify this, we fit the distribution curves for different methods and calculated the KL divergence from SND, as shown in Figure 6. The curve for MambaVC is noticeably closer to the SND with a smaller KL divergence , which indicates the Mamba-based hyper network can learn more accurately. We also investigate the hyper latent correlation and the relationship between and correlation, as shown in Figure 14.
6.2 Effective Receptive Field
The effective receptive field (ERF) denotes the region of the input that a neuron in a neural network "perceives". A larger receptive field enables the network to capture related information from a wider area. This characteristic aligns perfectly with the nonlinear encoder in visual compression, as it reduces redundancy in images through feature extraction and dimensionality reduction. Consequently, we are keenly interested in examining the receptive field sizes of MambaVC and its variants. As shown in Figure 7, MambaVC is the only model with a global ERF, while ConvVC has the smallest receptive field. This confirms that in high-resolution scenarios, MambaVC can leverage more pixels globally to eliminate redundancy, whereas SwinVC and ConvVC, with their limited receptive fields, can only utilize local information, leading to performance differences.
6.3 Quantize Deviation
Conclusions
In this paper, we introduced MambaVC, the first visual compression network based on the state-space model. MambaVC built a visual state space (VSS) block with 2D selective scanning (2DSS) mechanism to improve global context modeling and content compression. Experimental results showed that MambaVC achieves superior rate-distortion performance compared to CNN and Transformer variants while maintaining computational and memory efficiencies. These advantages are even more pronounced with high-resolution images, highlighting MambaVC’s potential and scalability in real-world applications. Compared to other designs, MambaVC exhibits stronger redundancy elimination, larger receptive fields, and lower quantization loss, revealing its comprehensive advantages for compression. We hope MambaVC can offer a basis for exploring SSMs in compression and inspire future works.
References
Appendix A Model Configurations
MambaVC The detailed architecture has been delineated in Section 3.2. For the number of channels and layers, we set them as and , respectively. Due to the high resolution of images in UHD, which slows down inference, we randomly select 20 images from the UHD dataset and crop their length to 3328 pixels along the center for use as the test set.
MambaVC-SSF For encoder/decoder and hyper encoder/decoder in SSF , there is a VSS Block following each upsampling or downsampling operation, except when generating the reconstructed image or latent with layer number .
A.2 Convolutional Variant
ConvVC The architecture of ConvVC are shown in Figure 9. Specifically, we replaced the VSS Block with the popular GDN layer , which has been proven effective in Gaussianizing the local joint statistics of natural images. To compensate for the limited effective receptive field of convolutions, we set all convolutional kernels to a size of 5. For architecture, our base model has the following parameters: .
A.3 Transformer Variant
SwinVC Among a large number of vision transformer variants, we select Swin Transformer as network components for its lower complexity and superior modeling capability. As shown in Figure 10, the layer number and window size are common to all experiments. For channels, we set .
SwinVC-SSF The original downsampling modules remain untouched. Following the structure akin to the image model, we utilize the Swin Transformer , albeit without any LayerNorm, instead appending a ReLU layer afterward. Both latent and hyper latent channels are set at 192. For I-frame compression, scale-space flow, and residual, we employ window sizes of 8, 4, and 8, respectively. The layer number is the same as MambaVC-SSF.
Appendix B Classical Standards
In this section, we provide the evaluation script used for traditional methods.
BPG444: We get BPG software from http://bellard.org/bpg/ and use command as follows:
VTM-15.0: VTM is sourced from https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM. The command is:
B.2 Video Compression
Appendix C More Results
Additional rate-distortion results on Kodak ,CLIC2020 and JPEG-AI are shown in Figure 11, Figure 12 and Figure 13.
C.2 Hyper Latent Correlation
Figure 14 illustrates the spatial correlation of the normalized prior latents. Horizontally comparing the different methods, MambaVC consistently shows the best performance across all . Vertically comparing the results, as the decreases, the proportion of distortion loss diminishes, leading the model to focus more on compression ratio and thus eliminate more redundancy.
C.3 Effective Receptive Field
In Figure 7, we present the receptive fields of latent after passing through the encoder . Additionally, we explore the receptive fields of the hyper latent after passing through the hyper encoder , as shown in Figure 15. Vertically comparing the methods, we observe that the receptive field expands as the network depth increases, suggesting a greater influence of surrounding areas on the value of each spatial point. Horizontally comparing the methods, MambaVC consistently demonstrates the largest receptive field among all approaches.