ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders

Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, Saining Xie

Introduction

Building on research breakthroughs in earlier decades , the field of visual recognition has ushered in a new era of large-scale visual representation learning. Pre-trained, large-scale vision models have become essential tools for feature learning and enabling a wide range of vision applications. The performance of a visual representation learning system is largely influenced by three main factors: the neural network architecture chosen, the method used for training the network, and the data used for training. In the field of visual recognition, progress in each of these areas contributes to overall improvements in performance.

Innovation in neural network architecture design has consistently played a major role in the field of representation learning. Convolutional neural network architectures (ConvNets) have had a significant impact on computer vision research by allowing for the use of generic feature learning methods for a variety of visual recognition tasks , rather than relying on manual feature engineering. In recent years, the transformer architecture , originally developed for natural language processing, has also gained popularity due to its strong scaling behavior with respect to model and dataset size . More recently, ConvNeXt architecture has modernized traditional ConvNets and demonstrated that pure convolutional models could also be scalable architectures. However, the most common method for exploring the design space for neural network architectures is still through benchmarking supervised learning performance on ImageNet.

In a separate line of research, the focus of visual representation learning has been shifting from supervised learning with labels to self-supervised pre-training with pre-text objectives. Among many different self-supervised algorithms, masked autoencoders (MAE) have recently brought success in masked language modeling to the vision domain and quickly become a popular approach for visual representation learning. However, a common practice in self-supervised learning is to use a predetermined architecture designed for supervised learning, and assume the design is fixed. For instance, MAE was developed using the vision transformer architecture.

It is possible to combine the design elements of architectures and self-supervised learning frameworks, but doing so may present challenges when using ConvNeXt with masked autoencoders. One issue is that MAE has a specific encode-decoder design that is optimized for the sequence processing capabilities of transformers, which allows the compute-heavy encoder to focus on visible patches and thus reduce the pre-training cost. This design may not be compatible with standard ConvNets, which use dense sliding windows. Additionally, if the relationship between the architecture and the training objective is not taken into consideration, it may be unclear whether optimal performance can be achieved. In fact, previous research has shown that training ConvNets with mask-based self-supervised learning can be difficult , and empirical evidence suggests that transformers and ConvNets may have different feature learning behaviors that can affect representation quality.

To this end, we propose to co-design the network architecture and the masked autoencoder under the same framework, with the aim of making mask-based self-supervised learning effective for ConvNeXt models and achieving results similar to those obtained using transformers.

In designing the masked autoencoder, we treat the masked input as a set of sparse patches and use sparse convolutions to process only the visible parts. The idea is inspired by the use of sparse convolutions in processing large-scale 3D point clouds . In practice, we can implement ConvNeXt with sparse convolutions, and at fine-tuning, the weights are converted back to standard, dense layers without requiring special handling. To further improve the pre-training efficiency, we replace the transformer decoder with a single ConvNeXt block, making the entire design fully convolutional. We have observed mixed results with these changes: the learned features are useful and improve upon the baseline results, but the fine-tuning performance is still not as good as the transformer-based model.

We then conduct a feature space analysis of different training configurations for ConvNeXt. We identify a potential issue of feature collapse at the MLP layer when training ConvNeXt directly on masked input. To address this issue, we propose adding a Global Response Normalization layer to enhance inter-channel feature competition. This change is most effective when the model is pre-trained with masked autoencoders, suggesting that reusing a fixed architecture design from supervised learning may be suboptimal.

In summary, we introduce ConvNeXt V2 which demonstrates improved performance when used in conjunction with masked autoencoders. We have found that this model significantly improves the performance of pure ConvNets across various downstream tasks, including ImageNet classification , COCO object detection and ADE20K segmentation . The ConvNeXt V2 models can be used in a variety of compute regimes and includes models of varying complexity: from an efficient 3.7M-parameter Atto model that achieves 76.7% top-1 accuracy on ImageNet, to a 650M Huge model that reaches a state-of-the-art 88.9% accuracy when using IN-22K labels.

Related Work

The design of ConvNets, which were first introduced in the 1980s and trained using back-propagation, has undergone numerous improvements in terms of optimization, accuracy, and efficiency over the years . These innovations have mainly been discovered through the use of supervised training on the ImageNet dataset. In recent years, some efforts have been made to perform architecture search using self-supervised pre-text tasks such as rotation prediction and colorization, as in the case of UnNAS . Recently, ConvNeXt conducted a comprehensive review of the design space and demonstrated pure ConvNets can be as scalable as the vision transformers , which have become the dominant architecture in many applications. ConvNeXt has particularly excelled in scenarios requiring lower complexity . Our ConvNeXt V2 model, which is powered by self-supervised learning, provides a simple way to upgrade existing models and achieve a significant boost in performance across a wide range of use cases.

Masked Autoencoders.

Masked image modeling, represented by masked autoencoders , is one of the latest self-supervised learning strategies. As a neural network pre-training framework, masked autoencoders have shown a broad impact on visual recognition. However, original masked autoencoders are not directly applicable to ConvNets due to their asymmetric encoder-decoder design. Alternative frameworks such as have attempted to adapt the approach for use with ConvNets, but with mixed results. MCMAE uses a few convolutional blocks as input tokenizers. To the best of our knowledge, there are no pre-trained models that show self-supervised learning can improve upon the best ConvNeXt supervised results.

Fully Convolutional Masked Autoencoder

Our approach is conceptually simple and runs in a fully convolutional manner. The learning signals are generated by randomly masking the raw input visuals with a high masking ratio and letting the model predict the missing parts given the remaining context. Our framework is illustrated in Figure 2, and we will now describe its main components in more detail.

We use a random masking strategy with a masking ratio of 0.6. As the convolutional model has a hierarchical design, where the features are downsampled in different stages, the mask is generated in the last stage and upsampled recursively up to the finest resolution. To implement this in practice, we randomly remove 60% of the 32×3232\times 32 patches from the original input image. We use minimal data augmentation, only including random resized cropping.

Encoder design.

We use ConvNeXt model as the encoder in our approach. One challenge in making masked image modeling effective is preventing the model from learning shortcuts that allow it to copy and paste information from the masked regions. This is relatively easy to prevent in transformer-based models, which can leave the visible patches as the only input to the encoder. However, it is more difficult to achieve this with ConvNets, as the 2D image structure must be preserved. While naive solutions involve introducing learnable masked tokens in the input side , these approaches decrease the efficiency of pre-training and result in a train and test time inconsistency, as there are no mask tokens at test time. This becomes especially problematic when the masking ratio is high.

To tackle this issue, our new insight is to view the masked image from a “sparse data perspective”, which was inspired by learning on sparse point clouds in 3D tasks . Our key observation is that the masked image can be represented as a 2D sparse array of pixels. Based on this insight, it is natural to incorporate sparse convolution into our framework to facilitate pre-training of the masked autoencoder. In practice, during pre-training, we propose to convert the standard convolution layer in the encoder with the submanifold sparse convolution, which enables the model to operate only on the visible data points . We note that the sparse convolution layers can be converted back to standard convolution at the fine-tuning stage without requiring additional handling. As an alternative, it is also possible to apply a binary masking operation before and after the dense convolution operation. This operation has numerically the same effect as sparse convolutions, is theoretically more computationally intensive, but can be more friendly on AI accelerators like TPU.

Decoder design.

We use a lightweight, plain ConvNeXt block as the decoder. This forms an asymmetric encoder-decoder architecture overall, as the encoder is heavier and has a hierarchy. We also considered more complex decoders such as hierarchical decoders or transformers , but the simpler single ConvNeXt block decoder performed well in terms of fine-tuning accuracy and reduced pre-training time considerably, demonstrated in Table 1. We set the dimension of the decoder to 512.

Reconstruction target.

We compute the mean squared error (MSE) between the reconstructed and target images. Similar to MAE , the target is a patch-wise normalized image of the original input, and the loss is applied only on the masked patches.

FCMAE.

We now present a Fully Convolutional Masked AutoEncoder (FCMAE) by combining the proposals described above. To evaluate the effectiveness of this framework, we use the ConvNeXt-Base model as the encoder and conduct a series of ablation studies. Throughout the paper, we focus on the end-to-end fine-tuning performance becuase of its practical relevance in transfer learning, and use that to assess the quality of the learned representation.

We pre-train and fine-tune using the ImageNet-1K (IN-1K) dataset for 800 and 100 epochs, respectively, and report the top-1 IN-1K validation accuracy for a single 224×224 center crop. Additional details about the experimental setup can be found in the appendix.

To understand the impact of using sparse convolution in our FCMAE framework, we first investigate how it affects the quality of the learned representation during masked image pre-training. Our empirical findings show that it is essential to prevent information leakage from the masked region in order to achieve good results.

Next, we compare our self-supervised approach to supervised learning. Specifically, we obtain two baseline experimental results: the supervised 100 epoch baseline using the same recipe and the 300 epoch supervised training baseline provided in the original ConvNeXt paper . We find that our FCMAE pre-training provides better initialization than the random baseline (i.e., 82.7 →\rightarrow 83.7), but it still needs to catch up to the best performance obtained in the original supervised setup.

This is in contrast to the recent success of masked image modeling using transformer-based models , where the pre-trained models significantly outperform the supervised counterparts. This motivates us to investigate the unique challenges faced by the ConvNeXt encoder during masked autoencoder pre-training, which we discuss next.

Global Response Normalization

In this section, we introduce a new Global Response Normalization (GRN) technique to make FCMAE pre-training more effective in conjunction with the ConvNeXt architecture. We first motivate our approach through both qualitative and quantitative feature analyses.

To gain more insight into the learning behavior, we first perform qualitative analysis in the feature space. We visualize the activations of a FCMAE pre-trained ConvNeXt-Base model and notice an intriguing “feature collapse” phenomenon: there are many dead or saturated feature maps and the activation becomes redundant across channels. We show some of the visualizations in Figure 3. This behavior was mainly observed in the dimension-expansion MLP layers in a ConvNeXt block .

Feature cosine distance analysis.

To further validate our observation quantitatively, we perform a feature cosine distance analysis. Given an activation tensor X∈RH×W×CX\in R^{H\times W\times C}, Xi∈RH×WX_{i}\in R^{H\times W} is the feature map of the ii-th channel. We reshape it as a HWHW dimensional vector and compute the average pair-wise cosine distance across the channels by 1C2∑iC∑jC1−cos(Xi,Xj)2\frac{1}{C^{2}}\sum_{i}^{C}\sum_{j}^{C}\frac{1-{cos}(X_{i},X_{j})}{2}. A higher distance value indicates more diverse features, while a lower value indicates feature redundancy.

To perform this analysis, we randomly select 1,000 images from different classes in the ImageNet-1K validation set and extract the high-dimensional features from each layer of different models, including the FCMAE models, the ConvNeXt supervised model and the MAE pre-trained ViT model . We then compute the distance per layer for each image and average the values across all images. The results are plotted in Figure 4. The FCMAE pre-trained ConvNeXt model exhibits a clear tendency towards feature collapse, consistent with our observations from the previous activation visualizations. This motivates us to consider ways to diversify the features during learning and prevent feature collapse.

Approach.

There are many mechanisms in the brain that promote neuron diversity. For example, lateral inhibition can help to sharpen the response of the activated neuron and increase the contrast and selectivity of individual neurons to the stimulus while also increasing the diversity of responses across the population of neurons. In deep learning, this form of lateral inhibition can be implemented by response normalization . In this work, we introduce a new response normalization layer called global response normalization (GRN), which aims to increase the contrast and selectivity of channels. Given an input feature, X∈RH×W×CX\in R^{H\times W\times C}, the proposed GRN unit consists of three steps: 1) global feature aggregation, 2) feature normalization, and 3) feature calibration.

First, we aggregate a spatial feature map XiX_{i} into a vector gxgx with a global function G(⋅)\mathcal{G}(\cdot):

This can be viewed as a simple pooling layer. We experimented with different functions in Table LABEL:tab:grn_spool. Interestingly, global average pooling, a widely used feature aggregator , did not perform well in our case. Instead, we found that using norm-based feature aggregation, specifically, using L2-norm, resulted in better performance. This gives us a set of aggregated values G(X)=gx={∣∣X1∣∣,∣∣X2∣∣,…,∣∣XC∣∣}∈RC\mathcal{G}(X)=gx=\{||X_{1}||,||X_{2}||,\ldots,||X_{C}||\}\in\mathcal{R}^{C} where G(X)i=∣∣Xi∣∣\mathcal{G}(X)_{i}=||X_{i}|| is a scalar that aggregates the statistics of the i-th channel.

Next, we apply a response normalization function N(⋅)\mathcal{N}(\cdot) to the aggregated values. Concretely, we use a standard divisive normalization as follows,

where ∣∣Xi∣∣||X_{i}|| is the L2-norm of the ii-th channel. To account for the increased number of channels at deeper layers, in practice, we also scale the normalized value by the channel count CC. Intuitively, for the i-th channel, Eqn. 2 computes its relative importance compared to all the other channels. Similar to other forms of normalization , this step creates a feature competition across channels by mutual inhibition. In Table LABEL:tab:grn_cnorm, we also examine the use of other normalization functions and find that the simple divisive normalization works best, though standardization (∣∣Xi∣∣−μ)/σ(||X_{i}||-\mu)/\sigma yields similar results when applied to the same L2-norm aggregated values.

Finally, we calibrate the original input responses using the computed feature normalization scores:

The core GRN unit is very easy to implement, requiring only three lines of code, and has no learnable parameters. The pseudo-code for the GRN unit is in Algorithm 1.

To ease optimization, we add two additional learnable parameters, γ\gamma and β\beta, and initialize them to zero. We also add a residual connection between the input and output of the GRN layer. The resulting final GRN block is Xi=γ∗Xi∗N(G(X)i)+β+XiX_{i}=\gamma*X_{i}*\mathcal{N}(\mathcal{G}(X)_{i})+\beta+X_{i}. This setup allows a GRN layer to initially perform an identity function and gradually adapt during training. The importance of residual connection is demonstrated in Table LABEL:tab:grn_residual.

ConvNeXt V2.

We incorporate the GRN layer into the original ConvNeXt block, as illustrated in Figure 5. We empirically found that LayerScale becomes unnecessary when GRN is applied and can be removed. Using this new block design, we create various models with varying efficiency and capacity, which we refer to as the ConvNeXt V2 model family. These models range from lightweight (e.g. Atto ) to compute-intensive (e.g. Huge) ones. Detailed model configurations can be found in the appendix.

Impact of GRN.

We now pre-train ConvNeXt V2 using the FCMAE framework and evaluate the impact of GRN. From visualization in Figure 3 and cosine distance analysis in Figure 4, we can observe that ConvNeXt V2 effectively mitigates the feature collapse issue. The cosine distance values are consistently high, indicating that feature diversity is maintained across layers. This behavior is similar to that of the MAE pre-trained ViT model . Overall, this suggests that ConvNeXt V2 learning behavior can resemble ViT, under a similar masked image pre-training framework.

Next, we evaluate the fine-tuning performance.

When equipped with GRN, the FCMAE pre-trained model can significantly outperform the 300 epoch supervised counterpart. GRN improves the representation quality by enhancing the feature diversity, which was absent in the V1 model but has proven crucial for masked-based pre-training. Note this improvement is achieved without adding additional parameter overhead or increased FLOPS.The additional affine parameters γ\gamma/β\beta are negligible.

Relation to feature normalization methods.

Can other normalization layers perform as well as the global response normalization (GRN) layer? In Table LABEL:tab:grn_vs_norms, we compare GRN with the three widely used normalization layers: Local Response Normalization (LRN) , Batch Normalization (BN) , and Layer Normalization (LN) . We observe that only GRN can significantly outperform the supervised baseline. LRN lacks global context as it only contrasts channels within nearby neighbors. BN normalizes spatially along the batch axis, which is unsuitable for masked inputs. LN implicitly encourages feature competition through global mean and variance standardization but does not work as well as GRN.

Relation to feature gating methods.

Another way to enhance competition across neurons is to use dynamic feature gating methods . In Table LABEL:tab:grn_vs_gates, we compare our GRN with two classic gating layers: squeeze-and-excite (SE) and convolutional block attention module (CBAM) . SE focuses on channel gating, while CBAM focuses on spatial gating. Both modules can increase the contrast of individual channels, similar to what GRN does. GRN is much simpler and more efficient as it does not require additional parameter layers (such as MLPs).

The role of GRN in pre-training/fine-tuning.

Finally, we examine the importance of GRN in pre-training and fine-tuning. We present results in Table LABEL:tab:grn_ptft where we either remove GRN from fine-tuning or add newly initialized GRN only at the time of fine-tuning. Either way, we observe a significant performance degradation, suggesting that keeping GRN in both pre-training and fine-tuning is important.

ImageNet Experiments

In this section, we present and analyze two key proposals, the FCMAE pre-training framework and ConvNeXt V2 architecture, which are co-designed to make masked-based self-supervised pre-training successful. We show these designs synergize well and provide a strong foundation for scaling the model to various sizes. Additionally, we compare our approach to previous masked image modeling approaches through experiments. Furthermore, we show that our largest ConvNeXt V2 Huge model, which has been pre-trained using the FCMAE framework and fine-tuned on the ImageNet-22K dataset, can achieve a new state-of-the-art of 88.9% top-1 accuracy on the ImageNet-1K dataset, using only publicly available data.

In this paper, we conduct a unique study that involves co-designing both the self-supervised learning framework (FCMAE) and the model architecture improvement (GRN layer), through an empirical study of their learning behavior. The results presented in Table 3 demonstrate the importance of this approach.

We found that using the FCMAE framework without modifying the model architecture has a limited impact on representation learning quality. Similarly, the new GRN layer has a rather small effect on performance under the supervised setup. However, the combination of the two results in a significant improvement in fine-tuning performance. This supports the idea that both the model and learning framework should be considered together, particularly when it comes to self-supervised learning.

Model scaling.

In this study, we evaluated a range of 8 models with different sizes, from a low-capacity 3.7M Atto model to a high-capacity 650M Huge model. We pre-trained these models using the proposed FCMAE framework and compared the fine-tuning results to the fully supervised counterparts.

The results, shown in Figure 1, demonstrate strong model scaling behavior, with consistently improved performance over the supervised baseline across all model sizes. This is the first time the benefit of masked image modeling has been demonstrated in such a broad model spectrum, both in terms of effectiveness and efficiency. The complete tabulated results can be found in the appendix.

Comparisons with previous methods.

We compare our approach to previous masked auto-encoder methods , which were all designed for transformer-based models. The results are summarized in Table 4. Our framework outperforms the Swin transformer pre-trained with SimMIM across all model sizes. Compared to the plain ViT pre-trained with MAE , our approach performs similarly up to the Large model regime, despite using much fewer parameters (198M vs 307M). However, in the huge model regime, our approach slightly lagged behind. This might be because a huge ViT model can benefit more from self-supervised pre-training. As we will see next, the gap might be closed with additional intermediate fine-tuning.

ImageNet-22K intermediate fine-tuning.

We also present ImageNet-22K intermediate fine-tuning results . The training process involves three steps: 1) FCMAE pre-training, 2) ImageNet-22K fine-tuning, and 3) ImageNet-1K fine-tuning. We use 3842384^{2} resolution images for pre-training and fine-tuning . We compare our results to the state-of-the-art architecture designs, including convolution-based , transformer-based , and hybrid designs . All these results were trained with ImageNet-22K supervised labels. The results are summarized in Table 5. Our method, using a convolution-based architecture, sets a new state-of-the-art accuracy using publicly available data only (i.e. ImageNet-1K and ImageNet-22K).

Transfer Learning Experiments

We now benchmark the transfer learning performance. First, we evaluate the impact of our co-design, i.e. comparing ConvNeXt V1 + supervised vs. ConvNeXt V2 + FCMAE. We also directly compare our approach with Swin transformer models pre-trained with SimMIM . The training and testing details are provided in the appendix.

We fine-tune Mask R-CNN on the COCO dataset and report the detection mAPbox and the segmentation mAPmask on the COCO val2017 set. The results are shown in Table 6. We see a gradual improvement as our proposals are applied. From V1 to V2, the GRN layer is newly introduced and enhances performance. Upon this, the model further benefits from better initialization when moving from supervised to FCMAE-based self-supervised learning. The best performances are achieved when both are applied together. Additionally, our final proposal, ConvNeXt V2 pre-trained on FCMAE, outperforms the Swin transformer counterparts across all model sizes, with the largest gap achieved in the huge model regime.

Semantic segmentation on ADE20K.

To summarize, we conduct experiments on the ADE20K semantic segmentation task using the UperNet framework. Our results show a similar trend to the object detection experiments, and our final model significantly improves over the V1 supervised counterparts. It also performs on par with the Swin transformer in the base and large model regimes but outperforms Swin in the huge model regime.

Conclusion

In this paper, we introduce a new ConvNet model family called ConvNeXt V2 that covers a broader range of complexity. While the architecture has minimal changes, it is specifically designed to be more suitable for self-supervised learning. Using our fully convolutional masked autoencoder pre-training, we can significantly improve the performance of pure ConvNets across various downstream tasks, including ImageNet classification, COCO object detection, and ADE20K segmentation.

We thank Ross Wightman for the initial design of the small-compute ConvNeXt model variants and the associated training recipe. We also appreciate the helpful discussions and feedback provided by Kaiming He.

Appendix

This appendix provides implementation details, including model configurations, pre-training and fine-tuning recipes, and sparse and dense encoding methods for FCMAE pre-training (see §A). In §B, we present complete fine-tuning accuracy comparisons between ConvNeXt V1 and V2 on ImageNet 1K and 22K. In §C, we perform analyses on the efficiency of sparse encoding and general feature analysis using the class selectivity index. Finally, in §D, we conduct additional ablation studies on the masking ratio and GRN component analysis. We also compare FCMAE (masked image modeling) with MoCo V3 (contrastive learning).

Appendix A Implementation Details

The basic models, i.e., Tiny (28M), Base (89M) and Large (198M), follow the same configurations of the stage, block (B), and channel (C) settings of the ConvNeXt V1 .

ConvNeXt V2-B: CC=128, BB=(3, 3, 27, 3)

ConvNeXt V2-L: CC=192, BB=(3, 3, 27, 3)

Given the same definitions above, we scale the model to provide a broad model size spectrum, targeting versatile scenarios. First, to obtain efficient models, we scale down as follows:

A, F, P, N denote Atto (3.7M), Femto (5.2M), Pico (9.1M), and Nano (15.6M) models designed originally in . Next, to introduce the large-capacity variant, we scale up as follows:

ConvNeXt V2-H: CC=352, BB=(3, 3, 27, 3)

H denotes Huge (659M) model, which is newly presented in this work.

A.2 ImageNet Experiments

All models share the same pre-training setup, as noted in Table 8. We use the linear lr scaling rule : lr = base_lr×\timesbatchsize / 256.

ImageNet-1K fine-tuning

As the learning capacity varies by model size, we adopt different fine-tuning recipes for each model. We summarize them in Table 9, 10 and 11. We see longer fine-tuning epochs help small models. We adopt two different learning-rate layer decay strategies in this work: group-wise , where we treat three sequential layers as a single “layer” and use the same decaying value for them, and the layer-wise , where we assign a distinct value for each layer, both following the standard decaying rule. The default is a layer-wise strategy, but we apply the group-wise decaying strategy to Base and Large models.

ImageNet-22K intermediate fine-tuning

We conduct ImageNet-22K intermediate fine-tuning with the FCMAE-pretrained ConvNeXt models. We use nano, tiny, base, large, and huge models. The setups are summarized in Table 12 and 13. Similarly, using larger layer-wise learning rate decay values for small models is helpful.

Sparse encoding implementations.

We propose two possible implementations to enable FCMAE pre-training: 1) sparse encoding using sparse convolution supported by external libraries , and 2) simulating sparse encoding with the masked dense convolution, which can be easily implemented by applying binary masks before and after the standard convolution operation. As they produce numerically identical outputs, both can be adopted depending on different use cases. In this work, we adopt sparse encoding on the GPU environment, where we use MinkowskiEngine library and PyTorch framework ; we use dense masked conv based encoding on TPU accelerators using Jax . The experiments in the main paper are all conducted on TPU (v3-256) pods and we release a PyTorch reproduction.

A.3 Object detection and segmentation on COCO

For COCO experiments, we use the MMDetection toolbox and the final model weights from ImageNet-1K pre-training as network initializations. All models are trained with a 3x schedule (36 epochs) and a batch size of 32. We utilize an AdamW optimizer with a learning rate of 1e-4, a weight decay of 0.05 and sweep layer-wise learning rate decay in {0.9, 0.95}, stochastic depth rate in {0.2, 0.3, 0.4, 0.5}. We employ a large-scale jittering augmentation (1024×\times1024 resolution, scale range [0.1, 2.0]). We use single-scale testing with soft-NMS during inference.

A.4 Semantic segmentation in ADE20K

For ADE20K experiments, we use the MMSegmentation toolbox. We use an AdamW optimizer with the following hyperparameters: a weight decay of 0.05, a batch size of 16 and sweep layer-wise decay rate {0.8, 0.9}, learning rate {1e-4, 2e-4, 3e-4}, stochastic depth rate {0.1, 0.2, 0.3, 0.4}. All models are trained for 160K iterations with an input resolution of 512×\times512. In inference, a multi-scale test using resolutions that are [0.75,0.875,1.0,1.125,1.25] of 512×\times2048 is employed.

Similar to , we initialized the segmentation models using model weights after supervised fine-tuning on ImageNet-1K, as we found its performance superior to using the self-supervised pre-trained weights directly.

Appendix B Complete comparisons with V1

In Tables 14 and 15, we present detailed experiment-level comparisons between ConvNeXt V1 and V2. In particular, Table 14 shows ImageNet-1K fine-tuning results using eight models: Atto, Femto, Nano, Pico, Tiny, Base, Large, and Huge, which range from low-compute (Atto, 3.7M) to large-capacity models (Huge, 660M). We see a consistent and significant improvement across all models. The best performance is achieved when the architecture is upgraded from V1 to V2 and the self-supervised learning framework FCMAE is used, demonstrating the effectiveness of the co-design. In Table 15, we present ImageNet-22K intermediate fine-tuning results. The pre-training and fine-tuning process consists of three steps: 1) FCMAE pre-training, 2) ImageNet-22K fine-tuning, and 3) ImageNet-1K fine-tuning. Here, we focus on five V2 models: Nano, Tiny, Base, Large and Huge. We see consistent improvement over the V1 counterparts. In particular, the V2 Base (86.8%/87.7%) and Large (87.3%/88.2%) models outperform the next-level model sizes of V1, which are the Large (86.6%/87.5%) and XLarge (87.0%/87.8%) models. The V2 Huge model also achieves a new state-of-the-art with a performance of 88.9%. Our proposal demonstrates that pure convolutional models can also be strong and scalable vision learners with mask-based pre-training.

Appendix C Further Analyses

One of the key design choices in our FCMAE framework is the use of sparse convolution during pre-training. The primary purpose is to block the flow of information from the masked region and facilitate masked autoencoder pre-training. As a byproduct, it also offers improved computational and memory efficiency during pre-training, as the kernels only apply to the visible pixels. However, we note that the sparse convolution libraries are not highly optimized for modern hardware, and the efficiency achieved usually depends on the frameworks used in practice.

To better understand the actual pre-training efficiency achieved using sparse convolution, we conducted benchmark experiments using a controlled setup with Minkowski Engine v0.5.4 and PyTorch . We simulated the pre-training masked input (image size 224×\times224, masking ratio 0.6, mask size 32×\times32) and compared the training throughput (image/s) and max GPU memory usage (G) between the sparse convolution-based and dense masked convolution-based encoders. While the results may vary depending on the experimental environment (we used PyTorch V1.8.0, CUDA 11.1, CuDNN 8.2, and NVIDIA RTX A6000 GPU), we observed a moderate increase in pre-training efficiency, with an average of 1.3×\times increase in throughput and a 2×\times decrease in max memory usage across the models. The gap becomes more salient as the model size increases.

Class Selectivity Index.

FCMAE pre-trained ConvNeXt V2 has a distinctive feature characteristic compared to V1. We conducted a class selectivity index analysis on the FCMAE pre-trained weights for ConvNeXt V1 and V2 to understand this. The class selectivity index is a metric that measures the difference between the highest class-conditional mean activity and all other class-conditional mean activities. The final normalized value lies between 0 and 1, with 1 indicating that a filter activates only for a single class and 0 indicating that the filter activates uniformly for all classes. In Figure 7, we plot the class selectivity index distribution for all intermediate layers in the model, using the output of every residual block. The distribution is closely matched between V1 and V2 in the early stages, but they begin to diverge in the deep layers, such as stage 3 layer 12. As the layer becomes deeper, the plot shows that V2 (bimodal) tends to include more class-generic features than V1 (unimodal). Since class-agnostic features are more transferrable , this leads to better fine-tuning performance in downstream tasks. We leave more explorations as a future study.

Appendix D Additional Experiments

The proposed Global Relation Network (GRN) consists of three steps: global feature aggregation, feature normalization, and feature calibration. The main paper demonstrates that the combination of L2-norm based aggregation and divisive normalization works well in practice. Table 16 verifies the individual contribution of these components using ConvNeXt V2-Base as the encoder. When either component is dropped, performance significantly decreases, and the training becomes unstable if feature normalization is not preceded by global aggregation. This supports the idea that both operations work together to make GRN effective.

Masking ratios.

We conduct a hyper-parameter analysis on the masking ratio for a mask size of 32×3232\times 32. The results, shown in Figure 8, suggest that a masking ratio in the range of 0.5 to 0.7 produces the best results, with a masking ratio of 0.6 providing the highest performance. The model’s performance declines at the two extremes of either removing or leaving 90% of the input information, although it is more robust when more information is retained.

Comparison with contrastive SSL.

In this work, we compare the performance of the two dominant self-supervised learning (SSL) approaches: contrastive learning and masked image modeling . Specifically, we compare the end-to-end fine-tuning performance of MoCoV3 , the current state-of-the-art contrastive learning method, with our proposed FCMAE framework using the same ConvNeXt V2-Base as the encoder. We follow the default pre-training and fine-tuning recipes for each approach and present the results below.

We use the 300-epoch supervised learning baseline as a reference. The above table shows that FCMAE leads to better representation quality than MoCo V3 and also outperforms the supervised baseline. This is consistent with the recent observations that masked image modeling offers superior results over contrastive learning-based SSL for end-to-end fine-tuning. In this work, this success was also made possible with pure ConvNets.

References