ConvMAE: Masked Convolution Meets Masked Autoencoders

Peng Gao, Teli Ma, Hongsheng Li, Ziyi Lin, Jifeng Dai, Yu Qiao

Introduction

Self-supervised learning frameworks, such as DINO , MOCO-V3 , MAE , unleash the potential of Vision Transformers (ViT) and achieve high performance on various downstream vision tasks . Among them, Mask Autoencoders (MAE) demonstrate superior learning ability and scalability. Motivated by BERT in natural language processing, MAE utilizes an asymmetric encoder and decoder architecture, in which masked tokens of the encoder are reconstructed by the decoder. Experiments show that MAE can learn discriminative and scalable representations from ImageNet-1K without relying on large-scale datasets, such as ImageNet-22K.

Local inductive bias and hierarchical representations are explored for boosting the performance of ViT. The combination of local convolution and global transformer operations leads to clear improvements on image classification , object detection , and semantic segmentation . In contrast to MAE , well-performing multi-scale backbones built upon local and global operations are mainly trained in supervised manner. A natural question is whether multi-scale backbone with local and global operations, which show promising performance on supervised learning can be exploited to enhance the masked auto-encoding paradigm .

In this paper, a simple and effective self-supervised learning framework, dubbed as ConvMAE, is proposed to train scalable representations by introducing hybrid convolution-transformer architectures and masked convolution into the masked auto-encoders. Although the modifications to the original MAE are minimal, ConvMAE shows great success on pretraining visual representations for boosting the performances of various tasks.

Different from MAE , the encoder of ConvMAE progressively abstracts the input image into multi-scale token embedding, while the decoder reconstructs the pixels corresponding to masked tokens. For high-resolution token embedding at early stages, convolutions blocks are adopted to encode local content. For low-resolution token embedding at late stage, transformer blocks are used to aggregate global context. The encoder therefore obtains both local and global FOV at different stages and generates discriminative multi-scale features. Note that the ConvMAE encoder is partly motivated by the strong hybrid convolution and transformer backbones, including Co-AtNet , Early Convolution , Container and Uniformer . However, previous hybrid convolution-transformer networks were either not explored for masked auto-encoding or show very similar performance to MAE . Instead of designing novel architectures, we focus on making basic hybrid convolution-transformer architectures work for mask auto-encoding and conduct extensive experiments to demonstrate its effectiveness on various downstream tasks.

The efficient and effective training of ConvMAE is enabled by a block-wise masking strategy with masked convolution . The masking strategy adopted in current mask-autoencoding frameworks, such as BEiT , MAE , SimMIM , cannot be naively used for ConvMAE as all tokens need to be kept in the later transformer stages. This leads to unaffordable computation cost for pretraining large and huge models, losing MAE’s efficiency advantage of omitting masked tokens in transformer encoder. In addition, directly pretraining with the convolution-transformer encoder causes pretraing-finetuning discrepancy as only visible tokens are processed during finetuning stages.

To tackle the issues, we focus on designing hybrid convolution-transformer architectures suitable for mask auto-encoding. Specifically, our ConvMAE adopts a block-wise masking strategy to first obtain a mask for the late stage in transformer and then progressively upsamples the mask to larger resolutions in early convolutional stages. In this way, tokens processed by late stages can be completely separated into masked tokens and visible tokens and inherit the computation efficiency of MAE. To prevent information leakage, the convolution blocks at early stages are equipped with masked convolutions, which avoid mixing up features of masked and visible regions in late stages to ensue the training effectiveness. Masked convolution has been well explored in sparse feature extraction and image inpainting . It can be naturally integrated into the hybrid convolution-transformer architecture to enable masked auto-encoding.

Our ConvMAE can naturally provide multi-scale features for object detection and semantic segmentation, which are required by modern detection and segmentation frameworks . Multi-scale features from the pretrained ConvMAE can significantly improve the performances of object detection and semantic segmentation compared with MAE. ConvMAE with masked-based autoencoding can even surpass the fully-supervised pretraining of Swin and MViT .

In summary, our contributions can be summarized below: (1) We present the strong and efficient self-supervised framework ConvMAE, which is easy to implement but show outstanding performances on different tasks. (2) The proposed ConvMAE naturally generates hierarchical representations and exhibit promising performances on object detection. (3) ConvMAE-Base improves the ImageNet finetuning accuracy by 1.4% compared with MAE-Base. On COCO 2017 with Mask-RCNN, ConvMAE-Base achieves 53.2% APboxAP^{\rm box} and 47.1% APmaskAP^{\rm mask} with a 25-epoch training schedule while MAE-Base attains 50.3% APboxAP^{\rm box} and 44.9% APmaskAP^{\rm mask} with 100 training epochs. On ADE20K with UperNet, ConvMAE-Base surpasses MAE-Base by 3.6 mIoU (48.1% vs. 51.7%).

Approach

Masked Autoencoders (MAE) is a self-supervised method for pretraining ViT by reconstructing masked RGB patches from visible patches. Although MAE has a simple design, it has been proven to be a strong and scalable pretraining framework for learning visual presentations. MAE consists of transformer-based encoder and decoder, where only visible patches are fed into the encoder and learnable mask tokens are processed by the decoder for image reconstruction to learn visual representations. As the encoder only needs to process a small portion of visible tokens, it alleviates the scalability problem to pretrain large vision models.

2 ConvMAE

ConvMAE is a simple and effective derivative of the popular MAE with minimal but effective modifications on the encoder design and the masking strategy. The goal of ConvMAE is to learn discriminative multi-scale visual representations and to prevent pretraining-finetuning discrepancy when applies MAE on convolution-transformer networks.

Directly applying the original masking strategy on the feature maps of the convolution-transformer encoder would make transformer layers keeping all tokens during the pretraining, jeopardizing the training efficiency. We introduce a hierarchical masking strategy coupled with masked convolution for the convolution stages to ensure only a small number of visible tokens are input into the transformer layers. The overall pipeline of ConvMAE is shown in Figure 1.

Block-wise Masking with Masked Convolutions. Mask auto-encoders, such as MAE and BEiT , adopt a random mask on the input tokens. However, the same strategy cannot be directly applied to our ConvMAE encoder. Uniformly masking stage-1 input tokens from the H4×W4\frac{H}{4}\times\frac{W}{4} feature maps would cause all tokens of stage-3 to have partially visible information and requires keeping all stage-3 tokens. Therefore, we propose to first generate the random mask to mask out p%p\% (e.g., 75%) of stage-3 input tokens and upsample the mask by 2 times and 4 times to obtain the corresponding block-wise masks for masking stage-2 and stage-1 inputs, respectively. The corresponding masked tokens in the three stages are dropped in the encoding process and are reconstructed by the decoder for feature learning. In this way, ConvMAE only needs to keep as few as 25% tokens in the time-consuming transformer blocks for training and the efficiency of ConvMAE is not compromised.

However, the 5×55\times 5 depthwise convolutions in the first two stages naturally lead to receptive fields larger than the masked patches and cause information leakage when reconstructing masked tokens. To avoid such information leakage and ensure the quality of pretraining, we adopt masked convolution in the first two stages, so that the masked regions would never be involved in the encoding process. The use of masked convolution is crucial to the superior performance of ConvMAE and the pretraining-testing discrepancy is prevented by removing partially masked tokens from stage.

The Multi-scale Decoder and Loss. The decoder of the original MAE takes as input both visible tokens EdE_{d} from the encoder and the mask tokens [Mask], and transform them in stacked transformer blocks for image reconstruction. Our ConvMAE encoder obtains multi-scale features E1E_{1}, E2E_{2}, E3E_{3}, captures both fine- and coarse-grained image information. To better supervise the pretraining of such multi-grained representations, we downsample E1E_{1} and E2E_{2} to the same size of E3E_{3} with stride-4 and stride-2 convolutions and fuse multi-grained tokens via a linear layer to obtain visible tokens EdE_{d} ,

where StrideConv(⋅,k){\rm StrideConv}(\cdot,k) represents stride-kk convolution. The multi-scale decoder is illustrated in the bottom-left part of Figure 1. The same losses from MAE are used for reconstructing masked image patches and only the reconstruction of masked patches are considered in the objective function.

3 ConvMAE for Object Detection and Semantic Segmentation

After pretraining, the proposed ConvMAE can naturally generate multi-scale feature maps, which can be processed by existing object detection and semantic segmentation heads.

As shown in Figure 2, to finetune ConvMAE for object detection, an E4E_{4} feature map of 1/321/32 input resolution is first obtained by 2×22\times 2 max pooling E3E_{3}. However, as the ConvMAE stage-3 has 11 global self-attention layers (in our ConvMAE-base model) with excessive computational cost, we follow Benchmarking ViT to replace all but 1st, 4th, 7th, 11th global self-attention layers in stage-3 to shifted-window local self-attention layers with alternatively shifted 7×77\times 7 windows. The modified local self-attention layers are still initialized by the pretrained global self-attention layers. A global relative position bias is shared between global transformer blocks. Similarly, a local relative position bias is shared by local transformer blocks. In this way, the heavy computational and GPU memory costs of the stage-3 are much mitigated. The multi-scale features E1,E2,E3,E4E_{1},E_{2},E_{3},E_{4} are then fed into the MaskRCNN head for object detection. To finetune ConvMAE for semantic segmentation, its stage-3 architecture is kept as the images in segmentation datasets have relatively smaller resolutions. The multi-scale features are feed into UperNet .

4 ConvMAE for Video Understanding

Attention based models have demonstrated superior performance on video understanding. Our ConvMAE can also be extended to serve as a strong video pretraining framework, dubbbed as VideoConvMAE, with simple modifications. Specifically, VideoConvMAE replaces image patch embedding with cube embedding, after which stage 1 and stage 2 perform local spatial-temporal feature fusion with masked 3D convolutions. Stage 3 still adopts stacked transformer blocks for spatial-temporal fusion. The spatial position embedding is extended to spatial-temporal embedding. Similar to the multi-scale decoder proposed in Section 2.2, outputs from stages 1, 2 and 3 are fused before feeding into a spatial-temporal transformer decoder for masked pixel reconstruction. Details about VideoConvMAE pretraining are in appendix B. Note that unlike previous approaches, which initialize models pretrained on images , our VideoConvMAE is pretrained from scratch on pure video datasets.

Experiments

To validate our proposed ConvMAE, we conduct experiments of image classification on ImageNet-1K dataset. The pretrained ConvMAE is also extensively tested on object detection and semantic segmentation. By default, we report performance of our the ConvMAE-base model with multi-scale decoder, which has similar parameters and FLOPs as the MAE-base.

Experimental Setup. ImageNet-1K consists of 1.3M images of 1k categories for image classification and is split to the training and validation sets. We pretrain our ConvMAE on ImageNet-1K training set. By default, we fix the mask ratio to 25% following the original MAE . The decoder is designed to have 8 transformer layers with 512 feature dimensions and 12 attention heads. We adopt a 1600-epoch cosine learning rate schedule with the first 40 epochs for warming up. The AdamW optimizer is utilized with a base learning rate of 1.5×10−41.5\times 10^{-4}, a weight decay of 0.05 and a batch size of 1024. Random cropping is employed as data augmentation during pretraining. After pretraining, the ConvMAE encoder is used for supervised finetuning on ImageNet-1K training set for 100 epochs using the cosine learning rate schedule. We follow the default finetuning parameters of the original MAE except for the layer-wise learning-rate decay parameters (0.65, 0.75, 0.85). For finetuning, we report the classification accuracy on the ImageNet validation set of the finetuned and pretrained (linear probe) ConvMAE encoders.

cccccc|cc P-Epochs &Masked Block 5×55\times 5 7×77\times 7 9×99\times 9 FT (%) FLOPs Conv Masking Conv Conv Conv 800 ✓ ✓ ✓ ✗ ✗ 84.6 1×1\times ✓ ✗ ✓ ✗ ✗ 84.2 1.7×1.7\times ✗ ✓ ✓ ✗ ✗ 81.5 1×1\times ✓ ✓ ✓ ✗ ✗ 84.5 0.997×0.997\times ✓ ✓ ✗ ✓ ✗ 84.4 1.003×1.003\times ✓ ✓ ✗ ✗ ✓ 84.6 1.007×1.007\times

cc|cc|cc|c P-Epochs Method FT (%) LIN (%) APboxAP^{box} APmaskAP^{mask} mIoU 200 ConvMAE-Base 84.1 N/A 50.2 44.8 48.1

w/ multi-scale decoder 84.4 N/A 50.8 45.4 48.5 1600 ConvMAE-Base 84.6 69.4 52.5 46.5 50.7 w/ multi-scale decoder 85.0 70.9 53.2 47.1 51.7

Convergence speed. We compare the convergence of ConvMAE and MAE in terms of ImageNet-1K finetuning, linear probing accuracy and COCO APboxAP^{box} in Figure 4. For fair comparison, ConvMAE and MAE are both pretrained for 1600 epochs. ConvMAE not only attains strong final results but also significantly increases convergence speed on various tasks. Specifically, ConvMAE can surpass the final performance of MAE at 58 epochs on ImageNet-1K finetuning. On COCO object detection, ConvMAE surpasses MAE at 16 epochs, indicating 6.6×6.6\times faster convergence speed.

Vision Transformer. Vision Transformer(ViT) achieved state-of-the-art results on various vision tasks. To increase the convergence speed and improve accuracy, well-explored locality inductive bias have been reintroduced into vision transformer , among which, hybrid architecture of convolution and transformer design can achieve state-of-the-art performance of a wide range of tasks. Our ConvMAE is highly motivated by the hybrid architecture design in vision backbones. Instead of designing new architectures, ConvMAE aim to unleash the powerful representation induced by hybrid architectures through MAE-style pretraining with several insightful modifications.

Self-supervised Representation Learning. Contrastive learning learn invariances by comparing augmented views of un-labeled images. Recently Mask-Autoencoding motivated by BERT raised to be a promising methodology. Mask-Autoencoding can learn strong representation through masked patch reconstruction with simple data augmentation. BEiT introduced Mask-Autoencoding into Vision Community. MAE introduced an asymmetric encoder and decoder architecture where masked tokens is skipped in computation-heavy encoder and only pass all tokens through a light-weight decoder. iBoT and Data2Vec , PeCo and MaskFeat explore different reconstruction targets. Different from previous improvements of Mask-autoencoding, ConvMAE introduce hierarchical representations architectures into MAE.

We propose a simple self-supervised learning framework named as ConvMAE which demonstrate the hybrid local-global blocks can boost the performance of MAE to generate discriminative multi-scale features . The computational efficiency and low pretraing-fineuning gap of original MAE can be well maintained under our ConvMAE. ConvMAE exhibits significantly improved performances on various vision tasks and can be easily implemented. We will study combining improved reconstruction targets with ConvMAE in the future.

Negative societal impact: We do not foresee nagative social impact from the proposed work.

Appendix

Model Scaling up and down. We design ConvMAE of different parameters scales to match those of MAE-small, MAE-base, MAE-large and MAE-huge. Detailed network architectures are in appendix. The finetuning performances are shown in Table 6. Compared with the original MAE of different scales, our ConvMAE of different scales consistently outperform its MAE counterparts on Imagenet finetuning. This suggests that ConvMAE can be an efficient learner for different paramter scales.