Siamese Image Modeling for Self-Supervised Vision Representation Learning

Chenxin Tao, Xizhou Zhu, Weijie Su, Gao Huang, Bin Li, Jie Zhou, Yu Qiao, Xiaogang Wang, Jifeng Dai

Introduction

Self-supervised learning (SSL) has been pursued in the vision domain for a long time . It enables us to pre-train models without human-annotated labels, which makes it possible to exploit huge amounts of unlabeled data. SSL has provided competitive results against supervised pre-training baselines in various downstream tasks, including image classification , object detection and semantic segmentation .

To effectively train models in the SSL manner, researchers design the so-called “pretext tasks” to generate supervision signals. One of the most typical frameworks is Instance Discrimination (ID), whose core idea is to pull together representations of different augmented views from the same image, and avoid representational collapse. Different variants of ID have been proposed, including contrastive learning , asymmetric networks , and feature decorrelation . A recent work has shown the intrinsic consistency among these methods via their similar gradient structures. For ID methods, the representations of each image are well separated, thus inducing good linear separability. However, as shown in , for transfer learning on detection tasks with Vision Transformers , ID is not superior to supervised pre-training, and even lags behind random initialization given enough training time.

Recently, another SSL framework has gradually attracted more attention, namely Masked Image Modeling (MIM) . MIM methods train the model to reconstruct the original content from a masked image. Such practice can help to learn the rich local structures within an image, leading to excellent performance in dense prediction tasks such as object detection . Nevertheless, MIM does not have good linear separability as ID, and usually performs poorly under the few-shot classification settings .

Both ID and MIM methods have their own strengths and weaknesses. We argue that this dilemma is caused by neglecting the representation requirements of either semantic alignment or spatial sensitivity. Specifically, MIM operates within each image independently, regardless of the inter-image relationship. The representations of semantically similar images are not well aligned, which further results in poor linear probing and few-shot learning performances of MIM. On the other hand, ID only uses a global representation for the whole image, and thus fails to model the intra-image structure. The spatial sensitivity of features is therefore missing, and ID methods usually produce inferior results on dense prediction.

To overcome this dilemma, we observe the key factors for semantic alignment and spatial sensitivity: (1) semantic alignment requires that images with similar semantics are projected into nearby representations. This can be achieved by matching different augmented views from the same image. Strong augmentations are also beneficial because they provide more invariance to the model; (2) spatial sensitivity needs modeling the local structures within an image. Predicting dense representations from masked images thus helps, because it models the conditional distribution of image content within each image. These observations motivate us to predict the dense representations of an image from a masked view with different augmentations.

To this end, we propose Siamese Image Modeling (SiameseIM), which reconstructs the dense representations of an augmented view, based on another masked view from the same image but with different augmentations (see Fig. 1). It adopts a Siamese network with an online and a target branch. The online branch consists of an encoder that maps the first masked view into latent representations, and a decoder that reconstructs the representations of the second view according to the relative positions between these two views. The target branch only contains a momentum encoder that encodes the second view into the prediction target. The encoder is made up of a backbone and a projector. After the pre-training, we only use the online backbone for downstream tasks.

As shown in Tab. 1, SiameseIM is able to surpass both MIM and ID methods over a wide range of evaluation tasks, including full-data fine-tuning, few-shot learning and linear probing on ImageNet , object detection on COCO and LVIS , semantic segmentation on ADE20k , as well as several robustness benchmarks . By gathering semantic alignment and spatial sensitivity in one model, SiameseIM can deliver superior results for all tasks. We also note that such improvements are more obvious on ADE20k(∼\sim3 points) and LVIS(∼\sim1.6 point for rare classes) datasets. These long-tailed datasets demands semantic alignment and spatial sensitivity at the same time, and SiameseIM thus can deliver superior performance on them.

Our contributions can be summarized as follows:

As a new form of SSL, SiameseIM is proposed to explore the possibilities of self-supervised pre-training. It displays for the first time that that using only a single dense loss is enough to learn semantic alignment and spatial sensitivity well at the same time;

Compared with MIM methods, SiameseIM shows that reconstructing another view helps to obtain good semantic alignment. This also suggests that MIM framework can be used to reconstruct other targets with proper guidance, which opens a possible direction for MIM pretraining;

Compared with ID methods, SiameseIM shows that dense supervision can be applied by matching the dense correspondence between two views strictly through their relative positions. We demonstrate dense supervision can bring a considerable improvement of spatial sensitivity;

SiameseIM is able to surpass both MIM and ID methods over a wide range of tasks. SiameseIM obtains more improvements in few-shot, long-tail and robustness-concerned scenarios.

Related work

Instance Discrimination (ID). The core idea of instance discrimination is to pull together different augmented views of the same image and avoid representational collapse . In this way, models can learn to separate the representation of each image, leading to decent linear separability. There are three typical types of instance discrimination methods, while Siamese networks are always employed. Contrastive Learning methods push apart views from different images (negative samples) to avoid representational collapse. Asymmetric Network methods explore to get rid of negative samples with the help of an asymmetric network design. In these methods, a predictor network is only appended after one branch of the siamese network, and the other branch is detached from the gradient back-propagation. Feature Decorrelation methods try to accomplish instance discrimination by reducing the redundancy among different feature dimensions. These different methods are then unified in UniGrad by revealing that they share similar gradient structures. A recent work has even successfully surpassed the performance of supervised learning with ResNets . There have also been works on applying instance discrimination to Vision Transformers , which demonstrate impressive performances.

The most common evaluation metric used for ID is linear probing, which trains a linear classifier on top of frozen representations. This metric concentrates on the linear separability of learned features. However, as has pointed out, the dense prediction performance of ID on Vision Transformers is not superior to supervised pre-training, especially on object detection tasks.

Some previous ID works have tried to introduce dense supervision to enhance local features. They employ different techniques to build dense correspondence between two views, including using Earth Mover’s Distance , matching feature similarity , finding nearest neighbors , using extra region proposal or flow modules . However, these works either only focus on detection result, or rely on an extra global loss to improve linear probing result, which still requires to find a trade-off between semantic alignment and spatial sensitivity.

Our work display for the first time that only using a dense loss is enough to learn these two properties well at the same time. Unlike previous methods, we utilizes the relative positions between two views to strictly align the spatial correspondence. In doing so, SiameseIM outperforms the original ID methods by a large margin on a strong detection baseline.

Masked Image Modeling (MIM). Masked image modeling intends to reconstruct image content from a masked image, which is motivated by the masked language modeling in NLP . iGPT first tries to reconstruct image pixels. ViT has also tried to predict the mean color of the masked patch. However, these preliminary attempts are not competitive with their supervised counterparts. BEiT reveals the power of MIM by predicting visual tokens from a pre-trained discrete VAE . MAE successfully performs pre-training via predicting raw pixels. It shows that the key point is to use a high masked ratio due to the high spatial redundancy. After that, different works continue to push the limit by improving the quality of prediction targets. Some works have shown that it is more effective to predict features rather than raw pixels for learning representations.

Unlike ID methods, MIM methods excel in transfer learning with full model fine-tuning with Vision Transformers, but lack good linearly-separated representations . For example, given enough training epochs, BEiT and MAE can surpass other pre-training paradigms on detection tasks. However, under few-shot scenes, MIM methods are not as data-efficient as ID methods because of their poor linear separability .

Our work demonstrates that MIM can also produce the same adequate linear separable representations as ID. This is achieved by predicting the representation of another augmented view from the same image, rather than reconstructing the original view. Through reconstructing another view, SiameseIM can even surpass the linear probing performance of ID methods.

Method

We depict our model of SiameseIM in Fig. 2. It takes two augmented views xax_{a} and xbx_{b} of the same image as inputs. SiameseIM aims to predict the dense representations of xbx_{b} based on that of xax_{a}. A Siamese network with an online and a target branch is used. The online branch is made up of an encoder that encodes the visible patches of xax_{a} into a latent representation, and a decoder that predicts the representation of xbx_{b} according to the relative positions between xax_{a} and xbx_{b}. The target branch only has a momentum encoder which takes xbx_{b} as input. The encoder consists of a backbone and a projector. After the pre-training, only the online backbone is used for downstream evaluation.

ID methods adopt two different augmented views, while MIM methods utilize a single view. In our method, similar to ID methods, we feed two different views as the inputs to the online and target branches, respectively. As will be shown in Section 4.3, different views can significantly increase the linear probing result without harming the performance of object detection.

Apart from the number of views, there are also differences in the augmentations of previous methods (see Fig. 1). ID methods tends to add stronger augmentations, which typically contain spatial and color augmentations. Whereas recently, MAE reports that color augmentations are not beneficial for MIM pre-training. We find that color augmentations have different effects under different training settings. They can provide more invariance when used with different views, but such effect will vanish if paired with the same view (more analysis can be found in Section 4.3). As a result, we reserve both the spatial and color augmentations from ID methods .

Another difference is that MIM masks out some patches of the input image for reconstruction, which we refer as mask augmentation. With mask augmentation, the task of dense prediction can model the conditional distribution of image content within each image. The representations are trained to capture the local structure, and thus are endowed with spatial sensitivity. Therefore, we also apply mask augmentation to the view of the online branch.

2 Prediction Targets

There can be multiple choices for the prediction targets. For example, ID methods select to predict the features of different augmented views, while MIM methods are designed to predict pixels or features of the same view. We empirically find that feature prediction is superior for different views (see Section 4.3). SiameseIM is thus designed to predict the features of another different augmented view from the same image. We shall describe how the prediction and target are calculated.

Here, mm indicates the learnable embedding of the mask token, which follows . pap_{a} is the position embeddings for xax_{a}, and pb(u,v)p_{b}^{(u,v)} denotes the positional embedding for the patch of xbx_{b} at location (u,v)(u,v), which will be introduced later. NhN_{h} and NwN_{w} denote the number of tokens along height and width dimensions in xbx_{b} (e.g., Nh=Nw=14N_{h}=N_{w}=14), respectively. N=Nh×NwN=N_{h}\times N_{w} is therefore the number of all tokens in xbx_{b}. Note that different from MIM methods, the mask tokens correspond to image patches from the different target view xbx_{b} instead of the input view xax_{a}.

Positional Embedding for Online Decoder is necessarily required to inform the decoder of the corresponding locations of each patch in xax_{a} and xbx_{b}. The decoder predicts the dense representations of xbx_{b} based on visible patches from xax_{a} and their corresponding locations. For all input patches to the online decoder, including visible patches from xax_{a} and mask tokens indicating xbx_{b}, their positional embeddings are calculated from the relative position with respect to the left-top origin of xax_{a}. Fig. 3 shows the detailed process. Suppose the (left, top, height, width) positional properties of the two cropped views xax_{a} and xbx_{b} in the original image are (i1,j1,h1,w1)(i_{1},j_{1},h_{1},w_{1}) and (i2,j2,h2,w2)(i_{2},j_{2},h_{2},w_{2}), respectively. The positions for xax_{a} and mask tokens indicating xbx_{b} are

3 Loss Function

Loss functions guide the training direction, and thus shape the characteristic of the learned representations. ID methods usually adopt a loss over the globally averaged feature to separate representations among images, while MIM employs a dense loss on image patches to learn representations within each individual image. Interestingly, we find that a dense loss only is enough to train both semantic alignment and spatial sensitivity well. As a result, we only employ a dense loss for training SiameseIM.

Once the prediction yby_{b} and target zbz_{b} have been calculated, we adopt a dense loss function for each predicted token. UniGrad is employed because it is a unified loss of ID methods and is also memory-friendly. To apply UniGrad on the dense level, we treat each token representation as an independent sample, i.e., ybi,i=1,…,Ny_{b}^{i},i=1,\dots,N. The corresponding positive sample is therefore zbiz_{b}^{i}, and the negative samples consist of all the tokens from the target branch. The dense loss can then be computed according toHere we remove the L2 normalization in the original UniGrad formulation, as we empirically find that not applying normalization to the online prediction can improve the performance (see Section 4.3).

where ybiy_{b}^{i} comes from the online prediction, its target zbiz_{b}^{i} is the positive sample, and all token representations from the target branch constitute the negative sample set N\mathcal{N}. Note that most ID methods use the InfoNCE loss , which will require O(∣N∣)\mathcal{O}(|\mathcal{N}|) memory to calculate the similarities. This is infeasible for the dense loss because of the vast number of negative sample patches. In contrast, UniGrad only consumes O(D2)\mathcal{O}(D^{2}) memory by first calculating the covariance matrix of negative samples.

4 Discussion

As illustrated in Fig. 1, we would like to further emphasize the differences between SiameseIM and previous works from three perspectives:

(1) Compared with MIM methods, SiameseIM reveals that it’s possible to reconstruct another augmented view rather than the same view. Such reconstruction can greatly enhance the semantic alignment of the model. We also show that strong augmentations can benefit this learning process;

(2) Compared with ID methods, SiameseIM shows that dense supervision can greatly improve spatial sensitivity, and the model can also learn semantic alignment well with a dense loss. By employing the relative positions between two views, SiameseIM is able to achieve strict spatial alignment and build dense correspondence without anbiguity. We also demonstrate a considerable boost on a strong detection baseline;

(3) For combining the best of MIM and ID methods, recent works have also made some attempts . Nevertheless, their efforts do not jump out of the default setting of both frameworks, i.e., they use different views only for ID, and apply MIM only on each view independently. These methods rely on both the global and dense loss as training objectives. In comparison, our work can naturally take the best of ID and MIM. SiameseIM enforces the similarity between different views from the dense level. Benefiting from our modeling, we can reveal the most important factors that influence the linear probing and object detection performances by gradually modifying ID or MIM methods to SiameseIM (see Section 4.3).

Experiments

ViT-B/16 is used as the backbone. Transformer encoder blocks with BatchNorm are adopted as the projector and decoder. Before calculating loss, if not specified, we follow MAE to apply LayerNorm without affine parameters to target, and no normalization to prediction. We set λ=0.02\lambda=0.02 in the loss function. During pre-training, we adopt the standard augmentation used in MoCo-v3 . For masking strategy, if not specified, we follow BEiT to use blockwise masking. We evaluate our model in various downstream tasks, including full data fine-tuning, few-shot learning and linear probing on ImageNet , object detection on COCO and LVIS , semantic segmentation on ADE20k , as well as several robustness benchmarks . Please refer to Appendix A for more implementation details.

2 Main Results

Image Classification. Tab. 2(a) shows the results of image classification tasks on ImageNet . For full data finetuning, SiameseIM surpasses both MIM and ID methods. For linear probing, our work outperforms MAE by 10 points, and MoCo-v3 by 1.3 poins. When only 1% data is available, SiameseIM can outperform MoCo-v3 by 1.7 points and MAE by 14.0 points. This validates that SiameseIM has obtained good semantic alignment. Moreover, SiameseIM can already deliver comparable results with previous works with only 400 epochs’ pretraining.

Common Object Detection. Tab. 2(b) reports the performance on COCO detection. Compared to MoCo-v3, our method obtains 4.2 points improvement. SiameseIM is also better than MIM methods . This comparison validates that our method gets good spatial sensitivity.

Semantic Segmentation. Tab. 2(b) also demonstrates segmentation results on ADE20k . SiameseIM can surpass all pure ID and MIM methods over 3.0 points. Moreover, our method obtains 49.6 mIoU with only 400 epochs’ pretraining, already on par with previous methods. Different from data-balanced ImageNet and COCO datasets, ADE20k contains classes that do not have enough labels. It thus demands both semantic alignment and spatial sensitivity of high quality. This shows the superiority of our method.

Long-tail Object Detection. Tab. 2(c) compares the results on LVIS object detection. SiameseIM performs on par with MAE on overall AP metric, and delivers 1.6 point gain on rare classes. Different from common object detection, long-tail object detection poses higher demand for semantic alignment because of rare classes. SiameseIM therefore displays larger improvement.

Robustness Evaluation. Tab. 2(d) shows the robustness evaluation on four datasets . Compared with MoCo-v3 , our method can bring an average of 4.5 gain. Compared with MAE , SiameseIM can leads by a large margin of an average of 6.1 points. The results suggest that SiameseIM helps to improve the robustness of representations both over ID and MIM methods.

3 Ablation Study

We carefully ablate the components of SiameseIM in this section to identify the most important factors for semantic alignment and spatial sensitivity. We focus on the linear probing and COCO detection results. We state the key observations as follows.

Predicting Pixels or Features. Tab. 3(ab) and (de) ablate what type of target to use. When the same view is used for input and target, we find that predicting raw pixels performs better than predicting features. On the contrary, it’s superior to predict features if different views are used. We suspect that, for different views, predicting pixels presents a much more difficult pretext task than using the same view, whereas predicting features simplifies this reconstruction because the network can help to filter irrelevant details and extract semantic information.

Different Views. We demonstrate the effectiveness of different views by comparing Tab. 3(af). Here, we choose the best-performed setting for both the same view or different views. It’s shown that different views significantly improve the linear probing performance by ∼\sim11 points. This justifies our claim that matching different augmented views is the key to obtain semantic alignment.

Color Augmentations. Tab. 3(ac) and (ef) reports the effects of color augmentations. We observe different effects with the same view or different views. Color augmentations can help linear probing to obtain 3.5 points gain for different views, but this improvement vanishes with the same view. This coincides with the phenomena in SimCLR and MAE . We presume that if the same view is adopted, the color augmentations used for the target will be leaked to the model, which spoils the color variation.

BN/LN for Projector and Decoder. We also study how different normalizations will influence the model. The commonly used normalization in Transformer blocks is LayerNorm (LN) , while BatchNorm (BN) proves to be important in ID methods . We therefore try to replace LNs with BNs. Note that to preserve the vanilla ViT backbone, we only conduct this replacement for the projector and decoder. Tab. 3(fg) displays that BN gives slightly better results on both linear probing and dense prediction.

Mask Type. Tab. 3(gh) compares two mask strategies. It shows that blockwise mask is beneficial for both downstream tasks. We think that it should be easier to reconstruct feature by just interpolating local features. Blockwise mask masks out continuous patches, which makes this hard and forces the model to capture long-range dependency.

Loss Normalization. Tab. 3(hi) ablates normalization in the loss function. MoCo-like normalization applies BN without affine parameters follow by l2 normalization to both target and prediction. MAE-like normalization only applies LN without affine parameters to the target. We find that MAE-like normalization performs better on linear probing and comparable on object detection campared with MoCo-like normalization, thus we adopt MAE-like loss normalization in our default setting.

Dense Supervision. Finally we study the role that the dense supervision plays in Tab. 3(gj). By adopting the dense loss, SiameseIM is able to get an improvement of 2.8 points on object detection and 2.3 points on instance segmentation. This comparison validates our observation that modeling dense representations from a masked image is beneficial for dense prediction tasks. We note that only using dense loss can also help to improve linear probing.

Conclusion

Different self-supervised learning (SSL) frameworks have their own strengths: Instance Discrimination (ID) possesses good semantic alignment, and Masked Image Modeling (MIM) has decent spatial sensitivity. In this study, we propose a new SSL framework, namely Siamese Image Modeling (SiameseIM), to show that it’s possible to obtain two properties at the same time using only a single dense loss. We observe that (1) semantic alignment can be learned by matching different augmented views; (2) spatial sensitivity can be obtained by modeling dense representations from masked images. As a result, we propose Siamese Image Modeling (SiameseIM), which predicts the dense representations of an augmented view, based on another masked augmented view from the same image. SiameseIM is able to outperform ID and MIM methods over a wide range of downstream tasks. We hope that SiameseIM can bring some insights and inspirations for self-supervised pre-training, and open new possibilities in this domain.

Limitations. The training of SiameseIM is less efficient compared to that of MAE. Using fewer tokens or smaller resolutions for the target branch may reduce the computation burden. Because this paper focuses on exploring the possibility of self-supervised pretraining, i.e., combining linear separability and spatial sensitivity within a single loss, and revealing the connection between ID and MIM methods, we expect to propose a more efficient way to perform SiameseIM pretraining in future work.

Potential negative societal impacts. Our method has similar problems of the SSL paradigm. It requires huge computational resources to conduct large scale pretraining, which may consume a lot of electricity. Furthermore, it may possess biases in its digested data, and therefore should be used with caution.

The work is partially supported by the National Natural Science Foundation of China under Grants No.U19B2044, No.61836011 and No.62022048.

References

Appendix A Implementation Details

For pre-training, we mainly follow the setting of MAE . Detailed hyper-parameters for pre-training are listed in Tab. 4.

Augmentation We use the strong augmentations from MoCo-v3 , including random resized cropping, horizontal flipping, color jittering, grayscale conversion, Gaussian blurring and solarization. For masking strategy, we use follow BEiT to use blockwise masking with a masking ratio of 60%.

Architecture. We use the standard ViT-B/16 as the backbone for both online and target branches. We stack 22 and 44 Transformer encoder blocks with BatchNorm as the projector and the decoder, respectively. Both the projector and decoder have 768768 embedding dimension and 1212 heads for each block. The EMA coefficient for target momentum encoder is initialized as 0.9950.995 and is applied with a cosine schedule from 0.9950.995 to 1.01.0. Before calculating loss, we follow MAE to apply LayerNorm without affine parameters to target, and no normalization to prediction.

A.2 Finetuning with 100% Data

We follow the finetuning setting of MAE except that we search for the optimal learning rate. Other hyper-parameters are listed in Tab. 5.

A.3 Linear Probing

We follow the linear probing setting of MAE while always searching for the optimal learning rate. Specifically, an extra BatchNorm layer without affine transformation is added before the final linear classifier. Other hyper-parameters are listed in Tab. 6.

A.4 Finetuning with 1% Data

For few-shot evaluation, we follow the practice in . Specifically, we freeze the backbone and extract representations for each image. Then the cyanure package is used to apply \l2\l_{2}-regularized logistic regression on the representations. Note that for MAE, we report partial finetuning result because it is better than just training a linear classifier .

A.5 COCO Detection

We follow to evaluate on COCO . We adjust the learning rate schedule so as to drop the learning rate once the performance saturates. Hyper-parameters are listed in Tab. 7.

A.6 Semantic Segmentation

We follow to use UperNet as the segmentation network. We use the open-source code from mmsegmentation and only change pretrained backbone. Hyper-parameters are listed in Tab. 8.

A.7 LVIS Detection

We follow to evaluate on LVIS . We adjust the learning rate schedule so as to drop the learning rate once the performance saturates. Hyper-parameters are listed in Tab. 9.

A.8 Robustness benchmarks

We finetune the model on original ImageNet using the setting in Tab. 5, and test it on different validation sets without further finetuning.

Appendix B Attribution of Assets

ImageNet is subject to the ImageNet terms of access . COCO 2017 is publicly available under the Creative Commons Attribution 4.0 License. As far as we know, they do not contain any personally identifiable information or offensive content.