Exploring Target Representations for Masked Autoencoders

Xingbin Liu, Jinghao Zhou, Tao Kong, Xianming Lin, Rongrong Ji

Introduction

Masked Image Modeling (MIM) has recently become an active research topic in the field of visual representation learning and establishes strong performance for vision recognition tasks, e.g., image classification, object detection, and semantic segmentation, which also surpasses traditional supervised learning mechanism. To be specific, MIM randomly masks a portion of the input and then reconstructs the masked portion according to the transformed target, formulated as

where “⊙\odot” means element-wise product; MM is the patch mask; “x⊙Mx\odot M” represents “unmasked patches” and vice versa; fθ(⋅)f_{\theta}(\cdot) is the learnable network to be pre-trained; T\mathcal{T} is the transformation function generating the reconstructed target. T\mathcal{T} can either be a parameterized network or a traditional image feature transformation method; M(⋅,⋅)\mathcal{M}(\cdot,\cdot) is the similarity measurement, e.g., l2l2-distance . A masked image passed through the network fθ(x⊙M)f_{\theta}(x\odot M) to reconstruct the visual representation of the intact image with transformation T(x⊙(1−M))\mathcal{T}(x\odot(1-M)).

A crucial problem of MIM is how to choose the reconstructed target, i.e., T(⋅)\mathcal{T}(\cdot) in Eq. 1. Previous methods use disparate teacher networks to generate the reconstruction target. BEiT employs a pre-trained DALL-E as the teacher network. In MaskFeat , authors use HOG , MoCo and DINO features to perform MIM; MVP employs a multi-modality model, CLIP , which is pre-trained by rich image-text pairs. MAE uses image pixels as the target, which functions likewise to a randomly initialized teacher network, as demonstrated in Sec. B.1. iBOT and data2vec use the exponential moving average (EMA) strategy to update teacher’s parameters ϕ\phi. Though different methods differ in their architectural designs and optimization, the choice of the teacher network lies crucial for each method and calls for a systematic study. In this work, we paraphrase a term Masked Knowledge Distillation (MKD) to focus our discussion on a special case of MIM where the target is generated by a parameterized network (teacher network), i.e., T(⋅)=hϕ(⋅)\mathcal{T}(\cdot)=h_{\phi}(\cdot). In this setting, T\mathbf{T} is the teacher network, and ff is the student network.

The purpose of our work is to investigate whether a careful design of the teacher network for MKD matters. Such exploration is nontrivial given that different teacher networks contain different knowledge we endued into the teacher network, which may induce diverse behaviors for the student networks. And the painstaking selection of the target representations in the field of MIM. To this end, we compare student networks distilled by four teacher networks with different computation pipelines, i.e., DINO for contrastive learning, MAE for masked autoencoding, DeiT for supervised learning, and DALL-E for autoregressive generation. Four teachers are all pre-trained on ImageNet-1K for a fair comparison. To our surprise, although the behaviors of the teacher networks are very different, the distilled student networks share similar characters after several stages of masked knowledge distillation: (i) the performance variance between student networks distilled from different teachers rapidly decreases. (ii) the model weights and output features across layers within the networks share similar properties.

Such observations indicate that the design of target representation is not essential for learning good visual representations when pre-trained with multi-stage, i.e., teacher networks do not matter with multi-stage masked knowledge distillation. Exceptionally, we use a randomly initialized model as teacher to perform multi-stage masked knowledge distillation, and find that it performs as well as those initialized by pre-trained models with the exact same settings! Using a random model as teachers not only avoids an extra pre-training stage, but also alleviates the painstaking selection of the target representations.

Based on the above studies and observations, we naturally propose to perform masked knowledge distillation with bootstrapped teachers, short as dBOT . Specifically, masked knowledge distillation is performed repeatedly in multiple stages. At the end of each stage, we assign the student’s weight to the teacher and re-initialize the student’s weight to continue masked knowledge distillation. With simple yet effective design that enables pre-training starting from randomly initialized teachers, dBOT achieves 84.5%, 86.6%, and 88.0% top-1 fine-tuning accuracy on ImageNet-1K with ViT-B/16, ViT-L/16, and ViT-H/14, respectively, significantly surpassing previous states of the art, MAE. Beyond that, dBOT achieves 52.7 and 56.0 APbox for object detection on COCO , as well as 49.5 and 54.5 mIoU for semantic segmentation on ADE20K , with ViT-B/16 and ViT-L/16 respectively. We also explore MKD with teachers of larger sizes, further boosting model performances on various visual tasks.

Related work

Self-supervised learning is an active research topic recently. Early practices revolve around contrastive learning where the model output features of images transformed by different data augmentations are pulled together. With the development of Masked Language Modeling (MLM) in language pre-training , researchers also introduce the training strategy of masked reconstruction to visual pre-training. BEiT uses the DALL-E to encode an image patch as the target for model reconstruction. iBOT uses an online teacher shifting the target from offline to online to make the target semantic meaningful. In addition to using the token obtained from offline or online model as reconstruct target, MAE , SimMIM , and MaskFeat achieve good performance in masked-image reconstruction using low-level pixels or HOG features. Among them, MAE uses an asymmetric encoder-decoder structure greatly increasing the training efficiency. data2vec demonstrates good generalizations on three modalities (vision, speech, and language) by reconstructing multiple neural network layer representations.

2 Knowledge Distillation

Knowledge distillation (KD) is widely employed in model knowledge compression , which improves the performance of the smaller student model by distilling the knowledge learned from a well-trained large teacher network. Further study on e.g. relational KD , contrastive KD , and latent feature KD is conducted to improve the performance of vanilla KD. Beyond its prominence in the field of supervised learning, KD recently cuts a figure in self-supervised learning. Concurrent work manages to adopt conventional feature distillation to match contrastive models with MIM-trained ones. Nevertheless, it shows negligible gains on MIM-trained models such as MAE. BEiT , MaskFeat and MVP could be seen as distilling knowledge from dVAE , HOG features and language-induced model CLIP within the discourse of MKD, respectively. Until now, there exists no work conferring a system-level study on the importance of how to choose adequate target representation or teacher networks to guide the learning of MKD. Beyond that, we propose using a randomly initialized model as the teacher and bootstraps the teacher for stages, demonstrating superiority over other practices.

Given the general form of masked knowledge distillation as shown in Eq. 1, in this section, we aim to investigate whether the careful design of the target, i.e., teacher network hϕ(⋅)h_{\phi}(\cdot), matters. Specifically, we want to answer three questions as follows:

Whether models distilled from different hϕ(⋅)h_{\phi}(\cdot) differ in terms of their transfer performances?

Whether distilled models differ in terms of their weights and outputs?

If hϕ(⋅)h_{\phi}(\cdot) does not matter, what matters more to close the gap between students distilled from different hϕ(⋅)h_{\phi}(\cdot)?

To answer these questions, we employ the standard masked autoencoder framework to give a system-level study, introduced next.

The architectural settings strictly follow . For the teacher network, we use the vanilla ViT with intact input. For the student network with masked input, we use the asymmetric encoder-decoder structure. The student’s output is further projected to a dimension the same as that of teacher’s embedding. During pre-training, we use Smooth L1 loss for the optimization of the student network, and the teacher network is kept fixed. Detailed settings are delayed to Sec. A.1. We pre-train models on ImageNet-1K and conduct evaluation under classification on ImageNet, object detection on COCO , and semantic segmentation on ADE20K .

1 Preliminary Study

We first investigate the effect of using networks initialized differently as teachers for masked knowledge distillation. Four canonical methods as pre-trained teachers are substantiated, each from a category distinguished based on their computation pipelines, i.e., DeiT for supervised learning, DINO for contrastive learning, DALL-E for autoregressive generation, and MAE for autoencoding. The results of initialized teacher at the 0th stage and of its distilled student at the 1st stage are shown in Table 1.

After the first stage of masked knowledge distillation, the student consistently outperforms teacher as shown in Table 1, yielding 1.8%, 1.0%, 2.4%, and 0.7% performance gains for four different hϕ(⋅)h_{\phi}(\cdot) respectively, demonstrating the effectiveness of masked knowledge distillation for visual representation learning. Although the performance order of different hϕ(⋅)h_{\phi}(\cdot) is reserved after the first stage of distillation, the students distilled from different hϕ(⋅)h_{\phi}(\cdot) have closer downstream performances compared to the original hϕ(⋅)h_{\phi}(\cdot). The performance variance drops from 2.24 to 0.37 after the first stage of distillation. Take MAE and DALL-E being the initialized teachers as an example, the performance range in classification, drops from 2.5% (83.6% vs. 81.1%) to 0.8% (84.3% vs. 83.5%), indicating that the performance gap is narrowed after the first stage. The conclusion holds true for experiments on object detection and semantic segmentation.

2 Distillation with Multiple Stages

Given the observations that better teacher generally induces better outperforming student, we are motivated to use the trained student as teacher to train new student repeatedly and study whether similar trend endures. If so, we would like to seek at what stage the performances saturate for different downstream tasks, as well as the discrepancy among the results incurred by different initialized teachers.

The performance gain is valid but decreases with multi-stage and eventually vanishes. Take MAE being the initialized teacher as an example, students outperform teachers by +0.7%, +0.1%, -0.1% for classification, +2.3, -0.2, -0.2 points for object detection, and +1.5, +0.8, -0.6 points, for semantic segmentation, from the 0th to the 3rd stage. Other teachers and downstream tasks share the same conclusion. Moreover, the performance gaps of students learned from different teachers decrease, especially after multi-stage, as shown by the performance variance at different stages in the last row of Table 1. Take classification tasks for instance, the variance decreases along with the training stage, i.e., 2.24, 0.37, 0.07, 0.04, which reveals that the choice of hϕ(⋅)h_{\phi}(\cdot) exerts little influence on the downstream performance. See Table 1 for results of more downstream tasks. To demonstrate models’ differences in terms of weights and outputs, we conduct a property analysis in Sec. 6. Similar properties are found, which verify our conclusion.

Since the choice of hϕ(⋅)h_{\phi}(\cdot) does not matter, an intuitive experiment is to see what will happen when we employ a random teacher, in which the parameters are randomly initialized at the 0th0^{th} stage. To our surprise, using a random teacher achieves performances comparably with other pre-trained teachers. Compared to a randomly initialized model, distilled students with multiple stages achieve 6.1%, 20.4, and 21.3 performance gain on classification, object detection and semantic segmentation respectively. Empirically, object detection and semantic segmentation require one more stage to saturate compared to classification. The saturated results are on par with those induced by pre-trained teachers, which enables us to train a state-of-the-art model more efficiently, without the need of an extra pre-training stage for the initialized teacher (e.g., contrastive learning as DINO).

MKD with Bootstrapped Teachers

The study in Sec. 3 motivates us to propose a multi-stage distillation pipeline for pre-training. The entire pre-training undergoes multiple stages split by breakpoints. For each stage, we fix teacher network to obtain a stable visual representation, guiding the learning of student network. The pre-trained student model is then used as a stronger teacher and distills its knowledge to a new subsequent student, providing richer visual representations. We re-initialize the student network at each breakpoint. The above process repeats itself - the teachers keep bootstrapped from the students, until a performance saturation on downstream tasks is observed. Hence, our strategy is to perform distillation with bootstrapped teachers. We illustrate our framework in Fig. 1(c) and the conceptual relations with the other two paradigms in Fig. 1. By noting mm as the momentum which indicates how fast the teacher’s parameters θt{\bm{\theta}}_{t} is updated from student’s parameters θs{\bm{\theta}}_{s}, i.e., θt=m⋅θt+(1−m)⋅θs{\bm{\theta}}_{t}=m\cdot{\bm{\theta}}_{t}+(1-m)\cdot{\bm{\theta}}_{s}, we present the following discussions.

One group of works leverages pre-trained teacher as in Fig. 1(a), i.e., BEiT . The teacher requires an extra stage of pre-training and is kept fixed with mm = 1. Ideally, pre-trained teachers bear additional knowledge which is prone to be more semantic meaningful, prompting student’s learning. Nonetheless, the pre-training of these teachers entails a completely different computation pipeline and often additional data , complicating its practical use. Another group as in Fig. 1(b) works with random teacher in dispense with pre-trained ones. Starting from randomness, the teachers in iBOT and data2vec , however, are bootstrapped from the student typically with m∈m\in (0, 1), e.g., 0.9998 as in . Although bootstrap induces improving quality of the teacher’s representation, the pipeline is plagued by its optimization instability and sensitivity towards hyper-parameters. We note that MAE uses identity mapping of pixels as the target, which is observed to function similarly as a fixed random teacher with mm = 1, as shown in Sec. B.1. Despite its simplicity, such practice eludes synergy between the teacher and the student. Comparatively, dBOT is with mm = 0 for every breakpoint and mm = 1 otherwise.

Experiments

We use different capacity Vision Transformers , i.e., ViT-B/16, ViT-L/16, and ViT-H/14 for dBOT. The input image of size 224×\times224 is first divided by a linear projection head into non-overlapping patch tokens total of 196 for ViT-B and ViT-L, and 256 for ViT-H. We exactly follow the common setup demonstrated in Sec. 3, e.g., a student with asymmetric encoder-decoder architecture, a teacher with intact input, etc.

Optimization.

The learning rate is first linearly increased to the initial learning rate for the first 40 epochs and then cosine annealed to 0. The initial learning rate is set as 1.5e-4 ×\times batch_size / 256, with batch size being 4096 for all models. We use the AdamW optimizer and Smooth L1 loss to optimize the parameters of student network. Stochastic drop rate are applied, 0.2 for ViT-B, 0.2 for ViT-L, and 0.3 for ViT-H. We use only center-crop and flipping for data augmentation. As shown in Table 1, the performance of different downstream tasks saturates at different stages. By default, we pre-train all models for classification with 2 stages, for object detection and semantic segmentation with 3 stages.

2 ImageNet Results

We primarily focus on the end-to-end fine-tuning performance and report the top-1 validation accuracy on ImageNet-1K dataset.

We sweep the base learning rate within a range with a batch size being 1024. We warm up the learning rate during the first 5 epochs to the initial learning rate and use a cosine schedule for the rest of the epochs. We average all the patch tokens output from the last transformer block and pass them into a linear projection head for classification. We fine-tune ViT-B for 100 epochs and ViT-L and ViT-H for 50 epochs in total.

Comparison with previous results.

We report the fine-tuning results on ImageNet-1K, mainly focusing on the comparison of the self-supervised and supervised methods. Supervised denotes the results reported in the MAE. As shown in Table 2, dBOT achieves remarkable results with different model capacities, demonstrating its scalability. We achieved top-1 evaluation accuracy of 84.5%, 86.6%, and 87.4% with ViT-B, ViT-L, and ViT-H, yielding gains of 0.9%, 0.7%, and 0.5% compared to MAE. When fine-tuned with an image size of 448, dBOT further achieves an accuracy of 88.0%, surpassing the results obtained by MAE.

Semi-supervised learning.

To investigate the label efficiency of dBOT, we also show the semi-supervised results on ImageNet-1K under different labeled data availability in Table 3. We focus on the comparison with self-supervised learning methods. The label-fraction sampling strategy follows . dBOT outperforms MAE by 1.7 and 1.4 points using 1% and 10% of the labels, respectively, showing a higher label efficiency.

3 Downstream Tasks

To further demonstrate the effectiveness, we consider dense prediction tasks: object detection, semantic segmentation, and instance segmentation, as well as classification tasks that transfer to smaller datasets.

We consider Cascade Mask R-CNN as the task head for object detection and instance segmentation with ViT-B and ViT-L on COCO . We report APbox and APmask for object detection and instance segmentation respectively. The results are demonstrated in Table 4. dBOT outperforms the previous self-supervised and supervised methods by a large margin, setting a new state-of-the-art result with both ViT-B and ViT-L. With ViT-B, dBOT achieves a APbox of 52.7 and a APmask of 45.7, outperforming the supervised baseline pre-training by 2.9 and 2.5 points, respectively. With ViT-L, such improvement is more prominent with 4.8 and 3.6 points respectively, showing the high scalability of dBOT for model capacity in downstream dense prediction tasks.

Semantic segmentation.

We adapt UperNet as the task head for semantic segmentation with ViT-B and ViT-L on ADE20K . We report the mIoU and mAcc for semantic segmentation, and the results are demonstrated in Table 5. We achieve the best performances on semantic segmentation compared to previous self-supervised methods by a nontrivial margin. dBOT improves mIoU from 47.4 to 49.5 with ViT-B, and 49.9 to 54.5 with ViT-L, yielding gains of 2.1 and 4.6 points respectively, compared to the supervised baseline. The improvement in semantic segmentation is as significant as in object detection.

Transfer learning.

To further investigate the generalizability of visual representations learned by dBOT. We study transfer learning performance by fine-tuning the pre-trained models on smaller datasets, including CIFAR10 (Cif10), CIFAR100 (Cif100), iNaturalist18 (iNa18), iNaturalist19 (iNa19), Flowers (Flwrs), and Cars . The results are shown in Table 6. dBOT achieves comparable, if not better, performances compared to previous best methods. Specifically, the improvement is significant on relatively larger datasets like iNaturalist18 and iNaturalist19, with 4.7% and 3.3% respectively compared to the supervised baseline.

4 Ablation Study

We study the influence of stage number by splitting total training epochs of 1600 into varying distillation stages, from 0 to 2. Results are shown in LABEL:tab:split_number. 2-stage distillation works the best (for classification task), achieving 84.5% accuracy. Splitting epochs to 3-stage brings 0.1% performance drop, while all splitting strategies obtain a top-1 accuracy higher than 83.6%, indicating its generalizability.

Epoch for each stage.

LABEL:tab:each_stage studies proper epochs needed for each stage in a 2-stage distillation pipeline. With the 2nd{}^{\textrm{nd}} stage distilling for 800 epochs, longer epochs for the 1st{}^{\textrm{st}} stage induces 0.2% improvement (84.3% vs. 84.5%). With the 1st{}^{\textrm{st}} stage distilling for 800 epochs, 800 epochs are enough for the 2nd{}^{\textrm{nd}} stage since 1200 epochs incur no gain. Evenly splitting the epochs in 2-stage masked knowledge distillation achieves the best performance.

Momentum update.

We use in dBOT a multi-stage distillation pipeline, which is to distill from a momentum encoder with mm being 0 for every breakpoint and 1 otherwise. We further investigate other momentum update strategies commonly used in self-supervised learning. Results are shown in LABEL:tab:momentum. The vanilla strategy works the best.

Target normalization.

We study whether patch tokens obtained by the self-attention blocks to be used as target representation should be passed through the Layer Normalization layer [LN]. The accuracy of models after 2-stage distillation is shown in LABEL:tab:reconstruct_target. Without passing through [LN], the patch tokens directly obtained from the transformer block make them less suitable as target representations to guide students’ learning.

Student initialization.

We study whether student’s weight should remain when entering the next stage of distillation. Specifically, we either keep the student’s weight unchanged or re-initialize the student at each breakpoint. As shown in LABEL:tab:student_weight, re-initializing the student’s weight works the best.

Mask ratio.

LABEL:tab:mask_ratio shows the influence of the mask ratio on end-to-end fine-tuning. The optimal mask ratio for dBOT is 75%, the same as that in MAE.

Property Analysis

We investigate the properties of models distilled from different teachers under certain criteria, analyzing models’ weights and outputs. Further, training efficiency is briefly discussed with previous methods.

We compute averaged attention distance , averaged over ImageNet-1K val set, for each attention head of different blocks to understand how local and global information flows into Transformers. Average attention distance for dBOT using DeiT, DINO, MAE, DALL-E, and random as teachers are illustrated in Fig. 2. The higher the attention distance, models’ attention over an image is more global. Although the average attention distance of disparate initialized teachers varies greatly, their distilled students after multi-stage distillation exhibit similar behaviors, e.g., models’ attention toward local or global contents. Additionally, dBOT achieves more local attention than previous works.

Singular value decomposition.

We computed the percentage of top-kk singular values of the embedding w.r.t each layer. The results are averaged over the ImageNet-1K val set. We showcase the results with kk varying from 1 to 5. Singular value decomposition for dBOT using DeiT, DINO, MAE, DALL-E, and random as teachers are shown in Fig. 3. The higher the percentage, the models’ output over an image is less correlated, indicating larger redundancy of its spatial representations thus less suitability for compression. Intuitively, random models at the 0th{}^{\textrm{th}} stage has the largest percentage given that pixel are merely randomly projected. The student networks distilled from different initialized teachers exhibit similar behaviors.

Unsupervised object detection.

We use unsupervised object localization to quantitatively evaluate the visual representation obtained by different models. We follow the evaluation practice proposed in with Correct Localization (CorLoc) on POC-VOC 2012 trainval sets, except that we conduct feature decomposition via SVD instead of Laplacian since we observe more stable behaviors with SVD. We first compute singular value decomposition for the patch feature obtained by the ViT-B last block. Then a sign operation is applied on the first eigenvector, obtaining a binary mask of an image. We then take the bounding box around the largest connected component, which is more like the foreground object instead of the background. Correct localization (CorLoc) is used to measure the results, evaluated on POC-VOC 2012 trainval sets. A box is considered to have correctly identified an object if it has more than 50% intersection-over-union with a ground truth bounding box. Quantitative results are demonstrated in Table 8. dBOT using different teachers achieves very similar results, with students consistently outperforming their teachers.

Training efficiency.

We compute the training time per epoch for different methods in Table 9. With an asymmetric encoder-decoder architecture (asym.) as the default setup, dBOT performs slower than MAE, but much faster than data2vec and BEiT. Such advantage turns more significant with models of larger size.

Distill from Bigger Teachers

Inspired by canonical practices in knowledge distillation , we use larger teachers to distill smaller students, showcasing the potential of MKD in general. Specifically, we attempt to use ViT-L/H as teacher networks to distill ViT-B, and ViT-H as the teacher network to distill ViT-L. All larger teachers are first distilled for 2 stages with the default setup. We resize the image to 196×\times196 for ViT-H/14 to keep the length of its output the same as that of ViT-B/L. While we do not find substantial gains on classification results, the results by distilling from ViT-H are significantly better for dense prediction tasks compared to the default setup, i.e., +0.8 points of APbox and +1.3 points of mIoU with ViT-B as the student. The performance gain in distilling ViT-L from ViT-H is diminished but still valid, i.e., +0.1 APbox and +0.7 mIoU. We also consider MKD with data-richer teachers, e.g. CLIP, as exploratory experiments and set new state-of-the-art results for self-supervised learning. Refer to Appendix C for details.

Conclusion

As a special case of MIM, we formulate MKD upon which an empirical investigation is conducted about the influence of different target representations on self-supervised masked autoencoders. The study concludes that it is not necessary to carefully choose the target representation to learn good visual representations if distillation is performed in multiple stages (i.e., with bootstrapped teachers). Instead of initializing teachers with pre-trained models, we resort to random ones for simple practice. Without an extra stage of pre-training, dBOT achieves favorable performance on image classification, object detection, and semantic segmentation. We hope our study and method will provide timely insights for self-supervised learning.

References

Appendix A Implementation Details

We show our default pre-training setup in the second colum of Table A1. We use Xavier Uniform to initialize the Vision Transformer . Note that we use asymmetry stochastic drop path rate for students and teachers.

Setup for distillation from bigger teachers.

We follow the default setup, except that we use a different setup for stages. We first train larger-size teachers for 2 stages (in all downstream tasks) and use those to distill new students for 1 stage (in all downstream tasks).

A.2 Classification

The default end-to-end fine-tuning recipe is shown in the second column of Table A2, following the common recipes of ViT tuning for self-supervised models. The same recipe is applied when distilling from bigger teachers.

A.3 Object Detection and Instance Segmentation

We adopt the vanilla ViT with Cascade Mask R-CNN as the task head on COCO dataset for object detection and instance segmentation, following the common setup . The default recipe is shown in Table A3. To cope with versatile image sizes, we add relative position embedding instead of interpolating the absolute position embedding obtained during pre-training. For a fair comparison, we applied the same setup and sweep the learning rate and stochastic drop path rate for different methods.

A.4 Semantic Segmentation

We use vanilla ViT and UperNet as the task head on ADE20K dataset for semantic segmentation, following the common setup . The default recipe is shown in Table A4. To cope with versatile image sizes, we add relative position embedding instead of interpolating the absolute position embedding obtained during pre-training. For a fair comparison, we applied the same setup and sweep the learning rate and layer-wise decay for different methods.

Appendix B Additional Experiments

MAE performs masked image modeling using the image pixel as the reconstruction target. We directly alter the target to patch tokens obtained from the image fed into a randomly initialized network. We select two patch tokens as the reconstruction target, one is the token obtained using the last transformer block, and the other is the token obtained using linear projection, i.e., without any transformer block. After 400 epoch pre-training of ViT-B, the top-1 accuracy of the model on ImageNet-1K obtained by the three different targets is shown below.

It can be derived that using the patch token obtained by a randomly initialized network as the target can achieve comparable results with a pixel as a target. A similar result proves that patch tokens obtained by a randomly initialized can also serve as a good reconstruction target.

B.2 Object Detection with Mask R-CNN

Additionally, we use Mask R-CNN structure with FPN for object detection and instance segmentation on COCO datasets. The results are shown in Table B5. dBOT outperforms other methods by a large margin, which is similar to the results using Cascade Mask R-CNN.

B.3 Linear Probing

We evaluate the linear probing performance of dBOT and MAE using ViT-B following the same setup as MAE, the results of which is shown below.

dBOT achieves comparable linear probing performances with MAE.

Appendix C Distill from Data-Richer Teachers

We explore to use models pre-trained with richer data (i.e., CLIP with 400M Image-Text pairs) as the initialized teacher to seek a potential upper-bound of MKD.

Compared to the default setup, there exist two major disparities of the pre-training recipes for models distilled from data-richer teachers, discussed next. The following practice is summarized as recipe detailed in Table A1.

Vanilla Architecture. We find that not using the asymmetric encoder-decoder architecture is optimal, as shown in Table C6. While an asymmetric architecture generates momentum for bootstrapping models similar to , which lies crucial for distillation with random teachers, it hurts the performance when distilling with stronger pre-trained teachers.

Hypothetically, the significance of the decoder in asymmetrical encoder-decoder architecture lies in the need for separate layers to decode low-level details when the targets contain little semantics (e.g., pixels and random mappings of pixels). Such a need is eased when the target contains high-level semantics (e.g., DINO and CLIP). The existence of the decoder, in this case, may even restrain the encoder to grasp full knowledge from the teacher, inducing degraded performances.

1-Stage MKD. We use different models as teachers to distill students for one stage with longer epochs, i.e., 1600. Results are shown in Table C7. Empirically, the performance gains for multi-stage MKD over 1-stage MKD decrease as teachers’ fine-tuning performance increases. Stronger teachers, such as DINO and MAE, induce similarly performed students with 1-stage MKD (1×\times1600) compared to 2-stage MKD (2×\times800).

Specifically, when using CLIP as the pre-trained teacher, the performance for 2-stage MKD is, to our surprise, 0.9% lower than that of 1-stage MKD. Understandably, although the fine-tuning result of the student after 1-stage distillation is better than that of CLIP, the student is essentially trained on IN1K and may not contain faithfully data information stored in the CLIP model. Therefore, strong teachers work well with 1-stage MKD, especially for models pre-trained on extra richer data.

C.2 Downstream Tasks

Implementation Details. For fine-tuning, we also use a slightly different recipe from default one with smaller learning rates and drop path, dubbed as recipe detailed in Table A2. For object detection, instance segmentation, and semantic segmentation, we follow the default setup detailed in Secs. A.3 and A.4.

Results. Results for downstream tasks are shown in Table C8. ViT-B distilled from CLIP-B achieves an 85.7% top-1 accuracy and a 52.9 mIoU, surpassing all previous arts. With CLIP-L as the teacher, ViT-H with image resolution 448448 achieves an 89.1% top-1 accuracy, setting a new state-of-the-art image recognition result.

C.3 Conflict with Main Conclusion

It can be observed that MKD with CLIP as the teacher performs much better than that with the random teacher and multi-stage distillation, which seems contradictory to our main conclusion that teacher networks do not matter with multi-stage masked knowledge distillation. Notably, CLIP is trained with 400M image text pairs (300×{\times} larger than ImageNet-1K), which is a drastically different setup from multi-stage distillation on ImageNet-1K only. Exploring CLIP as a target representation gains popularity recently but is beyond the main scope of this paper. We present these results to corroborate the validity and to explore the upper bound of MKD in general. We note that the exact solution to resolve the conflict is to perform multi-stage distillation using the CLIP’s in-house 400M data to which we have no access. It is hypothesized that two results should be matched in light of experiments on ImageNet-1K, which is left to future work.