Deconstructing Denoising Diffusion Models for Self-Supervised Learning
Xinlei Chen, Zhuang Liu, Saining Xie, Kaiming He
Introduction
Denoising is at the core in the current trend of generative models in computer vision and other areas. Popularly known as Denoising Diffusion Models (DDM) today, these methods learn a Denoising Autoencoder (DAE) that removes noise of multiple levels driven by a diffusion process. These methods achieve impressive image generation quality, especially for high-resolution, photo-realistic images —in fact, these generation models are so good that they appear to have strong recognition representations for understanding the visual content.
While DAE is a powerhouse of today’s generative models, it was originally proposed for learning representations from data in a self-supervised manner. In today’s community of representation learning, the arguably most successful variants of DAEs are based on “masking noise” , such as predicting missing text in languages (e.g., BERT ) or missing patches in images (e.g., MAE ). However, in concept, these masking-based variants remain significantly different from removing additive (e.g., Gaussian) noise: while the masked tokens explicitly specify unknown vs. known content, no clean signal is available in the task of separating additive noise. Nevertheless, today’s DDMs for generation are dominantly based on additive noise, implying that they may learn representations without explicitly marking unknown/known content.
Most recently, there has been an increasing interest in inspecting the representation learning ability of DDMs. In particular, these studies directly take off-the-shelf pre-trained DDMs , which are originally purposed for generation, and evaluate their representation quality for recognition. They report encouraging results using these generation-oriented models. However, these pioneering studies obviously leave open questions: these off-the-shelf models were designed for generation, not recognition; it remains largely unclear whether the representation capability is gained by a denoising-driven process, or a diffusion-driven process.
In this work, we take a much deeper dive into the direction initialized by these recent explorations . Instead of using an off-the-shelf DDM that is generation-oriented, we train models that are recognition-oriented. At the core of our philosophy is to deconstruct a DDM, changing it step-by-step into a classical DAE. Through this deconstructive research process, we examine every single aspect (that we can think of) of a modern DDM, with the goal of learning representations. This research process gains us new understandings on what are the critical components for a DAE to learn good representations.
Surprisingly, we discover that the main critical component is a tokenizer that creates a low-dimensional latent space. Interestingly, this observation is largely independent of the specifics of the tokenizer—we explore a standard VAE , a patch-wise VAE, a patch-wise AE, and a patch-wise PCA encoder. We discover that it is the low-dimensional latent space, rather than the tokenizer specifics, that enables a DAE to achieve good representations.
Thanks to the effectiveness of PCA, our deconstructive trajectory ultimately reaches a simple architecture that is highly similar to the classical DAE (Fig. 1). We project the image onto a latent space using patch-wise PCA, add noise, and then project it back by inverse PCA. Then we train an autoencoder to predict a denoised image. We call this architecture “latent Denoising Autoencoder” (l-DAE).
Our deconstructive trajectory also reveals many other intriguing properties that lie between DDM and classical DAE. For one example, we discover that even using a single noise level (i.e., not using the noise scheduling of DDM) can achieve a decent result with our l-DAE. The role of using multiple levels of noise is analogous to a form of data augmentation, which can be beneficial, but not an enabling factor. With this and other observations, we argue that the representation capability of DDM is mainly gained by the denoising-driven process, not a diffusion-driven process.
Finally, we compare our results with previous baselines. On one hand, our results are substantially better than the off-the-shelf counterparts (following the spirit of ): this is as expected, because these are our starting point of deconstruction. On the other hand, our results fall short of baseline contrastive learning methods (e.g., ) and masking-based methods (e.g., ), but the gap is reduced. Our study suggests more room for further research along the direction of DAE and DDM.
Related Work
In the history of machine learning and computer vision, the generation of images (or other content) has been closely intertwined with the development of unsupervised or self-supervised learning. Approaches in generation are conceptually forms of un-/self-supervised learning, where models were trained without labeled data, learning to capture the underlying distributions of the input data.
There has been a prevailing belief that the ability of a model to generate high-fidelity data is indicative of its potential for learning good representations. Generative Adversarial Networks (GAN) , for example, have ignited broad interest in adversarial representation learning . Variational Autoencoders (VAEs) , originally conceptualized as generative models for approximating data distributions, have evolved to become a standard in learning localized representations (“tokens”), e.g., VQVAE and variants . Image inpainting , essentially a form of conditional image generation, has led to a family of modern representation learning methods, including Context Encoder and Masked Autoencoder (MAE) .
Analogously, the outstanding generative performance of Denoising Diffusion Models (DDM) has drawn attention for their potential in representation learning. Pioneering studies have begun to investigate this direction by evaluating existing pre-trained DDM networks. However, we note that while a model’s generation capability suggests a certain level of understanding, it does not necessarily translate to representations useful for downstream tasks. Our study delves deeper into these issues.
On the other hand, although Denoising Autoencoders (DAE) have laid the groundwork for autoencoding-based representation learning, their success has been mainly confined to scenarios involving masking-based corruption (e.g., ). To the best of our knowledge, little or no recent research has reported results on classical DAE variants with additive Gaussian noise, and we believe that the underlying reason is that a simple DAE baseline (Fig. 2(a)) performs poorlyAccording to the authors of MoCo and MAE , significant effort has been devoted to DAE baselines during the development of those works, following the best practice established. However, it has not led to meaningful results (20% accuracy). (e.g., in 20% Fig. 5).
Background: Denoising Diffusion Models
Our deconstructive research starts with a Denoising Diffusion Model (DDM) . We briefly describe the DDM we use, following .
A diffusion process starts from a clean data point () and sequentially adds noise to it. At a specified time step , the noised data is given by:
where is a noise map sampled from a Gaussian distribution, and and define the scaling factors of the signal and of the noise, respectively. By default, it is set .
A denoising diffusion model is learned to remove the noise, conditioned on the time step . Unlike the original DAE that predicts a clean input, the modern DDM often predicts the noise . Specifically, a loss function in this form is minimized:
where is the network output. The network is trained for multiple noise levels given a noise schedule, conditioned on the time step . In the generation process, a trained model is iteratively applied until it reaches the clean signal .
DDMs can operate on two types of input spaces. One is the original pixel space , where the raw image is directly used as . The other option is to build DDMs on a latent space produced by a tokenizer, following . See Fig. 2(b). In this case, a pre-trained tokenizer (which is often another autoencoder, e.g., VQVAE ) is used to map the image into its latent .
Our study begins with the Diffusion Transformer (DiT) . We choose this Transformer-based DDM for several reasons: (i) Unlike other UNet-based DDMs , Transformer-based architectures can provide fairer comparisons with other self-supervised learning baselines driven by Transformers (e.g., ); (ii) DiT has a clearer distinction between the encoder and decoder, while a UNet’s encoder and decoder are connected by skip connections and may require extra effort on network surgery when evaluating the encoder; (iii) DiT trains much faster than other UNet-based DDMs (see ) while achieving better generation quality.
We use the DiT-Large (DiT-L) variant as our DDM baseline. In DiT-L, the encoder and decoder put together have the size of ViT-L (24 blocks). We evaluate the representation quality (linear probe accuracy) of the encoder, which has 12 blocks, referred to as “L” (half large).
DiT instantiated in is a form of Latent Diffusion Models (LDM) , which uses a VQGAN tokenizer . Specifically, this VQGAN tokenizer transforms the 2562563 input image (heightwidthchannels) into a 32324 latent map, with a stride of 8.
By default, we train the models for 400 epochs on ImageNet with a resolution of 256256 pixels. Implementation details are in Sec. A.
Our DiT baseline results are reported in Tab. 1 (line 1). With DiT-L, we report a linear probe accuracy of 57.5% using its L encoder. The generation quality (Fréchet Inception Distance , FID-50K) of this DiT-L model is 11.6. This is the starting point of our destructive trajectory.
Despite differences in implementation details, our starting point conceptually follows recent studies (more specifically, DDAE ), which evaluate off-the-shelf DDMs under the linear probing protocol.
Deconstructing Denoising Diffusion Models
Our deconstruction trajectory is divided into three stages. We first adapt the generation-focused settings in DiT to be more oriented toward self-supervised learning (Sec. 4.1). Next, we deconstruct and simplify the tokenizer step by step (Sec. 4.2). Finally, we attempt to reverse as many DDM-motivated designs as possible, pushing the models towards a classical DAE (Sec. 4.3). We summarize our learnings from this deconstructing process in Sec. 4.4.
While a DDM is conceptually a form of a DAE, it was originally developed for the purpose of image generation. Many designs in a DDM are oriented toward the generation task. Some designs are not legitimate for self-supervised learning (e.g., class labels are involved); some others are not necessary if visual quality is not concerned. In this subsection, we reorient our DDM baseline for the purpose of self-supervised learning, summarized in Tab. 1.
A high-quality DDM is often trained with conditioning on class labels, which can largely improve the generation quality. But the usage of class labels is simply not legitimate in the context of our self-supervised learning study. As the first step, we remove class-conditioning in our baseline.
Surprisingly, removing class-conditioning substantially improves the linear probe accuracy from 57.5% to 62.1% (Tab. 1), even though the generation quality is greatly hurt as expected (FID from 11.6 to 34.2). We hypothesize that directly conditioning the model on class labels may reduce the model’s demands on encoding the information related to class labels. Removing the class-conditioning can force the model to learn more semantics.
In our baseline, the VQGAN tokenizer, presented by LDM and inherited by DiT, is trained with multiple loss terms: (i) autoencoding reconstruction loss; (ii) KL-divergence regularization loss ;The KL form in does not perform explicit vector quantization (VQ), interpreted as “the quantization layer absorbed by the decoder” . (iii) perceptual loss based on a supervised VGG net trained for ImageNet classification; and (iv) adversarial loss with a discriminator. We ablate the latter two terms in Tab. 1.
As the perceptual loss involves a supervised pre-trained network, using the VQGAN trained with this loss is not legitimate. Instead, we train another VQGAN tokenizer in which we remove the perceptual loss. Using this tokenizer reduces the linear probe accuracy significantly from 62.5% to 58.4% (Tab. 1), which, however, provides the first legitimate entry thus far. This comparison reveals that a tokenizer trained with the perceptual loss (with class labels) in itself provides semantic representations. We note that the perceptual loss is not used from now on, in the remaining part of this paper.
We train the next VQGAN tokenizer that further removes the adversarial loss. It slightly increases the linear probe accuracy from 58.4% to 59.0% (Tab. 1). With this, our tokenizer at this point is essentially a VAE, which we move on to deconstruct in the next subsection. We also note that removing either loss harms generation quality.
In the task of generation, the goal is to progressively turn a noise map into an image. As a result, the original noise schedule spends many time steps on very noisy images (Fig. 3). This is not necessary if our model is not generation-oriented.
We study a simpler noise schedule for the purpose of self-supervised learning. Specifically, we let linearly decay in the range of (Fig. 3). This allows the model to spend more capacity on cleaner images. This change greatly improves the linear probe accuracy from 59.0% to 63.4% (Tab. 1), suggesting that the original schedule focuses too much on noisier regimes. On the other hand, as expected, doing so further hurts the generation ability, leading to a FID of 93.2.
Overall, the results in Tab. 1 reveal that self-supervised learning performance is not correlated to generation quality. The representation capability of a DDM is not necessarily the outcome of its generation capability.
2 Deconstructing the Tokenizer
Next, we further deconstruct the VAE tokenizer by making substantial simplifications. We compare the following four variants of autoencoders as the tokenizers, each of which is a simplified version of the preceding one:
Our deconstruction thus far leads us to a VAE tokenizer. As common practice , the encoder and decoder of this VAE are deep convolutional (conv) neural networks . This convolutional VAE minimizes the following loss function:
Here, is the input image of the VAE. The first term is the reconstruction loss, and the second term is the Kullback-Leibler divergence between the latent distribution of and a unit Gaussian distribution.
Next we consider a simplified case in which the VAE encoder and decoder are both linear projections, and the VAE input is a patch. The training process of this patch-wise VAE minimizes this loss:
Here denotes a patch flattened into a -dimensional vector. Both and are matrixes, where is the dimension of the latent space. We set the patch size as 1616 pixels, following .
We make further simplification on VAE by removing the regularization term:
As such, this tokenizer is essentially an autoencoder (AE) on patches, with the encoder and decoder both being linear projections.
Finally, we consider a simpler variant which performs Principal Component Analysis (PCA) on the patch space. It is easy to show that PCA is equivalent to a special case of AE:
in which satisfies ( identity matrix). The PCA bases can be simply computed by eigen-decomposition on a large set of randomly sampled patches, requiring no gradient-based training.
Latent dimension of the tokenizer is crucial for DDM to work well in self-supervised learning.
As shown in Tab. 2, all four variants of tokenizers exhibit similar trends, despite their differences in architectures and loss functions. Interestingly, the optimal dimension is relatively low ( is 16 or 32), even though the full dimension per patch is much higher (16163=768).
Surprisingly, the convolutional VAE tokenizer is neither necessary nor favorable; instead, all patch-based tokenizers, in which each patch is encoded independently, perform similarly with each other and consistently outperform the Conv VAE variant. In addition, the KL regularization term is unnecessary, as both the AE and PCA variants work well.
To our further surprise, even the PCA tokenizer works well. Unlike the VAE or AE counterparts, the PCA tokenizer does not require gradient-based training. With pre-computed PCA bases, the application of the PCA tokenizer is analogous to a form of image pre-processing, rather than a “network architecture”. The effectiveness of a PCA tokenizer largely helps us push the modern DDM towards a classical DAE, as we will show in the next subsection.
High-resolution, pixel-based DDMs are inferior for self-supervised learning.
Before we move on, we report an extra ablation that is consistent with the aforementioned observation. Specifically, we consider a “naïve tokenizer” that performs identity mapping on patches extracted from resized images. In this case, a “token” is the flatten vector consisting all pixels of a patch.
In Fig. 5, we show the results of this “pixel-based” tokenizer, operated on an image size of 256, 128, 64, and 32, respectively with a patch size of 16, 8, 4, 2. The “latent” dimensions of these tokenized spaces are 768, 192, 48, and 12 per token. In all case, the sequence length of the Transformer is kept unchanged (256).
Interestingly, this pixel-based tokenizer exhibits a similar trend with other tokenizers we have studied, although the optimal dimension is shifted. In particular, the optimal dimension is =48, which corresponds to an image size of 64 with a patch size of 4. With an image size of 256 and a patch size of 16 (=768), the linear probe accuracy drops dramatically to 23.6%.
These comparisons show that the tokenizer and the resulting latent space are crucial for DDM/DAE to work competitively in the self-supervised learning scenario. In particular, applying a classical DAE with additive Gaussian noise on the pixel space leads to poor results.
3 Toward Classical Denoising Autoencoders
Next, we go on with our deconstruction trajectory and aim to get as close as possible to the classical DAE . We attempt to remove every single aspect that still remains between our current PCA-based DDM and the classical DAE practice. Via this deconstructive process, we gain better understandings on how every modern design may influence the classical DAE. Tab. 3 gives the results, discussed next.
While modern DDMs commonly predict the noise (see Eq. 2), the classical DAE predicts the clean data instead. We examine this difference by minimizing the following loss function:
Here is the clean data (in the latent space), and is the network prediction. is a -dependent loss weight, introduced to balance the contribution of different noise levels . It is suggested to set as per . We find that setting works better in our scenario. Intuitively, it simply puts more weight to the loss terms of the cleaner data (larger ).
With the modification of predicting clean data (rather than noise), the linear probe accuracy degrades from 65.1% to 62.4% (Tab. 3). This suggests that the choice of the prediction target influences the representation quality.
Even though we suffer from a degradation in this step, we will stick to this modification from now on, as our goal is to move towards a classical DAE.We have revisited undoing this change in our final entry, in which we have not observed this degradation.
In modern DDMs (see Eq. 1), the input is scaled by a factor of . This is not common practice in a classical DAE. Next, we study removing input scaling, i.e., we set . As is fixed, we need to define a noise schedule directly on . We simply set as a linear schedule from 0 to . Moreover, we empirically set the weight in Eq. 3 as , which again puts more emphasis on cleaner data (smaller ).
After fixing , we achieve a decent accuracy of 63.6% (Tab. 3), which compares favorably with the varying counterpart’s 62.4%. This suggests that scaling the data by is not necessary in our scenario.
Thus far, for all entries we have explored (except Fig. 5), the model operates on the latent space produced by a tokenizer (Fig. 2 (b)). Ideally, we hope our DAE can work directly on the image space while still having good accuracy. With the usage of PCA, we can achieve this goal by inverse PCA.
The idea is illustrated in Fig. 1. Specially, we project the input image into the latent space by the PCA bases (i.e., ), add noise in the latent space, and project the noisy latent back to the image space by the inverse PCA bases (). Fig. 1 (middle, bottom) shows an example image with noise added in the latent space. With this noisy image as the input to the network, we can apply a standard ViT network that directly operate on images, as if there is no tokenizer.
Applying this modification on the input side (while still predicting the output on the latent space) has 63.6% accuracy (Tab. 3). Further applying it to the output side (i.e., predicting the output on the image space with inverse PCA) has 63.9% accuracy. Both results show that operating on the image space with inverse PCA can achieve similar results as operating on the latent space.
While inverse PCA can produce a prediction target in the image space, this target is not the original image. This is because PCA is a lossy encoder for any reduced dimension . In contrast, it is a more natural solution to predict the original image directly.
When we let the network predict the original image, the “noise” introduced includes two parts: (i) the additive Gaussian noise, whose intrinsic dimension is , and (ii) the PCA reconstruction error, whose intrinsic dimension is ( is 768). We weight the loss of both parts differently.
Formally, with the clean original image and network prediction , we can compute the residue projected onto the full PCA space: . Here is the -by- matrix representing the full PCA bases. Then we minimize the following loss function:
Here denotes the -th dimension of the vector . The per-dimension weight is 1 for , and 0.1 for . Intuitively, down-weights the loss of the PCA reconstruction error. With this formulation, predicting the original image achieves 64.5% linear probe accuracy (Tab. 3).
This variant is conceptually very simple: its input is a noisy image whose noise is added in the PCA latent space, its prediction is the original clean image (Fig. 1).
Lastly, out of curiosity, we further study a variant with single-level noise. We note that multi-level noise, given by noise scheduling, is a property motived by the diffusion process in DDMs; it is conceptually unnecessary in a classical DAE.
We fix the noise level as a constant (). Using this single-level noise achieves decent accuracy of 61.5%, a 3% degradation vs. the multi-level noise counterpart (64.5%). Using multiple levels of noise is analogous to a form of data augmentation in DAE: it is beneficial, but not an enabling factor. This also implies that the representation capability of DDM is mainly gained by the denoising-driven process, not a diffusion-driven process.
As multi-level noise is useful and conceptually simple, we keep it in our final entries presented in the next section.
4 Summary
In sum, we deconstruct a modern DDM and push it towards a classical DAE (Fig. 6). We undo many of the modern designs and conceptually retain only two designs inherited from modern DDMs: (i) a low-dimensional latent space in which noise is added; and (ii) multi-level noise.
We use the entry at the end of Tab. 3 as our final DAE instantiation (illustrated in Fig. 1). We refer to this method as “latent Denoising Autoencoder”, or in short, l-DAE.
Analysis and Comparisons
Conceptually, l-DAE is a form of DAE that learns to remove noise added to the latent space. Thanks to the simplicity of PCA, we can easily visualize the latent noise by inverse PCA.
Fig. 7 compares the noise added to pixels vs. to the latent. Unlike the pixel noise, the latent noise is largely independent of the resolution of the image. With patch-wise PCA as the tokenizer, the pattern of the latent noise is mainly determined by the patch size. Intuitively, we may think of it as using patches, rather than pixels, to resolve the image. This behavior resembles MAE , which masks out patches instead of individual pixels.
Fig. 8 shows more examples of denoising results based on l-DAE. Our method produces reasonable predictions despite of the heavy noise. We note that this is less of a surprise, because neural network-based image restoration has been an intensively studied field.
Nevertheless, the visualization may help us better understand how l-DAE may learn good representations. The heavy noise added to the latent space creates a challenging pretext task for the model to solve. It is nontrivial (even for human beings) to predict the content based on one or a few noisy patches locally; the model is forced to learn higher-level, more holistic semantics to make sense of the underlying objects and scenes.
Notably, all models we present thus far have no data augmentation: only the center crops of images are used, with no random resizing or color jittering, following . We further explore a mild data augmentation (random resized crop) for our final l-DAE:
which has slight improvement. This suggests that the representation learning ability of -DAE is largely independent of its reliance on data augmentation. A similar behavior was observed in MAE , which sharply differs from the behavior of contrastive learning methods (e.g., ).
All our experiments thus far are based on 400-epoch training. Following MAE , we also study training for 800 and 1600 epochs:
As a reference, MAE has a significant gain (4%) extending from 400 to 800 epochs, and MoCo v3 has nearly no gain (0.2%) extending from 300 to 600 epochs.
Thus far, our all models are based on the DiT-L variant , whose encoder and decoder are both “ViT-L” (half depth of ViT-L). We further train models of different sizes, whose encoder is ViT-B or ViT-L (decoder is always of the same size as encoder):
We observe a good scaling behavior w.r.t. model size: scaling from ViT-B to ViT-L has a large gain of 10.6%. A similar scaling behavior is also observed in MAE , which has a 7.8% gain from ViT-B to ViT-L.
Finally, to have a better sense of how different families of self-supervised learning methods perform, we compare with previous baselines in Tab. 4. We consider MoCo v3 , which belongs to the family of contrastive learning methods, and MAE , which belongs to the family of masking-based methods.
Interestingly, l-DAE performs decently in comparison with MAE, showing a degradation of 1.4% (ViT-B) or 0.8% (ViT-L). We note that here the training settings are made as fair as possible between MAE and l-DAE: both are trained for 1600 epochs and with random crop as the data augmentation. On the other hand, we should also note that MAE is more efficient in training because it only operates on unmasked patches. Nevertheless, we have largely closed the accuracy gap between MAE and a DAE-driven method.
Last, we observe that autoencoder-based methods (MAE and l-DAE) still fall short in comparison with contrastive learning methods under this protocol, especially when the model is small. We hope our study will draw more attention to the research on autoencoder-based methods for self-supervised learning.
Conclusion
We have reported that l-DAE, which largely resembles the classical DAE, can perform competitively in self-supervised learning. The critical component is a low-dimensional latent space on which noise is added. We hope our discovery will reignite interest in denoising-based methods in the context of today’s self-supervised learning research.
We thank Pascal Vincent, Mike Rabbat, and Ross Girshick for their discussion and feedback.
A Implementation Details
We follow the DiT architecture design . The DiT architecture is similar to the original ViT , with extra modifications made for conditioning. Each Transformer block accepts an embedding network (a two-layer MLP) conditioned on the time step . The output of this embedding network determines the scale and bias parameters of LayerNorm , referred to as adaLN . Slightly different from , we set the hidden dimension of this MLP as 1/4 of its original dimension, which helps reduce model sizes and save memory, at no accuracy cost.
The original DiTs are trained with a batch size of 256. To speed up our exploration, we increase the batch size to 2048. We perform linear learning rate warm up for 100 epochs and then decay it following a half-cycle cosine schedule. We use a base learning rate blr = 1e-4 by default, and set the actual lr following the linear scaling rule : blr batch_size / 256. No weight decay is used . We train for 400 epochs by default. On a 256-core TPU-v3 pod, training DiT-L takes 12 hours.
Our linear probing implementation follows the practice of MAE . We use clean, -sized images for linear probing training and evaluation. The ViT output feature map is globally pooled by average pooling. It is then processed by a parameter-free BatchNorm layer and a linear classifier layer, following . The training batch size is 16384, learning rate is (cosine decay schedule), weight decay is 0, and training length is 90 epochs. Randomly resized crop and flipping are used during training and a single center crop is used for testing. Top-1 accuracy is reported.
While the model is conditioned on in self-supervised pre-training, conditioning is not needed in transfer learning (e.g., linear probing). We fix the time step value in our linear probing training and evaluation. The influence of different values (out of 1000 time steps) is shown as follows:
We note that the value determines: (i) the model weights, which are conditioned on , and (ii) the noise added in transfer learning, using the same level of . Both are shown in this table. We use = 10 and clean input in all our experiments, except Tab. 4 where we use the optimal setting.
Fixing also means that the -dependent MLP layers, which are used for conditioning, are not exposed in transfer learning, because they can be merged given the fixed . As such, our model has the number of parameters just similar to the standard ViT , as reported in Tab. 4.
The DiT-L has 24 blocks where the first 12 blocks are referred to as the “encoder” (hence ViT-L) and the others the “decoder”. This separation of the encoder and decoder is artificial. In the following table, we show the linear probing results using different numbers of blocks in the encoder, using the same pre-trained model:
The optimal accuracy is achieved when the encoder and decoder have the same depth. This behavior is different from MAE’s , whose encoder and decoder are asymmetric.
B Fine-tuning Results
In addition to linear probing, we also report end-to-end fine-tuning results. We closely followed MAE’s protocol . We use clean, -sized images as the inputs to the encoder. Globally average pooled outputs are used as features for classification. The training batch size is 1024, initial learning rate is , weight decay is 0.05, drop path is 0.1, and training length is 100 epochs. We use a layer-wise learning rate decay of 0.85 (B) or 0.65 (L). MixUp (0.8), CutMix (1.0), RandAug (9, 0.5), and exponential moving average (0.9999) are used, similar to . The results are summarized as below:
Overall, both autoencoder-based methods are better than MoCo v3. l-DAE performs similarly with MAE with ViT-B, but still fall short of MAE with ViT-L.