EfficientTrain: Exploring Generalized Curriculum Learning for Training Visual Backbones
Yulin Wang, Yang Yue, Rui Lu, Tianjiao Liu, Zhao Zhong, Shiji Song, Gao Huang
Introduction
The success of modern visual backbones is largely fueled by the interest in exploring big models on large-scale benchmark datasets . In particular, the recent introduction of vision Transformers (ViTs) scales up the number of model parameters to more than 1.8 billion, with the training data expanding to 3 billion samples . Although state-of-the-art accuracy has been achieved, this huge-model and high-data regime results in a time-consuming and expensive training process. For example, it takes 2,500 TPUv3-core-days to train ViT-H/14 on JFT-300M , which may be unaffordable for practitioners in both academia and industry. Additionally, the power consumption leads to significant carbon emissions . Due to both economic and environmental concerns, there has been a growing demand for reducing the training cost of modern deep networks.
In this paper, we contribute to this issue by revisiting the idea of curriculum learning , which reveals that a model can be trained efficiently by starting with the easier aspects of a given task or certain easier subtasks, and increasing the difficulty level gradually. Most existing works implement this idea by introducing easier-to-harder examples progressively during training . However, obtaining a light-weighted and generalizable difficulty measurer is typically non-trivial . In general, these methods have not exhibited the capacity to be a universal efficient training technique for modern visual backbones.
In contrast to prior works, this paper seeks a simple yet broadly applicable efficient learning approach with the potential for widespread implementation. To attain this goal, we consider a generalization of curriculum learning. In specific, we extend the notion of training curricula beyond only differentiating between ‘easier’ and ‘harder’ examples, and adopt a more flexible hypothesis, which indicates that the discriminative features of each training sample comprise both ‘easier-to-learn’ and ‘harder-to-learn’ patterns. Instead of making a discrete decision on whether each example should appear in the training set, we argue that it would be more proper to establish a continuous function that adaptively extracts the simpler and more learnable discriminative patterns within every example. In other words, a curriculum may always leverage all examples at any stage of learning, but it should eliminate the relatively more difficult or complex patterns within inputs at earlier learning stages. An illustration of our idea is shown in Figure 1.
Driven by our hypothesis, a straightforward yet surprisingly effective algorithm is derived. We first demonstrate that the ‘easier-to-learn’ patterns incorporate the lower-frequency components of images. We further show that a lossless extraction of these components can be achieved by introducing a cropping operation in the frequency domain. This operation not only retains exactly all the lower-frequency information, but also yields a smaller input size for the model to be trained. By triggering this operation at earlier training stages, the overall computational/time cost for training can be considerably reduced while the final performance of the model will not be sacrificed. Moreover, we show that the original information before heavy data augmentation is more learnable, and hence starting the training with weaker augmentation techniques is beneficial. Finally, these theoretical and experimental insights are integrated into a unified ‘EfficientTrain’ learning curriculum by leveraging a greedy-search algorithm.
One of the most appealing advantages of EfficientTrain may be its simplicity and generalizability. Our method can be conveniently applied to most deep networks without any modification or hyper-parameter tuning, but significantly improves their training efficiency. Empirically, for the supervised learning on ImageNet-1K/22K , EfficientTrain reduces the wall-time training cost of a wide variety of popular visual backbones (e.g., ConvNeXt , DeiT , PVT , Swin , and CSWin ) by more than , while achieving competitive or better performance compared with the baselines. Importantly, our method is also effective for self-supervised learning (e.g., MAE ).
Related Work
Curriculum learning is a training paradigm inspired by the organized learning order of examples in human curricula . This idea has been widely explored in the context of training deep networks from easier data to harder data . Typically, a pre-defined or automatically-learned difficulty measurer is deployed to differentiate between easier and harder samples, while a scheduler is defined to determine when and how to introduce harder training data. Our method is also based on the ‘starting small’ spirit , but we always leverage all the training data simultaneously. Our work is also related to curriculum by smoothing and curriculum dropout , which do not perform example selection as well. However, our method is orthogonal to them since we reduce the training cost by modifying the model inputs, while they regularize deep features during training (e.g., via anti-aliasing smoothing or feature dropout).
Progressive or modularized training. Deep networks can be trained efficiently by increasing the model size during training, e.g., a growing number of layers , a growing width , or a dynamically changed network connection topology . These methods are mainly motivated by that smaller models are more efficient to train at earlier epochs. This idea is also explored in language models , recommendation systems and graph ConvNets . Locally supervised learning, which trains different model components using tailored objectives, is a promising direction as well .
A similar work to us is progressive learning (PL) , which down-samples the images to save the training cost. Nevertheless, our work differs from PL in several important aspects: 1) EfficientTrain is drawn from a distinctly different motivation of generalized curriculum learning, based on which we present a novel frequency-inspired analysis; 2) we introduce a cropping operation in the frequency domain, which is not only theoretically different from the down-sampling operation in PL (see: Proposition 1), but also outperforms it empirically (see: Table 13 (b)); 3) from the perspective of system-level comparison, we design an EfficientTrain curriculum, achieving a significantly higher training efficiency than PL on a variety of state-of-the-art models (see: Tables 9). In addition, FixRes shows that a smaller training resolution may improve the accuracy by fixing the discrepancy between the scale of training and test inputs. However, we do not borrow gains from FixRes as we adopt a standard resolution at the final stages of training. Our method is actually orthogonal to FixRes (see: Table 9).
Frequency-based analysis of deep networks. Our observation that deep networks tend to capture the low-frequency components first is inline with , but the discussions in focus on the robustness of ConvNets and are mainly based on some small models and tiny datasets. Towards this direction, several existing works also explore decomposing the inputs of models in the frequency domain in order to understand or improve the robustness of the networks. In contrast, our aim is to improve the training efficiency of modern deep visual backbones.
A Generalization of Curriculum Learning
As uncovered in previous research, machine learning algorithms generally benefit from a ‘starting small’ strategy, i.e., to first learn certain easier aspects of a task, and increase the level of difficulty progressively . The dominant implementation of this idea, curriculum learning, proposes to introduce gradually more difficult examples during training . In specific, a curriculum is defined on top of the training process to determine whether or not each sample should be leveraged at a given epoch (Figure 1 (a)).
On the limitations of sample-wise curriculum learning. Although curriculum learning has been widely explored from the lens of the sample-wise regime, its extensive application is usually limited by two major issues. First, differentiating between ‘easier’ and ‘harder’ training data is non-trivial. It typically requires deploying additional deep networks as a ‘teacher’ or exploiting specialized automatic learning approaches . The resulting implementation complexity and the increased overall computational cost are both noteworthy weaknesses in terms of improving the training efficiency. Second, it is challenging to attain a principled approach that specifies which examples should be attended to at the earlier stages of learning. As a matter of fact, the ‘easy to hard’ strategy is not always helpful . The hard-to-learn samples can be more informative and may be beneficial to be emphasized in many cases , sometimes even leading to a ‘hard to easy’ anti-curriculum .
Our work is inspired by the above two issues. In the following, we start by proposing a generalized hypothesis for curriculum learning, aiming to address the second issue. Then we demonstrate that an implementation of our idea naturally addresses the first issue.
Generalized curriculum learning. We argue that simply measuring the easiness of training samples tends to be ambiguous and may be insufficient to reflect the effects of a sample on the learning process. As aforementioned, even the difficult examples may provide beneficial information for guiding the training, and they do not necessarily need to be introduced after easier examples. To this end, we hypothesize that every training sample, either ‘easier’ or ‘harder’, contains both easier-to-learn or more accessible patterns, as well as certain difficult discriminative information which may be challenging for the deep networks to capture. Ideally, a curriculum should be a continuous function on top of the training process, which starts with a focus on the ‘easiest’ patterns of the inputs, while the ‘harder-to-learn’ patterns are gradually introduced as learning progresses.
A formal illustration is shown in Figure 1 (b). Any input data will be processed by a transformation function conditioned on the training epoch before being fed into the model, where is designed to dynamically filter out the excessively difficult and less learnable patterns within the training data. We always let . Notably, our approach can be seen as a generalized form of the sample-wise curriculum learning. It reduces to example-selection by setting .
Overview. In the rest of this paper, we will demonstrate that an algorithm drawn from our hypothesis dramatically improves the implementation efficiency and generalization ability of curriculum learning. We will show that a zero-cost criterion pre-defined by humans is able to effectively measure the difficulty level of different patterns within images. Based on such simple criteria, even a surprisingly straightforward implementation of introducing ‘easier-to-harder’ patterns yields significant and consistent improvements on the training efficiency of modern visual backbones.
The EfficientTrain Approach
To obtain a training curriculum following our aforementioned hypothesis, we need to solve two challenges: 1) identifying the ‘easier-to-learn’ patterns and designing transformation functions to extract them; 2) establishing a curriculum learning schedule to perform these transformations dynamically during training. This section will demonstrate that a proper transformation for 1) can be easily found in both the frequency and the spatial domain, while 2) can be addressed with a greedy search algorithm. Implementation details of the experiments in this section: see Appendix A.
Image-based data can naturally be decomposed in the frequency domain . In this subsection, we reveal that the patterns in the lower-frequency components of images, which describe the smoothly changing contents, are relatively easier for the networks to learn to recognize.
Ablation studies with the low-pass filtered input data. We first consider an ablation study, where the low-pass filtering is performed on the data we use. As shown in Figure 2 (a), we map the images to the Fourier spectrum with the lowest frequency at the centre, set all the components outside a centred circle (radius: ) to zero, and map the spectrum back to the pixel space. Figure 2 (b) illustrates the effects of . The curves of accuracy v.s. training epochs on top of the low-pass filtered data are presented in Figure 3. Here both training and validation data is processed by the filter to ensure the compatibility with the i.i.d. assumption.
Lower-frequency components are captured first. The models in Figure 3 are imposed to leverage only the lower-frequency components of the inputs. However, an appealing phenomenon arises: their training process is approximately identical to the original baseline at the beginning of training. Although the baseline finally outperforms, this tendency starts midway in the training process, instead of from the very beginning. In other words, the learning behaviors at earlier epochs remain unchanged even though the higher-frequency components of images are eliminated. Moreover, consider increasing the filter bandwidth , which preserves progressively more information about the images from the lowest frequency. The separation point between the baseline and the training process on low-pass filtered data moves towards the end of training. To explain these observations, we postulate that, in a natural learning process where the input images contain both lower- and higher-frequency information, a model tends to first learn to capture the lower-frequency components, while the higher-frequency information is gradually exploited on the basis of them.
More evidences. Our assumption can be further confirmed by a well-controlled experiment. Consider training a model using original images, where lower/higher-frequency components are simultaneously provided. In Figure 4, we evaluate all the intermediate checkpoints on low-pass filtered validation sets with varying bandwidths. Obviously, at earlier epochs, only leveraging the low-pass filtered validation data does not degrade the accuracy. This phenomenon suggests that the learning process starts with a focus on the lower-frequency information, even though the model is always accessible to higher-frequency components during training. Furthermore, in Table 1, we compare the accuracies of the intermediate checkpoints on low/high-pass filtered validation sets. We find the accuracy on the low-pass filtered validation set grows much faster at earlier training stages, even though the two final accuracies are the same.
Frequency-based curricula. Returning to our hypothesis in Section 3, we have shown that lower-frequency components are naturally captured earlier. Hence, it would be straightforward to consider them as a type of the ‘easier-to-learn’ patterns. This begs a question: can we design a training curriculum, which starts with providing only the lower-frequency information for the model, while gradually introducing the higher-frequency components? We investigate this idea in Table 2, where we perform low-pass filtering on the training data only in a given number of the beginning epochs. The rest of the training process remains unchanged.
Learning from low-frequency information efficiently. At the first glance, the results in Table 2 may be less dramatic, i.e., by processing the images with a properly-configured low-pass filter at earlier epochs, the accuracy is moderately improved. However, an important observation is noteworthy: the final accuracy of the model can be largely preserved even with aggressive filtering (e.g., ) performed in 50-75% of the training process. This phenomenon turns our attention to training efficiency. At earlier learning stages, it is harmless to train the model with only the lower-frequency components. These components incorporate only a selected subset of all the information within the original input images. Hence, can we enable the model to learn from them efficiently with less computational cost than processing the original inputs? As a matter of fact, this idea is feasible, and we may have at least two approaches.
1) Down-sampling. Approximating the low-pass filtering in Table 2 with image down-sampling may be a straightforward solution. Down-sampling preserves much of the lower-frequency information, while it quadratically saves the computational cost for a model to process the inputs . However, it is not an operation tailored for extracting lower-frequency components. Theoretically, it preserves some of the higher-frequency components as well (see: Proposition 1). Empirically, we observe that this issue degrades the performance (see: Table 13 (b)).
2) Low-frequency cropping (see: Figure 5). We propose a more precise approach that extracts exactly all the lower-frequency information. Consider mapping an image into the frequency domain with the 2D discrete Fourier transform (DFT), obtaining an Fourier spectrum, where the value in the centre denotes the strength of the component with the lowest frequency. The positions distant from the centre correspond to higher-frequency. We crop a patch from the centre of the spectrum, where is the window size (). Since the patch is still centrosymmetric, we can map it back to the pixel space with the inverse 2D DFT, obtaining a new image , namely
where , and denote 2D DFT, inverse 2D DFT and centre-cropping. The computational or the time cost for accomplishing Eq. (1) is negligible on GPUs.
Notably, achieves a lossless extraction of lower-frequency components, while the higher-frequency parts are strictly eliminated. Hence, feeding into the model at earlier training stages can provide the vast majority of the useful information, such that the final accuracy will be minimally affected or even not affected. In contrast, importantly, due to the reduced input size of , the computational cost for a model to process is able to be dramatically saved, yielding a considerably more efficient training process.
Our claims can be empirically supported by Table 3, where we replace the low-pass filtering in Table 2 with the low-frequency cropping. Even such a straightforward implementation yields favorable results: the training cost can be saved by 20% while a competitive final accuracy is preserved. This phenomenon can be interpreted via Figure 6: at intermediate training stages, the models trained using the inputs with low-frequency cropping can learn similar deep representations to the baseline with a significantly reduced cost, i.e., the bright parts are clearly above the white lines.
Suppose that , and that , where down-sampling is realized by a common interpolation algorithm (e.g., nearest, bilinear or bicubic). Then we have two properties.
a) is only determined by the lower-frequency spectrum of (i.e., ). In addition, the mapping to is reversible. We can always recover from .
b) has a non-zero dependency on the higher-frequency spectrum of (i.e., the regions outside ).
2 Easier-to-learn Patterns: Spatial Domain
Apart from the frequency domain operations, extracting ‘easier-to-learn’ patterns can also be attained through spatial domain transformations. For example, deep networks (e.g., ViTs and ConvNets ) are typically trained with strong and delicate data augmentation techniques. We argue that the augmented training data provides a combination of both the information from original samples and the information introduced by augmentation operations. The original patterns may be ‘easier-to-learn’ as they are drawn from real-world distributions. This assumption is supported by the observation that data augmentation is mainly influential at the later stages of training .
To this end, following our generalized formulation of curriculum learning in Section 3, a curriculum may adopt a weaker-to-stronger data augmentation strategy during training. We investigate this idea by selecting RandAug as a representative example, which incorporates a family of common spatial-wise data augmentation transformations (rotate, sharpness, shear, solarize, etc.). In Table 4, the magnitude of RandAug is varied in the first half training process. One can observe that this idea improves the accuracy, and the gains are compatible with low-frequency cropping.
3 A Unified Training Curriculum
Finally, we integrate the techniques discussed above (i.e., low-frequency cropping at earlier epochs and weaker-to-stronger RandAug) to design a unified efficient training curriculum. We first set the magnitude of RandAug to be a linear function of the epoch : , with other data augmentation unchanged. Although being simple, this setting yields consistent empirical improvements. We adopt following the common practice .
Then we propose a greedy-search algorithm (Alg. 1) to determine the schedule of the bandwidth during training for low-frequency cropping. We divide the full training process into several stages and solve for a value of for each stage (a staircase approximation of the continuous curriculum learning schedule; see Appendix B.2 for more discussions). Alg. 1 starts from the last stage, minimizing under the constraint of not degrading the performance compared to the baseline accuracy . Here is obtained by training a model with a fixed , and does not change throughout Alg. 1. In Alg. 1, ValidationAccuracy(·) refers to training a new model with the given choices of , and evaluating the accuracy. The input is used here. For implementation, we only execute Alg. 1 for a single time. We obtain a schedule on top of Swin-Tiny under the standard 300-epoch training setting with , and directly adopt this schedule for other models or other training settings. Notably, the cost of Alg. 1 is to train a relatively small model (e.g., Swin-Tiny) for a small number of times (e.g., 7 to obtain our curriculum), where we can solve the minimization problems in Alg. 1 via simple linear search.
Derived from the aforementioned procedure, our finally proposed curriculum is presented in Table 5, which is named as EfficientTrain. Notably, despite its simplicity, it is general and surprisingly effective. It can be directly applied to most visual backbones under various training settings (e.g., different training budgets, varying amounts of training data, and supervised/self-supervised learning algorithms) without tuning additional hyper-parameters, and contributes to a significantly improved training efficiency.
Experiments
Datasets. Our main experiments are based on the large-scale ImageNet-1K/22K datasets, which consist of 1.28M/14.2M images in 1K/22K classes. We verify the transferability of pre-trained models on MS COCO , CIFAR , Flowers-102 , and Stanford Dogs .
Models. A wide variety of visual backbones are considered in our experiments, including ResNet , ConvNeXt , DeiT , PVT , Swin and CSWin Transformers. We adopt the training pipeline in , where EfficientTrain only modifies the terms mentioned in Table 5. Unless otherwise specified, we report the results of our implementation for both our method and the baselines. More implementation details can be found in Appendix A.
Training various visual backbones on ImageNet-1K. Table 6 presents the results of applying our method to train representative deep networks on ImageNet-1K. EfficientTrain achieves consistent improvements across different models, i.e., it reaches a competitive or better validation accuracy compared to the baselines (e.g., 83.6% v.s. 83.4% on the CSWin-Small network), while saving the training cost by . Importantly, the practical speedup is consistent with the theoretical results. The detailed training runtime (GPU-hours) is deferred to Appendix B.1.
ImageNet-22K pre-training. Our method exhibits excellent scalability with a growing training data scale or an increasing model size. In Table 7, a number of large backbones are pre-trained on ImageNet-22K w/ or w/o EfficientTrain, and evaluated by being fine-tuned to ImageNet-1K. Note that pre-training accounts for the vast majority of the total computation/time cost in this procedure. Our method performs at least on par with the baselines, while achieving a significant training speedup. A highlight from the results is that EfficientTrain saves the real training time considerably, e.g., 162 GPU-days (307.7 v.s. 469.5) for CSWin-Large, corresponding to 20 days for a 8-GPU node.
Adapted to varying epochs. EfficientTrain can conveniently adapt to a varying length of training schedule, i.e., by simply scaling the indices of epochs in Table 5. As shown in Table 6 (a), the advantage of our method is even more significant with fewer training epochs, e.g., it outperforms the baseline by 0.9% (76.4% v.s. 75.5%) for the 100-epoch trained DeiT-Small (the speedup is the same as 300-epoch). We attribute this to the greater importance of efficient training algorithms in the scenarios of limited training resources. Table 6 (b) further shows that our method can significantly improve the accuracy with the same training wall-time as the baseline (e.g., by 1.8% for DeiT-Tiny).
Adapted to any final input size . To adapt to a final input size , the value of for the three stages of EfficientTrain can be simply adjusted to . As shown in Table 6 (c), our method outperforms the baseline by large margins for in terms of training efficiency.
Orthogonal to pre-training + fine-tuning. In particular, in some cases, existing works find it efficient to fine-tune pre-trained models to a target test input size . Here our method can be directly leveraged for more efficient pre-training (e.g., in Table 7).
Comparisons with state-of-the-art efficient training methods are summarized in Table 9. EfficientTrain outperforms all the recently proposed baselines in terms of both accuracy and training efficiency. Moreover, the simplicity of our method enables it to be conveniently applied to different models and training settings, which may be non-trivial for other methods (e.g., the network-growing method ).
Orthogonal to FixRes. FixRes reveals that there exists a discrepancy in the scale of images between the training and test inputs, and thus the inference with a larger resolution will yield a better test accuracy. However, EfficientTrain does not leverage the gains of FixRes. We adopt the original inputs (e.g., ) at the final training stage, and hence the finally-trained model resembles the -trained networks, while FixRes is orthogonal to it. This fact can be confirmed by the evidence in both Table 9 (see: FixRes v.s. EfficientTrain + FixRes on top of the state-of-the-art CSWin Transformers ) and Table 7 (see: Input Size=).
Results with Masked Autoencoders (MAE). Since our method only modifies the training inputs, it can also be applied to self-supervised learning algorithms. As a representative example, we deploy EfficientTrain on top of MAE in Table 10. Our method reduces the training cost significantly while preserving a competitive accuracy.
Downstream image recognition tasks. In Table 11, we fine-tune the models trained with EfficientTrain to downstream classification datasets. Following , the CIFAR images are resized to , and hence their discriminative patterns are concentrated in the lower-frequency components. On the contrary, Flowers-102 and Stanford Dogs are fine-grained visual recognition datasets where the high-frequency clues contain important discriminative information. As shown in Table 11, EfficientTrain yields competitive transfer learning accuracy on both types of datasets. In other words, although our method learns to exploit the lower/higher-frequency information via an ordered curriculum, the finally obtained models can leverage both types of information effectively.
Object detection & instance segmentation. Table 12 investigates fine-tuning our pre-trained models to other computer vision tasks. Our method reduces the pre-training cost by , but performs at least on par with the baselines in terms of the detection/segmentation performance.
Ablation study. In Table 13 (a), we show that the major gain of training efficiency comes from low-frequency cropping, which effectively reduces the training cost at the price of a slight drop of accuracy. On top of this, linear RandAug further improves the accuracy. Moreover, replacing low-frequency cropping with image down-sampling consistently degrades the accuracy (see: Table 13 (b)), since down-sampling cannot strictly filter out all the higher-frequency information (see: Proposition 1), yielding a sub-optimal implementation of our idea. In addition, as shown in Table 13 (c), the schedule of found by Alg. 1 outperforms the heuristic design choices (e.g., the linear schedule in ).
Curves of val. accuracy during training are shown in Figure 7. The horizontal axis denotes the wall-time training cost. The low-frequency cropping in EfficientTrain is performed on both the training and test inputs. Our method is able to learn discriminative representations more efficiently at earlier epochs.
This paper investigated a novel generalized curriculum learning approach. The proposed algorithm, EfficientTrain, always leverages all the data at any training stage, but only exposes the ‘easier-to-learn’ patterns of each example at the beginning of training, and gradually introduces more difficult patterns as learning progresses. Our method significantly improves the training efficiency of state-of-the-art deep networks on the large-scale ImageNet-1K/22K datasets, for both supervised and self-supervised learning.
This work is supported in part by the National Key R&D Program of China under Grant 2021ZD0140407, the National Natural Science Foundation of China under Grants 62022048 and 62276150, the National Defense Basic Science and Technology Strengthening Program of China, Beijing Academy of Artificial Intelligence (BAAI), and Huawei Technologies Ltd.
Appendix for “EfficientTrain: Exploring Generalized Curriculum Learning for Training Visual Backbones”
Appendix A Implementation Details
Dataset. We use the data provided by ILSVRC2012https://image-net.org/index.php . The dataset includes 1.2 million images for training and 50,000 images for validation, both of which are categorized in 1,000 classes.
Training. Our approach is developed on top of a state-of-the-art training pipeline of deep networks, which incorporates a holistic combination of various model regularization & data augmentation techniques, and is widely applied to train recently proposed models . Our training settings generally follow from , while we modify the configurations of weight decay, stochastic depth and exponential moving average (EMA) according to the recommendation in the original papers of different models (i.e., ConvNeXt , DeiT , PVT , Swin Transformer and CSWin Transformer )The training of ResNet follows the recipe provided in .. The detailed hyper-parameters are summarized in Table 13.
The baselines presented in Table 6 directly use the training configurations in Table 13. Based on Table 13, our proposed EfficientTrain curriculum performs low-frequency cropping and modifies the value of in RandAug during training, as introduced in Table 5. The results in Tables 6 and 9 adopt a varying number of training epochs on top of Table 6.
In addition, the low-frequency cropping operation in EfficientTrain leads to a varying input size during training. Notably, visual backbones can naturally process different sizes of inputs with no or minimal modifications. Specifically, once the input size varies, ResNets and ConvNeXts do not need any change, while vision Transformers (i.e., DeiT, PVT, Swin and CSWin) only need to resize their position bias correspondingly, as suggested in their papers. Our method starts the training with small-size inputs and the reduced computational cost. The input size is switched midway in the training process, where we resize the position bias for ViTs (do nothing for ConvNets). Finally, the learning ends up with full-size inputs, as used at test time. As a consequence, the overall computational/time cost to obtain the final trained models is effectively saved.
Inference. Following , we use a crop ratio of 0.875 and 1.0 for the inference input size of 2242 and 3842, respectively.
A.2 ImageNet-22K Pre-training
Dataset and pre-processing. In our experiments, the officially released processed version of ImageNet-22Khttps://image-net.org/data/imagenet21k_resized.tar.gz is used. The original ImageNet-22K dataset is pre-processed by resizing the images (to reduce the dataset’s memory footprint from 1.3TB to 250GB) and removing a small number of samples. The processed dataset consists of 13M images in 19K classes. Note that this pre-processing procedure is officially recommended and accomplished by the official website.
Pre-training. We pre-train CSWin-Base/Large and ConvNeXt-Base/Large on ImageNet-22K. The pre-training process basically follows the training configurations of ImageNet-1K (i.e., Table 13), except for the differences presented in the following. The number of training epochs is set to 120 with a 5-epoch linear warm-up. For all the four models, the maximum value of the increasing stochastic depth regularization is set to 0.1 . Following , the initial learning rate for CSWin-Base/Large is set to 2e-3, while the weight-decay coefficient for CSWin-Base/Large is set to 0.05/0.1. Following , we do not leverage the exponential moving average (EMA) mechanism. To ensure a fair comparison, we report the results of our implementation for both baselines and EfficientTrain, where they adopt exactly the same training settings (apart from the configurations modified by EfficientTrain itself).
Fine-tuning. We evaluate the ImageNet-22K pre-trained models by fine-tuning them and reporting the corresponding accuracy on ImageNet-1K. The fine-tuning process of ConvNeXt-Base/Large follows their original paper . The fine-tuning of CSWin-Base/Large adopts the same setups as ConvNeXt-Base/Large. We empirically observe that this setting achieves a better performance than the original fine-tuning pipeline of CSWin-Base/Large in .
A.3 Object Detection and Segmentation on COCO
Our implementation of RetinaNet follows from . Our implementation of Cascade Mask-RCNN is the same as .
A.4 Experiments in Section 4
In particular, the experimental results provided in Section 4 are based on the training settings listed in Table 13 as well, expect for the specified modifications (e.g., with the low-passed filtered inputs). The computing of CKA feature similarity follows .
Appendix B Additional Results
The detailed wall-time training cost for the models presented in Table 6 of the paper is reported in Table 14. The numbers of GPU-hours are benchmarked on NVIDIA 3090 GPUs. The batch size for each GPU and the total number of GPUs are configured conditioned on different models, under the principle of saturating all the computational cores of GPUs.
B.2 On the Continuous Selection of B𝐵B
Notably, the basic formulation behind EfficientTrain considers a continuous function of , i.e., . Theoretically, we can obtain an optimal curriculum if we directly solve for a strictly continuous . However, directly solving for a continuous function is computationally intractable. To achieve a reasonable trade-off between the computational cost and the effectiveness of the solution, we adopt an approximation approach, i.e., approximating the continuous function of with a staircase function. Specifically, we divide the training process into stages and solve for a value of for each stage, where we set and obtained the EfficientTrain curriculum.
Importantly, such approximation works reasonably well. As shown in our paper, its solution (EfficientTrain) considerably improves the training efficiency of deep networks, and exhibits superior generalizability across different backbone architectures and various training settings. Besides, as shown in Table 15, further approaching solving for a continuous function of (e.g., ) only yields limited gains.
Appendix C Proof of Proposition 1
In this section, we theoretically demonstrate the difference between two transformations, namely low-frequency cropping and image down-sampling. In specific, we will show that from the perspective of signal processing, the former perfectly preserves the lower-frequency signals within a square region in the frequency domain and discards the rest, while the image obtained from pixel-space down-sampling contains the signals mixed from both lower- and higher- frequencies.
Similarly, the inverse 2D discrete Fourier transform is defined by
Denote the low-frequency cropping operation parametrized by the output size as , which gives outputs by simple cropping:
Note that here , and this operation simply copies the central area of with a scaling ratio. The scaling ratio is a natural term from the change of total energy in the pixels, since the number of pixels shrinks by the ratio of .
C.2 Propositions
Proof. The proof of this proposition is simple and straightforward. Take Fourier transform on both sides of the above transformation equation, we get
Denote the spectral of as and similarly . According to our definition of the cropping operation, we know that
Hence, the spectral information of simply copies ’s low frequency parts and conducts a uniform scaling by dividing .
Proof. Taking Fourier transform on both sides, we have
For any , according to the definition we have
while at the same time we have the inverse DFT for :
Plugging (*2) into (*1), it is easy to see that essentially each is a linear combination of the original signals . Namely, it can be represented as
Therefore, we can compute the dependency weight for any given tuple as
where , same for . Further deduction shows
Denote , which is a constant conditioned on . Then we know
Similar to the proof of Proposition 1.2, by plugging (*2) into (*4), it is easy to see that is a linear combination of the signals from the original image. Namely, we have
Given any , can be computed as
Denote , which is a constant conditioned on . Then we know
Since , we have . Thus, we have when .
Appendix D More Discussions
Potential impacts. The de-facto guarantee for the state-of-the-art performance of modern deep networks (e.g., vision Transformers) incorporates an increasing model size, the large-scale training data, and a sufficiently long training procedure with delicate regularization techniques. However, the establishment of this regime comes at an intensive and unaffordable computational cost for training. Towards this direction, EfficientTrain proposes a simple, easy-to-use, but effective learning approach to reduce the training cost of visual backbones. Our work may benefit real-world applications in terms of accelerating the designing and validating of deep learning architectures or algorithms. Under environmental considerations, it will also help to reduce the carbon emission caused by training large deep learning models. For the research community, EfficientTrain may potentially motivate the researchers to focus on the generalized formulation of curriculum learning.
Limitations and future work. Currently, the EfficientTrain algorithm mainly focuses on training models with images. In the future, we will focus on extending our method to leveraging videos or texts. In addition, it would be interesting to explore whether we can extract the ‘easier-to-learn’ information from the lens of the spatial or temporal redundancy of vision data . We will also focus on exploring facilitating the efficient training of deep networks by leveraging dynamic network architectures .