Fixing the train-test resolution discrepancy: FixEfficientNet

Hugo Touvron, Andrea Vedaldi, Matthijs Douze, Hervé Jégou

Introduction

In order to obtain the best possible performance from Convolutional neural nets (CNNs), the training and testing data distributions should match. However, in image recognition, data pre-processing procedures are often different for training and testing: the most popular practice is to extract a rectangle with random coordinates from the image to artificially increase the amount of training data. This Region of Classification (RoC) is then resized to obtain an image, or crop, of a fixed size (in pixels) that is fed to the CNN. At test time, the RoC is instead set to a square covering the central part of the image, which results in the extraction of a center crop. Thus, while the crops extracted at training and test time have the same size, they arise from different RoCs, which skews the data distribution seen by the CNN.

Over the years, training and testing pre-processing procedures have evolved, but so far they have been optimized separately . Touvron et al. show that this separate optimization has a detrimental effect on the test-time performance of models. They address this problem with the FixRes method, which jointly optimizes the choice of resolutions and scales at training and test time, while keeping the same RoC sampling.

We apply this method to the recent EfficientNet architecture, which offers an excellent compromise between number of parameters and accuracy. This evaluation paper shows that properly combining FixRes and EfficientNet further improves the state of the art . Noticeably,

We report the best performance without external data on ImageNet (top1: 85.7%);

We report the best accuracy (top1: 88.5%) with external data on ImageNet, and with ImageNet with Reallabels ;

We achieve state-of-the-art compromises between accuracy and number of parameters, see Figure 1;

We validate the significance of our results on the ImageNet-v2 test set, an improved evaluation setup that clearly separates the validation and test sets. FixEfficientNet achieves the best performance.

This paper is organized as follows. In Section 2 we introduce the corrected training procedure for EfficientNet, that produces FixEfficientNet. Section 3 analyzes our extensive evaluation and compare FixEfficientNet with the state of the art. Section 4 concludes the paper.

Training with FixRes: updates

Recent research in image classification tends towards larger networks and higher resolution images . For instance, the state-of-the-art in the ImageNet ILSVRC 2012 benchmark is currently held by the EfficientNet-L2 architecture with 480M parameters using 800×\times800 images for training. Similarly, the state-of-the-art model learned from scratch is currently EfficientNet-B8 with 88M parameters using 672×\times672 images for training. In this note, we focus on the EfficientNet architecture due to its good accuracy/cost trade-off and its popularity.

Data augmentation is routinely employed at training time to improve model generalization and reduce overfitting. In this note, we use the same augmentation setup as in the original FixRes paper . In addition, we have integrated label smoothing, which is orthogonal to the approach. FixRes is a very simple fine-tuning that re-trains the classifier or a few top layers at the target resolution. Therefore, it has several advantages:

it is computationally cheap, the back-propagation is not performed on the whole network;

it works with any CNN classification architecture and is complementary with the other tricks mentioned above;

it can be applied on a CNN that comes from a possibly non reproducible source.

Experiments

We experiment on the ImageNet-2012 benchmark , and report standard performance metrics (top-1 and top-5 accuracies) on a single image crop.

We focus on the EfficientNet architectures. In the literature, wo versions provide the best performance: EfficientNet trained with adversarial examples , and EfficientNet trained with Noisy student pre-trained in a weakly-supervised fashion on 300 million unlabeled images.

We start from the EfficientNet models in rwightman’s GitHub repository . These models have been converted from the original Tensorflow to PyTorch.

Training. We mostly follow the FixRes training protocol. The only difference is that we combine the FixRes data-augmentation with label smoothing during the fine-tuning.

2 Comparison with the state of the art

Table 1 and Table 2 compare our results with those of the EfficientNet reported in the literature. All our FixEfficientNets outperform the corresponding EfficientNet (see Figure 1). As a result and to the best of our knowledge, our FixEfficientNet-L2 surpasses all other results reported in the literature. It achieves 88.5% Top-1 accuracy and 98.7% Top-5 accuracy on the ImageNet-2012 validation benchmark .

Clean labels. In order to complement this evaluation, Table 3 present the results with the ImageNet clean labels proposed by Beyer et all. . With 90.9% Top-1 accuracy and 98.8% Top-5 accuracy FixEfficientNet-L2 surpasses all other results reported in the literature with this labels.

3 Significance of the results

Several runs of the same training incur variations of about 0.1 accuracy points on Imagenet due to random initialization and mini-batch sampling. In general, since the Imagenet 2012 test set is not available, most works tune the hyper-parameters on the validation set, ie. there is no distinction between validation and test set. This setting, while widely adopted, is not legitimate and can cause overfitting to go unnoticed.

EfficientNets employ Neural Architecture Search, which significantly enlarges the hyper-parameter space. Additionally, the ImageNet validation images were used to filter the images from the unlabelled set . Therefore the pre-trained models may benefit from more overfitting on the validation set. We quantify this in the experiments presented below.

Since we use pre-trained EfficientNet for our initialization, our results are comparable to those from the Noisy Student , which uses the same degree of overfitting, but not directly with other semi-supervised approaches like that of Yalniz et al. .

4 Evaluation on ImageNet-V2

The ImageNet-V2 dataset was introduced to overcome the lack of a test split in the Imagenet dataset. ImageNet-V2 consists of 3 novel test sets that replace the ImageNet test set, which is no longer available. They were carefully designed to match the characteristics of the original test set. One of these test sets, Matched Frequency is the closest to the ImageNet validation set. To ensure that observed improvements are not due to overfitting, we evaluate all our models on the Matched Frequency version of the ImageNet-v2 dataset. We evaluate the other methods in the same way. We present the results in Tables 4 and 5.

The original study of shows that there is significant overfitting of various models to the Imagenet 2012 valuation set, but that it does not impact the relative order of the models.

Quantifying the overfitting on Imagenet. As mentioned earlier, several choices in the Noisy Student method are prone to overfitting. We verify this hypothesis and quantify its extent by comparing the relative accuracy of this approach with another semi-supervised approach both on ImageNet and ImageNet-V2 .

Without overfitting, models performing similarly on Imagenet should also have similar performances on ImageNet-V2 . However, for a comparable performance on ImageNet, when evaluating on ImageNet-V2, the Billion scale models of Yalniz et al. outperform the EfficientNets from Noisy Student. For example, FixResNeXt-101 32x4d has the same performance as EfficientNet-B3 on ImageNet but on ImageNet-V2 FixResNeXt-101 32x4d is better (+0.7% Top-1 accuracy).

This shows that the EfficientNet Noisy student tends to overfit and does not generalize as well as the (prior) semi-supervised work or other works of the literature. Figure 2 illustrates this effect. The FixRes fine-tuning procedure is neutral with respect to overfitting: overfitted models remain overfitted and conversely.

Comparison with the state of the art. Despite overfitting, EfficientNet remains very competitive on ImageNet-V2, as reported in Table 6. Interestingly, the FixEfficientNet-L2 that we fine-tuned from EfficientNet establishes the new state of the art with additional data on this benchmark.

Conclusion

The ”Fixing Resolution” is a method that improves the performance of any model. It is a method that is applied as a fine-tuning step after the conventional training, during a few epochs only, which makes it very flexible. It is easily integrated into any existing training pipeline. In our paper we proposed a thorough evaluation of the combination of the current state-of-the-art models, namely EfficientNet, with this improved training method.

We provide an open-source implementation of our method http://github.com/facebookresearch/FixRes.

References