Vision Transformer for Small-Size Datasets

Seung Hoon Lee, Seunghyun Lee, Byung Cheol Song

INTRODUCTION

Convolutional neural networks (CNNs), which are effective in learning visual representations of image data, have been the main-stream in the field of computer vision (CV) . Meanwhile, in the field of Natural Language Processing (NLP), the so-called Transformer based on self-attention mechanism has achieved tremendous success . So, in the CV field, there have been attempts to combine the self-attention mechanism with CNNs . These studies have succeeded in proving that the self-attention mechanism also works for the image domain. Recently, it was reported that Vision Transformer (ViT) , which applied a standard Transformer composed entirely of self-attention to image data, showed better performance than ResNet and EfficientNet in the image classification task. This made Transformer receive a lot of attention in the CV field.

ViT rarely uses convolutional filters, i.e., the core of CNNs. Convolutional filters were usually used only for their tokenization. Thus, ViT structurally lacks locality inductive bias than CNNs, and they require a too large amount of training data to obtain acceptable visual representation . For example, just to learn a small-size dataset, ViT had to precede pre-training on a large-size dataset such as JFT-300M . In order to alleviate the burden of pre-training, several ViTs which can learn a mid-size dataset such as ImageNet from scratch have been proposed. Such data-efficient ViTs tried to increase the locality inductive bias in terms of network architecture. For example, some adopted a hierarchical structure like CNNs to leverage various receptive fields , and the others tried to modify the self-attention mechanism itself . However, learning from scratch on mid-size datasets still requires significant costs. Moreover, learning small-size datasets from scratch is very challenging considering the trade-off between dataset capacity and performance. Therefore, we need to study ViT that can learn small-size datasets by sufficiently increasing the locality inductive bias.

Through observations, we found two problems that decrease locality inductive bias and limit the performance of the ViT. The first problem is poor tokenization. ViT divides a given image into non-overlapping patches of equal size, and linearly projects each patch to a visual token. Here, the same linear projection is applied to each patch. So, tokenization of the ViT has the permutation invariant property, which enables a good embedding of relations between patches . On the other hand, non-overlapping patches allow visual tokens to have a relatively small receptive field. Usually, tokenization based on non-overlapping patches has a smaller receptive field than tokenization based on overlapping patches with the same down-sampling ratio. Small receptive fields cause ViT to tokenize with too few pixels. As a result, the spatial relationship with adjacent pixels is not sufficiently embedded in each visual token. The second problem is the poor attention mechanism. The feature dimension of image data is far greater than that of natural language and audio signal, so the number of embedded tokens is inevitably large. Thus, the distribution of attention scores of tokens becomes smooth. In other words, we face the problem that ViTs cannot attend locally to important visual tokens. The above two main problems cause highly redundant attentions that cannot focus on a target class. This redundant attention makes it easy for ViT to normally concentrate on the background and not capture the shape of the target class well (see Fig. 5).

This paper presents two solutions to effectively improve the locality inductive bias of ViT for learning small-size datasets from scratch. First, we propose Shifted Patch Tokenization (SPT) to further utilize spatial relations between neighboring pixels in the tokenization process. The idea of SPT was derived from Temporal Shift Module (TSM) . TSM is effective temporal modeling which shifts some temporal channels of features. Inspired by this, we propose effective spatial modeling that tokenizes spatially shifted images together with the input image. SPT can give a wider receptive field to ViT than standard tokenization. This has the effect of increasing the locality inductive bias by embedding more spatial information in each visual token. Second, we propose Locality Self-Attention (LSA), which allows ViT to attend locally. LSA mitigates the smoothing phenomenon of attention score distribution by excluding self-tokens and by applying learnable temperature to the softmax function. LSA induces attention to work locally by forcing each token to focus more on tokens with large relation to itself. Note that the proposed SPT and LSA can be easily applied to various ViTs in the form of add-on modules without structural changes and can effectively improve performance (see Fig. 1 and Table 5).

Our experiments show that the proposed method improves the performance of various ViTs both qualitatively and quantitatively. First, Fig. 5 illustrates that when SPT and LSA are applied to the ViTs, object shapes are better captured. From a quantitative aspect, SPT and LSA improve image classification performance. For example, in the experiment on Tiny-ImageNet, the classification accuracy is improved by an average of 2.962.96%, and a maximum of 4.084.08% (see Table 2). Also, SPT and LSA improve the performance of ViTs up to 1.06%1.06\% in the mid-size dataset such as ImageNet (see Table 3). The main contribution points of this paper are as follows:

To sufficiently embed spatial information between neighboring pixels, we propose new tokenization based on spatial feature shifting. The proposed tokenization can give a wider receptive field to visual tokens. This dramatically improves the performance of the ViTs.

We propose a locality attention mechanism to solve or attenuate the smoothing problem of the attention score distribution. This mechanism significantly improves the performance of ViTs with only a small parameter increase and the addition of simple operations.

RELATED WORK

Recently, several data-efficient ViTs have been proposed to alleviate the dependence of ViT on large-size datasets. These ViTs can learn mid-size datasets from scratch. For example, DeiT improved the efficiency of ViTs by employing data augmentations and regularizations and realized knowledge distillation by introducing the distillation token concept. T2T used a tokenization method that flattened overlapping patches and applied a transformer. This makes it possible to learn local structure information around a token. PiT produced various receptive fields through spatial dimension reduction based on the pooling structure of a convolutional layer. CvT replaced both linear projection and multi-layer perceptron with convolutional layers. Also, like PiT, CvT generated various receptive fields only with a convolutional layer. Swin Transformer presented an efficient hierarchical transformer that gradually reduces the number of tokens through patch merging while using attention calculated in non-overlapping local windows. CaiT employed LayerScale, which converges well even in training ViTs with a large depth. In addition, the transformer layer of the CaiT is divided into a patch-attention layer and a class-attention layer, which is effective for class embedding.

However, the ViT for small-size datasets has not been reported yet. Therefore, this paper proposes tokenization using more spatial information and also a high-performance attention mechanism, which allows ViTs to effectively learn small-size datasets from scratch.

PROPOSED METHOD

This section specifically describes two key ideas for increasing the locality inductive bias of ViTs: SPT and LSA. First, Fig. 2(a) depicts the concept of SPT. SPT spatially shifts an input image in several directions and concatenates them with the input image. Fig. 2(a) is an example of shifting in four diagonal directions. Next, patch partitioning is applied like standard ViTs. Then, for embedding into visual tokens, three processes are sequentially performed: patch flattening, layer normalization , and linear projection. As a result, SPT can embed more spatial information into visual tokens and increase the locality inductive bias of ViTs.

Fig. 2(b) explains the second idea, LSA. In general, a softmax function can control the smoothness of the output distribution through temperature scaling . LSA primarily sharpens the distribution of attention scores by learning the temperature parameters of the softmax function. Additionally, the self-token relation is removed by applying the so-called diagonal masking, which forcibly suppresses the diagonal components of the similarity matrix computed by Query and Key. This masking relatively increases the attention scores between different tokens, making the distribution of attention scores sharper. As a result, LSA increases the locality inductive bias by making ViT’s attention locally focused.

Before a detailed description of the proposed SPT and LSA, this section briefly reviews the tokenization and formulation of the self-attention mechanism of standard ViT .

Next, we obtain patch embeddings by linearly projecting each vector into the space of the hidden dimension of the transformer encoder. Each patch embedding corresponds to a visual token input to the transformer encoder, so this series of processes is called tokenization, i.e., T\mathcal{T}. This is defined by:

Note that the receptive fields of visual tokens in ViT are determined by tokenization. In the transformer encoder running after the tokenization step, the number of visual tokens does not change, so the receptive field cannot be adjusted there. and the tokenization (Eq. 2) of standard ViT is the same as the operation of the non-overlapping convolutional layer with the same size of kernel and stride. So, the receptive field size of visual tokens can be calculated by the following equation given in :

where rtokenr_{token} and rtransr_{trans} stand for the receptive field sizes of tokenization and transformer encoder, respectively. jj and kk are the stride and kernel size of the convolutional layer, respectively. As mentioned earlier, the receptive field is not adjusted in the transformer encoder, so rtrans=1r_{trans}=1. Thus, rtokenr_{token} is the same as the kernel size. Here, the kernel size is the patch size of ViT.

At this time, let’s investigate whether rtokenr_{token} is of sufficient size. For instance, we compare rtokenr_{token} with the receptive field size of the last feature of ResNet50 when training on the ImageNet dataset consisting of images of 224×224224\times 224. The patch size of standard ViT is 16, so rtokenr_{token} of visual tokens is also 16. On the other hand, the receptive field size of the ResNet50 feature amounts to 483 . As a result, the visual tokens of ViTs have a receptive field size that is about 30 times smaller than that of the ResNet50 feature. We interpret this small receptive field of tokenization as a major factor in the lack of local inductive bias. Therefore, Sec. 3.2 proposes the SPT to leverage rich spatial information by increasing the receptive field of tokenization.

2 Shifted Patch Tokenization

This section first describes the overall formulation of SPT (Sec. 3.2.1) and applies the proposed SPT to the patch embedding layer and the pooling layer, i.e., two main tokenizations for ViTs (Sec. 3.2.2 and Sec. 3.2.3).

2.2 Patch Embedding Layer

This section describes how to use SPT as a patch embedding layer. We concatenate a class token to visual tokens and then add positional embedding. Here the class token is the token with representation information of the entire image, and the positional embedding gives positional information to the visual tokens. If a class token is not used, only positional embedding is added to the output of SPT. How to apply the SPT to the patch embedding layer is formulated as follows:

2.3 Pooling Layer

3 Locality Self-Attention Mechanism

This section describes the LSA. The core of LSA is the diagonal masking (Sec. 3.3.1) and the learnable temperature scaling (Sec. 3.3.2).

3.2 Learnable Temperature Scaling

The second technique for LSA is the learnable temperature scaling, which allows ViT to determine the softmax temperature by itself during the learning process. Fig. 3 shows the average learned temperature according to depth when the softmax temperature is used as the learnable parameter in Eq. 5. Note that the average learned temperature is lower than the constant temperature of standard ViT. In general, the low temperature of softmax sharpens the score distribution. Therefore, the learnable temperature scaling sharpens the distribution of attention scores. Based on Eq. 5, the LSA with both diagonal masking and learnable temperature scaling applied is defined by:

EXPERIMENT

This section verifies that the proposed method improves the performance of various ViTs through several experiments. Sec. 4.1 describes the settings of the following experiments. Sec. 4.2 quantitatively shows that the proposed method effectively improves various ViTs and reduces the gap with CNNs. Finally, Sec. 4.3 demonstrates that the ViTs are qualitatively enhanced by visualizing the attention scores of the final class token.

The proposed method was implemented in Pytorch . In the small-size dataset experiment (Table 2), The details of throughput measurement are as follows: The inputs were Tiny-ImageNet, and the batch size was 128, and the GPU was RTX 2080 Ti.

For small-size dataset experiments, CIFAR-10, CIFAR-100 , Tiny-ImageNet , and SVHN were employed and ImageNet was employed for the mid-size dataset experiment.

1.2 Model Configurations

In the small dataset experiment, in the case of ViT, the depth was set to 9, the hidden dimension was set to 192, and the number of heads was set to 12. This configuration was determined experimentally. And in the ImageNet experiment, we used the ViT-Tiny suggested by DeiT . In the case of PiT, T2T, Swin and CaiT, the configurations of PiT-XS, T2T-14, Swin-T and CaiT-XXS24 presented in the corresponding papers were adopted as they were, respectively. The performance of ViT improves as the number of tokens increases, but the computational cost increases quadratically. We were able to experimentally observe that it was effective when both the number of visual tokens in ViT without pooling and the number of tokens in the intermediate stage of ViT with pooling are 64, considering this trade-off. Accordingly, we modified the baseline models. In small-size dataset experiments, the patch size of the patch embedding layer was set to 88 and the patch size of ViTs using pooling layers such as Swin and PiT was set to 1616. In the ImageNet dataset experiment, the patch size was set to be the same as that used in each paper. Also, the hidden dimension of MLP was set to twice that of the transformer in the small dataset experiment, and the configuration used in each paper was applied in the ImageNet experiment.

1.3 Training Regime

According to DeiT, various techniques are required to effectively train ViTs. Thus, we applied data augmentations such as CutMix , Mixup , Auto Augment , Repeated Augment to all models. In addition, regularization techniques such as label smoothing , stochastic depth , and random erasing were employed. Meanwhile, AdamW was used as the optimizer. Weight decays were set to 0.05, batch size to 128 (however, 256 for ImageNet), and warm-up to 10 (however, 5 for ImageNet). All models were trained for 100 epochs, and cosine learning rate decay was used. In the small-size dataset experiments, the initial learning rate of ViT and CNNs was set to 0.003, and that of the remaining models was set to 0.001. On the other hand, in the ImageNet experiment, the initial learning rate was set to 0.00025 for all models.

2 QUANTITATIVE RESULT

This section presents the experimental results for small-size datasets and the ImageNet dataset. In the small-size dataset experiment, Throughput, FLOPs, and the number of parameters were measured in Tiny-ImageNet.

Table 3 shows performance when training a mid-size dataset ImageNet from scratch. In ViT, SPT was applied only to patch embeddings, and in PiT and Swin, SPT was applied to both patch embedding and pooling layers. We could observe that the proposed method is sufficiently effective for ImageNet. For example, the performance was improved by the proposed method as much as +1.60%+1.60\% for ViT, +1.44%+1.44\% for PiT, and +1.06%+1.06\% for Swin. As a result, we find that the proposed method noticeably improves the ViTs even on mid-size datasets.

2.2 Ablation Study

This section describes the ablation study on the proposed method. ViT was used for this experiment.

Let’s look at the effect of learnable temperature scaling and diagonal masking, two key elements of LSA, on overall performance. Table 4 shows that learnable temperature scaling and diagonal masking effectively resolves the smoothing phenomenon of attention score distribution (see Fig. 4). For example, learnable temperature scaling and diagonal masking in Tiny-ImageNet improved performance by +0.88%+0.88\% and +1.22%+1.22\%, respectively. Considering that the LSA applied with both techniques shows a performance improvement of +1.43%+1.43\%, we can claim that the contribution of each is sufficiently large and the two techniques produce a synergy.

Table 5 shows that SPT and LSA can dramatically improve performance by increasing the locality inductive bias of ViT independently. In particular, in Tiny-ImageNet, SPT and LSA improved performance by +1.43%+1.43\% and +3.60%+3.60\%, respectively. When both techniques were applied, the performance improvement was +4.00%+4.00\%. This proves the competitiveness and synergy of the two key element technologies.

3 QUALITATIVE RESULT

Fig. 5 visualizes the attention scores of the final class token when SPT and LSA were applied to various ViTs. When the proposed method was applied, we can observe that the object shape is better captured as the attention, which was dispersed in the background, is concentrated on the target class. In particular, this phenomenon is evident in the CaiT of the first row, the T2T of the second row, the ViT of the third row, and the PiT of the last row. Therefore, we can find that the proposed method effectively increases the locality inductive bias and induces the attention of the ViTs to improve.

CONCLUSION

To train ViT on small-size datasets, this paper presents two novel techniques to increase the locality inductive bias of ViT. First, SPT embeds rich spatial information into visual tokens through specific transformation. Second, LSA induces ViT to attend locally through softmax with learnable parameters. The SPT and LSA can achieve significant performance improvement independently, and they are applicable to any ViTs. Therefore, this study proves that ViT learns small-size datasets from scratch and provides an opportunity for ViT to develop further.

References

Supplementary

This section investigates the various shifting strategies that SPT can employ. Specifically, we explored the shift direction and shift intensity (shift ratio), which have the most impact on performance.

We examined the following three shift directions. The first is the 4 cardinal directions consisting of up, down, left and right directions (Fig. 1(a)). The second is 4 diagonal directions including up-left, up-right, down-left and down-right (Fig. 1(b)). The last is the 8 cardinal directions including all the preceding directions (Fig. 1(c)). Table 1 shows top-1 accuracy in small-size datasets such as CIFAR-10, CIFAR-100, SVHN, and Tiny-ImageNet for each shift direction. This experiment adopted a model applying SPT to standard ViT. 4 cardinal directions showed the best performance in CIFAR-10 and SVHN. On the other hand, 4 diagonal directions and 8 cardinal directions provided the best performance in CIFAR-100 and Tiny-ImageNet, respectively. This shows that the shift direction is somewhat dependent on the characteristics of datasets. For example, in CIFAR-10 or CIFAR-100, the target class tends to be in the center of the image, whereas other datasets do not. The location of the target class has some degree of correlation with the shift direction, and the correlation can affect the performance. However, since the performance difference was experimentally marginal, in this paper, the shift direction in the experiment was fixed to 4 diagonal directions.