FlexiViT: One Model for All Patch Sizes
Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov, Mathilde Caron, Simon Kornblith, Xiaohua Zhai, Matthias Minderer, Michael Tschannen, Ibrahim Alabdulmohsin, Filip Pavetic
Introduction
Vision Transformers (ViTs) cut images into non-overlapping patches and perform all computations on tokens created from these patches. This “patchification” procedure represents a significant shift away from the previously dominant convolutional neural network (CNN) approach , where an image is processed with small local and typically overlapping filters. Patchification has unlocked new capabilities, such as (random) dropping of image patch tokens , adding specialized tokens for new tasks or mixing image tokens with tokens from other modalities .
Despite the importance of patchification for ViT models, the role of the patch size has received little attention. While the original ViT paper works with three patch sizes (3232, 1616, and 1414 pixels), many follow-up works fix the patch size at 1616 pixels . In this work, we show that the patch size provides a simple and effective lever to change the compute and predictive performance of a model, without changing model parametrization. For example, a ViT-B/8 model achieves top-1 accuracy on ImageNet1k with 156 GFLOPs and 85 M parameters, while a ViT-B/32 model achieves only accuracy with 8.6 GFLOPs and 87 M parameters. Despite the major difference in performance and compute, these models have essentially the same parametrization. However, standard ViT models perform well only at the patch size that they have been trained at. Tuning the patch size therefore requires complete re-training of the model.
To overcome this limitation, we propose FlexiViT, a flexible ViT which matches or outperforms standard fixed-patch ViTs across a wide range of patch sizes with no added cost. To train FlexiViT, we randomize the patch size during training, and resize the positional and patch embedding parameters adaptively for each patch size, as shown in Figure 1. These simple modifications are already sufficient for strong performance, but we also propose a optimized resizing operation and a training procedure based on knowledge distillation which achieves even better results.
We demonstrate the efficiency of FlexiViT models in many downstream tasks, such as image classification, transfer learning, panoptic and semantic segmentation, image-text retrieval and open-world recognition, and provide a general recipe for flexifying existing ViT-based training setups. Furthermore, we show that flexibility of the backbone, i.e. strong performance across patch sizes, is often preserved even after fine-tuning with a fixed patch size. We leverage this observation to perform resource-efficient transfer learning: we finetune the model cheaply with a large patch size, but then deploy it with a small patch size for strong downstream performance. We further show that flexible patch size can be used to accelerate pre-training.
To explain the effectiveness of FlexiViT, we analyze the model’s representations. We find that the representations are often similar across different patch sizes, especially in the deeper layers. Finally, we show that FlexiViT outperforms alternative architectural ways of controlling the performance-compute trade-off in ViT models.
Related work
Several recent works explore improving ViT’s efficiency by exploiting patchification. Some suggest removing tokens, either in randomized or structured fashion throughout training. Others aim to quantify a token’s importance and remove the least important ones, during or after training. trained a cascade of Transformers using increasing number of tokens to allow early exiting during inference. Conversely, we always keep all tokens and do not discard any information. It may be possible to combine such approaches with FlexiViT in future work.
Another related line of work looks at changing input resolution during training, typically in order to speed-up pre-training or as data augmentation for self-supervised learning of ViTs . We do not explore the data augmentation inpact of Flexi training to avoid an explosion in scope. The aforementioned models all work only at their single, final resolution, while our ai is to have one model that works well at all trained patch-sizes.
More similar to our approach, the Neural Architecture Search (NAS) field is converging towards training one “supernet” from which individual, differently-shaped “subnets” can be extracted . Since these works aim for changes in most or all model dimensions, they usually involve multiple specialized architectural additions. SuperViT is most related to FlexiViT as it patchifies an image at multiple scales, passes all these patches to ViT, while dropping random tokens to reduce the sequence length. In contrast to the aforementioned works, our sharpened focus on ViT’s patch size only, allows benefiting from existing pretrained models, future ViT improvements, and is an easy drop-in to any existing training procedure.
Matryoshka representation learning proposes training models whose output vector contains meaningful sub-vectors. This can be seen as the complement of FlexiViT.
Making ViT flexible
In this section we show that standard ViT models are not flexible, and introduce the FlexiViT model and training procedure in the supervised image classification setting. We perform all experiments in this section on the public ImageNet-21k dataset . We use the base (ViT-B) scale model and unregularized light2 setting from , and train the models for 90 epochs following .
FlexiViT is based on the Vision Transformer (ViT) architecture . Here, we briefly describe the ViT architecture and introduce the necessary notation.
In summary, for a given image size , the patch size determines the length of the input sequence to the Transformer model: smaller patch sizes correspond to longer input sequences and slower, more expressive models. Following , we denote ViT models as ViT-/, where is the model scale (small, medium, base, large, …) and is the patch size. Note that there are only two parts of the model where the parameter vectors depend on the patch size: the patch embedding weights and the position embedding . In the following sections, we will develop a flexible ViT model which works simultaneously for any patch size.
2 Standard ViTs are not flexible
We first show that evaluating a standard pre-trained ViT model at different patch sizes yields poor performance. In order to change the patch size, we simply resize the patch embedding weights and the position embeddings with bilinear interpolation. tf.image.resize(input, res, method=’bilinear’). For the position embeddings, this resize approach was already proposed in the original ViT paper to fine-tune at higher resolution.
The result is shown in Figure 3, where we see that the performance of standard ViT models (dashed and dotted lines) rapidly degrades as the inference-time patch size departs from the one used during training.
3 Training flexible ViTs
In Figure 3 we also show the performance of our FlexiViT-B model (solid line), which matches both ViT-B/16 and ViT-B/30 when evaluated at their training patch sizes, and significantly outperforms them for all other patch sizes. This model was trained in the same setting as the ViT-B/16 and ViT-B/30 models, except that at each step of training, the patch size was chosen uniformly at random from a set of pre-defined patch sizes.We sample patch sizes uniformly in most experiments. Some early runs used a distribution which slightly favors intermediate patch sizes. Later experiments showed that the distribution makes little difference (Appendix C). We therefore did not re-run the early experiments. In order to do so, two small changes to the model and training code are necessary.
First, the model needs to define an underlying parameter shape for and . The learnable parameters are of that shape, and resized on-the-fly as part of the model’s forward pass. We show in Appendix B that the exact shape of these underlying learnable parameters does not matter much, and we use an underlying size of for patches and for position embeddings in all experiments.
Second, to have a large variety of patch sizes that perfectly tile the image, we use an image resolution of 240² px, which allows for patch sizes 240, 120, 60, 48, 40, 30, 24, 20, 16, 15, 12, 10, 8, 6, 5, 4, 2, 1, of which we use all between 48 and 8, inclusive.Perfect tiling may not be strictly necessary, and it may be fine to use arbitrary patch sizes and ignore a small border of the image. For simplicity, we focus on the perfect tiling setting. At each iteration we sample from the uniform distribution over these patch sizes.
These are all the changes necessary to flexify an existing ViT training procedure. Algorithm 1 summarizes them.
Note that changing the patch size is related to, but not identical to, changing the image size. The patch size is purely a change to the model while changing the image size may drastically reduce the available information. This distinction is further explored in Section 3.4.
We explore two alternative ways to flexify ViTs in Section 7: flexible depth and flexible patch stride. Both of them have merits, but patch size works best.
4 How to resize patch embeddings
One way to achieve this equality is to normalize the tokens right after their embedding, either explicitly or by using a LayerNorm module. However, this approach requires changing the model architecture and is not compatible with existing pre-trained ViTs. Further, it does not exactly preserve the patch embeddings. As we will show, there is a more principled way of achieving this goal, which is compatible with existing pre-trained models and does not require any architectural change.
First, we note that the linear resize operation introduced in Section 3.2 can be represented by a linear transformation:
Intuitively, we would like to find a new set of patch-embedding weights such that the tokens of the resized patch match the tokens of the original patch. Formally, we want to solve the optimization problem:
where and is some distribution over the patches. In case when we are increasing the patch size, i.e. , we can use where is the pseudoinverse of :
This way we match the patch embeddings exactly for all .
In the case of downsampling, i.e. when , the solution to the optimization problem in Eq. (2) will in general depend on the patch distribution . In Appendix A.2, we show that for , we recover the pseudoinverse as the optimal solution We can also target the patch distribution in the data in place of , producing a resize operation which depends on the data. In our preliminary experiments, we did not observe significant benefits from this approach. . To sum up, we define PI-resize (pseudoinverse resize) as:
To experimentally validate the effectiveness of PI-resize and compare it to several alternative heuristics, including standard linear resize, we load a pre-trained ViT-B/8 model from and evaluate it after resizing both the image and the model, thus preserving its sequence length . The results, shown in Figure 4, demonstrate that PI-resize maintains nearly constant performance when upsampled, and degrades gracefully when downsampling. None of the heuristics works as well as thoughtful PI-resize across the board.
For completeness, in Appendix A.1 we experimentally compare the remaining ways of dealing with variable patch sizes when one does not care about maintaining model compatibility. These methods include fixed normalization, LayerNorm, and learning separate parameters for each patch size. Adding a LayerNorm works best, but otherwise, PI-resize and bilinear resize are among the best techniques.
5 Connection to knowledge distillation
Knowledge distillation is a popular technique, where a typically smaller student model is trained to mimic the predictions of a typically larger teacher model. This can significantly improve the performance of the student model compared to standard label-supervised training .
It was recently shown that knowledge distillation corresponds to a much more challenging optimization problem than standard supervised training , and that initializing the student close to the teacher simplifies alleviates this . Unfortunately, this solution is impractical since the teacher usually has a different (larger) architecture than the student . However, with FlexiViT, we can initialize a student FlexiViT with the weights of a powerful ViT teacher and significantly improve distillation performance.
Unless otherwise stated, the model we use for the remaining experiments in this paper is a FlexiViT-B initialized and distilled from the powerful ViT-B/8 model of . At initialization, we PI-resize the teacher’s patch embedding weights to , and bilinearly resample its position embeddings to . We then train the student model following the FunMatch approach, minimizing the KL-divergence between the predictions of the teacher and the student FlexiViT with a randomized patch size:
where is the distribution over classes for the FlexiViT model on an input with patch size , is the predictive distribution of the teacher on the exact same input, is the training data distribution with random flips, crops, and mixup, and is the distribution over patch sizes used for training the FlexiViT model.
Figure 5 compares the effect of distilling using teacher initialization to random initialization and to supervised training from labels. The comparison was performed for 90 epochs and shows considerable benefits of this unique initialization capability of FlexiViT. Since distillation needs patience , we additionally run for 300 and 1000 epochs, shown as pale green curves in the figure. FlexiViT matches the teacher’s performance at small patch sizes, and teacher initialization provide a large improvement in accuracy at the largest patch sizes. In the following sections, we use the FlexiViT that was trained for 300 epochs and train two fixed ViT-B/30 and ViT-B/16 models in the same setting (including the initialization) as baselines.
6 FlexiViT’s internal representation
Does FlexiViT process inputs with different patch sizes in similar ways? We investigate this by analyzing the model’s internal representations. We apply minibatch centered kernel alignment (CKA) , a widely-used approach for comparing representations within and across neural networks. For visualization purposes, we apply an arccosine transform to transform CKA/cosine similarity to proper metrics and then perform t-SNE.
Results are shown in Figure 6. Feature map representations are similar across grid sizes from the first layer until the MLP sublayer of block 6. At the MLP sublayer of block 6, layer representations diverge, before converging again at the final block. By contrast, CLS token representations remain aligned across grid sizes. Thus, although internal representations of a substantial portion of FlexiViT differ by grid size, output representations are generally aligned.
Using pre-trained FlexiViTs
We have shown that ViTs can be trained flexibly without significant loss of upstream performance. Next, we verify that pre-trained FlexiViTs are still comparable to individual fixed patch-size ViTs when transferred to other tasks. We check this by transferring the single pre-trained FlexiViT with its patch size fixed to either or to during transfer. We compare FlexiViT to ViT-B/16 and a ViT-B/30 models that were pre-trained using the same distillation setup as FlexiViT (Section 3.5), but with a fixed patch size. We perform this transfer on the following set of diverse tasks.
For each task, we provide more details along with many more results, all with the same take-away, in Appendix E.
Classification We fine-tune on small- (Pet , Flowers ) and medium-scale (CIFAR10, CIFAR100 , Food101 , SUN397 ) classification datasets following the setup of at px resolution.
Locked-image Tuning (LiT) We follow to train a text model contrastively for the frozen FlexiViT, which we evaluate in terms of 0-shot classification and retrieval.
Open-vocabulary detection We test the transferability of FlexiViT to object detection using OWL-ViT , an open-vocabulary object detector based on image-text models such as LiT or CLIP . We evaluate its zero-shot open-vocabulary detection performance on LVIS .
Panoptic segmentation The Universal Vision Model (UViM) is a general-purpose modeling approach for vision . We train UViM on the COCO panoptic segmentation dataset and use FlexiViT as initialization for the image encoder in UViM.
Semantic segmentation We transfer to semantic segmentation following Segmenter’s linear decoder setup . We report mean IoU for single scale evaluation and evaluate on Cityscapes and ADE-20k .
The results of these transfer experiments are summarized in Figure 7. Across the diverse set of tasks, a single FlexiViT model roughly matches the two fixed ViT models, barely lagging behind at large patch size and leading to a small or significant improvement at smaller patch size.
These results confirm that there is no significant downside in using a pre-trained FlexiViT, as opposed to pre-training multiple ViTs for different patch sizes.
2 Resource-efficient transfer via flexibility
FlexiViT enables a new way of making transfer learning more resource efficient, saving accelerator memory and compute. This is possible because, surprisingly, flexibility is largely retained even after transfer at a fixed patch size. We can therefore perform transfer training cheaply with large input patches (small input grid), but later deploy the resulting model using small patch sizes (large input grid). We preform experiments by transferring a FlexiViT-B model (pretrained on ImageNet-21k with distillation) to the ImageNet-1k dataset, and use a similarly pretrained fixed ViT-B/30 model as the baseline. The pretrained FlexiViT works well at larger grid sizes even after fixed-size transfer. For example, we can perform relatively cheap finetuning at grid size. When evaluated at grid size, the model achieves 81.8% accuracy, but when evaluated at the grid size, it achieves 85.3% top-1 accuracy gaining 3.5% accuracy at no additional training cost (Figure 8). More details on the finetuning setup can be found in the Appendix D.
Flexifying existing training setups
So far, we have focused on flexifying models during pre-training. We now show that existing pre-trained models can also be flexified during transfer to downstream tasks. Below, we flexify a diverse set of existing training setups.
We use the same set of 6 transfer datasets from Section 4, with the same settings. We again show the results for SUN397 in Figure 12 and all other datasets in Appendix E. The difference is that we now also randomize the patch size during transfer, and evaluate the single resulting model at different patch sizes (x-axis, three groups of bars).
Flexible transfer of FlexiViT works best, but flexifying a fixed model during transfer also works surprisingly well, considering the very short training and low learning rate used for transfer. The baseline of a fixed-size model transferred at a fixed patch size and evaluated at that same size is indicated by a small horizontal line.
2 Multimodal image-text training
Next, we discuss two ways to flexify multimodal image-text training: FlexiLiT and FlexiCLIP. In FlexiLiT, we train a text tower to produce text embeddings that align well with visual embeddings from various patch sizes (B/flexi). LiT baselines with direct use of either FlexiViT models at fixed resolutions, or ViT models are provided. Figure 12 shows zero-shot image to text retrieval results on the Flickr30k dataset. FlexiLiT-B/flexi performs the best on average, while LiT with FlexiViT-B/30 and FlexiViT-B/16 both get very close results. Flexification additionally provides the possibility of fast transfer as discussed in Section 4.2. The LiT-ViT baselines shown in Figure 12 match FlexiLiT on the sequence length it has been trained for, but performance drops quickly when using a different sequence length during inference. We observe similar conclusions with a from-scratch image-text training setup, i.e. FlexiCLIP (see Appendix G for more results).
3 Open-vocabulary detection
Beyond image-level tasks, we find that flexification also works for object detection training. We modify the training of OWL-ViT to introduce flexible patch sizes as described in Algorithm 1. Similar to classification, flexible OWL-ViT detection models perform close to or better than fixed-size models at any patch size during inference (Figure 12). In addition, we find that for detection, the optimal patch size is not necessarily the smallest. When evaluated on a set of 35 detection datasets , inference-time tuning of the patch size leads to improved results over evaluation at the smallest patch size (Appendix E). This makes flexification especially valuable for detection.
4 Training times and flexification
Analyzing FlexiViTs
Attention relevance patterns across scales We find that decreasing the patch size results in attention relevance to concentrate into a larger number of smaller areas throughout the image. In Figure 13 (top) we observe that attention can significantly change at different scales.
Relation of token representations across scales As we decrease FlexiViT’s patch size, each token “splits” into multiple tokens. A natural question is how token representations at larger patch sizes relate to token representations at smaller patch size. To answer this question, we measure cosine similarity between the representation of a “seed” token at the center of a feature map at one patch size and representations of other tokens at the same and different patch sizes. As shown in Figure 13 (bottom), we are indeed able to find correspondences between tokens across scales.
Ensembling We explored whether it is possible to improve prediction accuracy by ensembling the predictions of the same FlexiViT at multiple scales. We find that, in terms of total compute spent, it is nearly always better to run a single FlexiViT at that compute budget than to ensemble multiple smaller ones. Full results are provided in Appendix J.
Shape or texture bias ViT’s bias towards using shape or texture features has been shown to largely depend on its patch size . In Appendix K, we show that FlexiViT evaluated at each patch size has a similar texture bias to a ViT trained and evaluated at that same patch size.
Model and dataset size Throughout the paper, we focus on FlexiViT models of the base size (-B) trained on 12 M images. In order to validate that neither of these two settings are required, we train FlexiViT-S,B,L models on ImageNet-1k (1.2 M images) using the ImageNet-1k DeiT III model as teacher. We can see in Fig 2 that a single FlexiViT-L model matches or outperforms all three DeiT III models and EfficientNetV2. However, there is still a point at which it becomes more effective to change model width than patch size. Numerical results and evaluation on ImageNet-ReaL/v2/A/R are in Appendix F, a version of Fig. 2 using GFLOPs is in Appendix P.
Discussion of alternatives
Changing the input patch size is not the only way to trade off sequence length and compute in ViTs. We explore two alternatives in our core setup: distillation on ImageNet-21k.
Varying patch embedding stride One alternative is to fix the patch size and change its sampling stride, i.e. extract overlapping patches to increase sequence length. Intuitively, the advantage of this approach is that the intrinsic patch size is fixed and we avoid any special care when computing patch embeddings. Results in Figure 14 suggest varying the stride works almost as well, only slightly lagging behind our baseline.
Varying model depth Another alternative is adding flexibility in terms of depth, i.e. number of layers. Depth pruning has been explored in the context of NLP and more recently also for ViTs . Depth pruning differs fundamentally from FlexiViT: it scales linearly in depth, uses a subset of parameters, and allows progresively refining a prediction. We randomize the depth by attaching the shared head to various intermediate layers. We also tried separate heads, which worked worse. The results in Figure 14 show that FlexiViT provides a significantly better compute-accuracy tradeoff than depth pruning.
Conclusion
FlexiViT is a simple and efficient way of trading off compute and predictive performance with a single model, enabled by the unique patch embedding strategy of ViTs. FlexiViT can be used to significantly reduce pre-training costs by only training a single model for all scales at once, and performs well at a variety of downstream tasks. There are many exciting directions for future work, and we hope that our results inspire the community to explore additional creative applications of patchification.
Acknowledgments
We thank Daniel Keysers for good feedback on a draft of the paper, Geoffrey Hinton for a nudge to pursue this project early on, and our respective teams at Google for encouraging creative and independent research.
We would also like to thank the following artists for making their photographs (used for visualizations) freely available through unsplash.com: Tanya Patrikeyeva, Markus Spiske, Matheus Bardemaker, Alexandru Sofronie, Chris Smith, Kajetan Sumila, Julee Juu, Mike Erskine, Piermanuele Sberni, Feyza Yıldırım and Sixteen Miles Out.
We thank the anonymous reviewers for good feedback that further improves our paper.
Finally, the raccoon picture used as background in Figure 31 is from publicdomainvectors.org.
References
Appendix A More details on flexible patch-sizes
In this section, we further elaborate on many details of flexible patch-sizes. We provide results for alternative ways of dealing with flexible patch-sizes when one does not care about preserving model architecture in Appendix A.1. We provide a detailed derivation of PI-resize in Appendix A.2, and show some PI-resize matrices in Appendix A.3. We further show some visualizations of patch-embedding weights, both raw and resized, in Appendix A.4
Besides bilinear resizing (called Vanilla) or PI-resizing the patch-embedding weights to deal with variable patch-sizes, there are a few other alternatives which we discuss and compare here.
Untied weights for each size, i.e. having separate trainable parameter buffers for each patch-size.
Untied, =init is the same as above, but initializing all patch embedding weights to the same (PI-resized) values as a reference initialization. In this setting, the model is still compatible with standard ViT models at initialization time, however, the parameters can, and do, diverge during training, resulting in a non-standard ViT architecture.
Normalize simply l2-normalizes the tokens computed by the patch-embedding individually to unit-norm. This solves any norm-related issues in a simple, parameter-free way, but is incompatible with pre-trained standard ViT models.
Token-LN and Image-LN add a LayerNorm right after the patch-embedding and differ only in which axis they perform the normalization. Again, this solves the norm-related issues, but does add learnable parameters and is incompatible with pre-trained standard ViT models.
A comparison of all these variants is performed in the label-supervised training setup on ImageNet-21k following , but training for 90 epochs. The result, presented in Figure 15, indicates that plain resizing and PI-resizing are among the best solutions, but have the added benefit of resulting in standard ViT models. Furthermore, not visible in this figure, both Untied variants displayed slightly unstable training curves in the first half of training, while all other variants train smoothly.
A.2 PI-resize derivation
We can rewrite the objective function in Eq. (2) as follows:
Finally, we note that the pseudoinverse matrix recovers a least squares solution to a linear system of equations:
A.3 Visualization of some PI-resize matrices
We visualize an upscaling and a downscaling matrix for both bilinear and PI-resize operations for a visual comparison in Figure 16.
A.4 Visualization of patch-embedding weights
PCA of patch embeddings (see ) in Figs. 33, 35, 35, 36 and 37.
Appendix B The “underlying” parameter shapes
FlexiViT does introduce two new hyper-parameters which were not present in the original ViT architecture: the size of the underlying patch-embedding weights and position-embeddings (i.e. learned params, before resize).
However, both of these parameters have (maybe surprisingly) little influence on the final performance of the model, as long as they are in a “reasonable” range. We verified these in two different settings.
For the patch-size parameter, since it is affected by the resize method used, we perform the ablation in the full FlexiViT setup. Figure 17 (b) shows that there is no notable difference across all evaluation sizes, and hence we stick to the (initially arbitrary) default of 32 across all experiments.
For the position embedding parameter, we ran an early experiment with ViT-S/16 trained from scratch on ImageNet-1k following . Figure 17 (c) shows that in general, even for plain ViT training, this approach could be taken and training curves are mostly unaffected by in-graph bilinear resizing of position embeddings.
Appendix C Distribution of patch-size sampling
During roughly the first half of this project, we sampled patch-size from a distribution which is not uniform, but samples patch-sizes between 16 and 30 up to three times more than patch-sizes outside this range. This “triangular” distribution was based purely on gut-feeling, and once we verified that it is no better than uniform sampling (the experiment is shown in Figure 17 (a)), we decided to use a uniform distribution for simplicity. We further decided to avoid the costly re-running of all experiments we did so far, and thus some experiments in the main paper were done with the “triangular” distribution. However, in all direct comparisons presented throughout the paper, the curves being directly compared always were trained in the exact same way.
Appendix D Detais on resource-efficient transfer
For finetuning FlexiViT models on the Imagenet-1k dataset we generally follow the transfer learning setup from . We use SGD momentum optimizer, with the initial learning rate of and cosine learning rate decay. We also reduce the learning rate for the pretrained parameters by a factor of 10. We optimize for steps with inception crop and flip left-right augmentation, using batch size 512 and input image size of .
Appendix E Using pre-trained FlexiViT models: more details and results
In this section, we provide more details and full results for the scenario described in Section 4: using pre-trained FlexiViT models.
For more details on flexified training procedures discussed in Section 5, we redirect to Appendix M for flexified transfer learning, Appendix G for flexified contrastive image-text learning (LiT and CLIP), and Appendix N for flexified open-vocabulary detection (OWL-ViT), including results on ELEVATER.
For transfer, we follow the simple BiT-HyperRule . In short, we transfer for a relatively brief number of steps (500 for flowers and pets, 2500 for food101 and sun, and 10 000 for CIFAR), using the SGD optimizer with a momentum of , no weight decay, no dropout, and no other augmentations besides flips and random crops. We initialize the new classification layer to all-zeros and use a short learning-rate warmup, both of these having as effect to better preserve the pre-trained weights. The only setting which differs from is that we do sweep the learning-rate across for each task individually. We show results for all 6 datasets we used in Figure 18.
E.2 Using FlexiViT models in LiT for image-text tasks
We use the same 4B image-text pairs dataset as in to train the LiT models, and use identical hyper parameters as LiT models. The only difference is to use the FlexiViT model here, instead of a standard ViT model. FlexiViT models are transferred at a fixed sequence length, i.e. or , with image resolution. We report zero-shot classification results on ImageNet , zero-shot image-to-text / text-to-image retrieval results on MS-COCO and Flickr30K , in Figure 21.
E.3 Using FlexiViT models in OWL-ViT for zero-shot detection
To evaluate FlexiViT backbones for open-vocabulary detection (Figure 7), we compare OWL-ViTs initialized with either fixed or flexible LiT-B models. Specifically, the backbones start with either a fixed or flexibly pre-trained ViT image model, which is then then frozen and LiT-tuned with a text model at a fixed patch size (the same as final evaluation, i.e. and respectively). OWL-ViTs using these backbones are then trained at a fixed patch size ( or ) at a resolution of on Objects365 and Visual Genome as in the original paper . We report mean average precision (AP) on LVIS .
E.4 Using FlexiViT models in UViM for panoptic segmentation
Apart from using FlexiViT model weights for the initialization, we follow the setup of the original UViM setup as close as possible. In particular, we train the model on the COCO panoptic dataset and report the standard PQ metric on the official validation split. We train the model for 200 epochs, using the custom adafactor optimizer variant with the base learning rate of . The learning rate for the pretrained part of the model is decreased by a factor of 10. The input image size is . More details on the training setting can be found in the UViM paper and official repository of the UViM model https://github.com/google-research/big_vision/tree/main/big_vision/configs/proj/uvim.
E.5 Using FlexiViT models in Segmenter for semantic segmentation
We follow the experimental setup of Segmenter for end-to-end finetuning of Vision Transformer with linear decoder. For data augmentation during training, we apply random resizing of the image with a random ratio between 0.5 and 2.0, photometric augmentation and random horizontal flipping. We randomly crop images to resolution with padding, therefore preserving aspect ratio. We use the resolution for both Cityscapes and ADE20k. We train for 127 epochs with minibatch size of 16 (resulting in 160k iterations on ADE20k). We use the “poly” learning rate decay schedule and sweep the base learning rate in for all of our runs. Weight-decay is kept fixed at . At evaluation time, we use the sliding-window with a resolution to handle varying image sizes during inference. Table 3 row 6 in reports mIoU. Average of 6 runs in our codebase in the same setting gives mIoU. The results on ADE20k are provided in Table 7.
Appendix F Full numerical ImageNet-1k-only results
We provide full numerical results of the FlexiViTs trained purely on ImageNet-1k and presented in Figure 2, including on additional robustness and OOD test-sets in Tabs. 1, 2, 3 and 4.
Appendix G FlexiLiT and FlexiCLIP results
FlexiLiT follows exactly the same setup as described in Section E.2, but randomizes patch sizes during LiT training. We show more FlexiLiT results in Figure 21.
For FlexiCLIP, we simply replace the pre-trained and frozen backbone in FlexiLiT with a random initialized and unfrozen backbone, which corresponds to the uu setting in and is equivalent to CLIP . Figure 21 shows the same conclusions in this setting. It is reassuring that the patch size randomization does not hinder learning both image and text representations from scratch simultaneously.
Appendix H Accelerate pre-training
We set to be the desired/target patch size during evaluation. Hence, we compare a curriculum-based approach of pretraining ViTs (denoted FasterViT) with the standard ViT/B/16 architecture in which the patch size is fixed to throughout training. Except for the variable patch sizes and the embedding layers, both architecture are otherwise identical.
Appendix I Further analysis of cosine similarities between token representations across scales
In Section 6, we measure cosine similarity between the representation of a seed token at one scale and representations of other tokens at other scales, demonstrating that the most similar tokens at other scales are those that represent the same spatial location. in Figure 24, we provide additional results for seed tokens at other grid sizes and from additional blocks. Results are consistent with those in the main text.
Appendix J Ensembling FlexiViT predictions across scales
We ensemble FlexiViT models by averaging models’ logits.We have also explored ensembling based on averaging output probabilities. We find that results are nearly identical, but on average slightly worse. When ensembling all models, we attain 51.7% precision@1 on our ImageNet-21K validation set, which is slightly worse than the accuracy achieved at the largest grid size/smallest patch size (52.0%).
We further explore ensembles of pairs of models in Figure 26. Agreement between models evaluated at large patch sizes is relatively low, with models evaluated at the largest two patch sizes (/48 and /40) agreeing on only 67.4% of examples (Figure 26 middle). Nonetheless, ensembling these models provides no accuracy improvement over simply using FlexiViT-B/40; both strategies achieve 45.8% ImageNet-21K precision@1 (Figure 26 left).
When comparing the computational cost of ensembles of FlexiViT predictions across scales to applying FlexiViT at a single scale, a single scale is nearly always better (Figure 26 right). The only configuration where the accuracy of a two-scale ensemble exceeds the accuracy of a single scale with the same computational footprint is the ensemble of FlexiViT-B/10 and /12, which together have a similar computational cost to FlexiViT-B/8 (/10 + /12: 186.8 GFLOPs; /8: 184.5 GFLOPs). The ensemble attains marginally higher accuracy (52.1% vs. 52.0%), but this improvement is unlikely to be statistically significant nor practically meaningful.
Appendix K Shape and texture bias of FlexiViT
When confronted with images with conflicting shape and texture, ImageNet-trained models tend to produce labels that match their textures, whereas humans instead tend to assign labels that match their shapes . We evaluated FlexiViT using the same dataset as , which was generated using the style transfer. In the dataset constructed by , images have both shape and texture labels. We define the shape accuracy as the percentage of images for which the top-1 prediction matches the shape label, and texture accuracy as the percentage of images for which the top-1 prediction matches the texture label. As in , we define shape bias by taking the ratio of the number of images classified according to their shape label to the number of images classified correctly according to either the shape or texture label and converting this ratio to a percentage. In other words, shape bias the ratio of shape accuracy to the sum of shape and texture accuracy . To evaluate ImageNet-21K models on the dataset of , we use the mapping from WordNet IDs to the dataset classes provided by . We take the top-1 class among the ImageNet-21K classes for which a mapping exists.
In Figure 25, we show that larger patch sizes lead to greater shape bias compared to smaller patch sizes. However, larger patch sizes have greater shape bias primarily because their texture accuracy is lower, rather than because their shape accuracy is higher. The shape biases of FlexiViT-B/16 and FlexiViT-B/30 are similar to the shape biases of ViT-B/16 and ViT-B/30 models trained at a single scale (stars).
Appendix L Attention relevances
We provide the attention relevance maps for the same image as shown in Figure 13 in the main paper, but for all three classes present in the image, in Figure 28.
We further provide the attention relevance maps for a random selection of 10 royalty-free images obtained from unsplash.com for the class that was predicted by all models in Figure 32 at the end of the Appendix.
Appendix M Flexifying transfer-learning
When flexifying transfer-learning, we run the exact same setup as when transferring pre-trained FlexiViT models, described in Section E.1, except that we now randomize the patch size during transfer too.
This minor change shows good synergy when combined with using a pre-trained FlexiViT model (green bars), and even enables flexifying plain ViT models during transfer to some degree (orange and olive bars).
Appendix N Flexifying open-vocabulary detection (OWL-ViT)
For flexifying OWL-ViT (Section 5.3), we use LiT-uu backbones , i.e. CLIP-style models in which the image and text encoders are contrastively pre-trained together (as in the OWL-ViT paper ). We flexify detection training as described in Algorithm 1 and use resolution . For flexible detection training, we use patch sizes from to , i.e. omitting and , which would use excessive memory at the higher resolution. Other training settings are as in the OWL-ViT paper . We find that flexifying both the image-text pre-training and the detection training (Figure 12, pale green line) works slightly better than flexifying just the detection training.
For evaluating on the ELEVATER set of datasets, we use the model for which both image-text pre-training and detection training were flexified (Figure 31).
Appendix O Details on flexible depth, stride alternatives
Common setup The setup largely follows that introduced in Section 3.5: distilling the ViT-B/8 model from while simultaneously using it for initialization of the student. Distillation is performed on ImageNet-21k, the labels are ignored and the KL-divergence between student and teacher is the only loss, following . Training is performed for 90 epochs with uniform sampling of patch sizes.
Flexible stride We train a FlexiViT model variant, where we flexibly change window stride when extracting image patches, but keep the patch size fixed at . In order to perfectly match default grid sizes, and perfectly cover the whole image and avoid padding, we perform minimally required image resize. For example, to get grid size we resize the image to size (from ) and apply stride 30, or to get grid size we resize the image to size and apply stride 9.
Flexible depth For every batch, we sample a depth uniformly from and perform a forward pass up to layer . We then apply the classification head, which is shared across all depths, to the class token of layer and compute the loss. We also explored sampling a depth per-example, but this led to unstable training except when sampling from all layers . Finally, we also tried using a head per depth as well as a class token per depth, which did not lead to any significant improvement.
Appendix P Figure 2 with GFLOPs
Following the Efficiency Misnomer , we also provide a copy of Figure 2 using GFLOPs as the x-axis instead of inference time as Figure 29. This confirms that FLOPs do not always directly translate to wall-clock time in all situations.
Appendix Q Extrapolating patch-size
We take the model from the main paper’s Section 3.3 and run inference at even smaller patchsizes, see Figure 30. We observe that performance slowly starts to deteriorate. This means that the model does not learn to generalize to patch-sizes or sequence lengths beyond those seen during the training. Note that we do not claim extrapolation capabilities, but rather that training many patchsizes into a single model works without loss of quality.
Appendix R Full tabular results
We provide Tabs. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13 and 12 which contain the numerical results from all plots from the main paper.
Appendix S Configuration file for Fig 2
Algorithm LABEL:alg:config shows the big_visionhttps://github.com/google-research/big_vision config for training the FlexiViT models from Figure 2.