XCiT: Cross-Covariance Image Transformers

Alaaeldin El-Nouby, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, Hervé Jegou

Introduction

Transformers architectures have provided quantitative and qualitative breakthroughs in speech and natural language processing (NLP). Recently, Dosovitskiy et al. 2021 established transformers as a viable architecture for learning visual representations, reporting competitive results for image classification while relying on large-scale pre-training. Touvron et al. 2020a have shown on par or better accuracy/throughput compared to strong convolutional baselines such as EfficientNets when training transformers on ImageNet-1k using extensive data augmentation and improved training schemes. Promising results have been obtained for other vision tasks, including image retrieval , object detection and semantic segmentation , as well as video understanding .

One major drawback of transformers is the time and memory complexity of the core self-attention operation, that increases quadratically with the number of input tokens, or similarly number of patches in computer vision. For w×hw{\times}h images, this translates to a complexity of O(w2h2){\mathcal{O}}(w^{2}h^{2}), which is prohibitive for most tasks involving high-resolution images, such as object detection and segmentation. Various strategies have been proposed to alleviate this complexity, for instance using approximate forms of self-attention , or pyramidal architectures which progressively downsample the feature maps . However, none of the existing solutions are fully satisfactory, as they either trade complexity for accuracy, or their complexity remains excessive for processing very large images.

We replace the self-attention, as originally introduced by Vaswani et al. 2017, with a “transposed” attention that we denote as “cross-covariance attention” (XCA). Cross-covariance attention substitutes the explicit full pairwise interaction between tokens by self-attention among features, where the attention map is derived from the cross-covariance matrix computed over the key and query projections of the token features. Importantly, XCA has a linear complexity in the number of patches. To construct our Cross-Covariance Image Transformers (XCiT), we combine XCA with local patch interaction modules that rely on efficient depth-wise convolutions and point-wise feedforward networks commonly used in transformers, see Figure 1. XCA can be regarded as a form of a dynamic 1 ⁣× ⁣11\!\times\!1 convolution, which multiplies all tokens with the same data-dependent weight matrix. We find that the performance of our XCA layer can be further improved by applying it on blocks of channels, rather than directly mixing all channels together. This “block-diagonal” shape of XCA further reduces the computational complexity with a factor linear in the number of blocks.

Given its linear complexity in the number of tokens, XCiT can efficiently process images with more than thousand pixels in each dimension. Notably, our experiments show that XCiT does not compromise the accuracy and achieves similar results to DeiT and CaiT in comparable settings. Moreover, for dense prediction tasks such as object detection and image segmentation, our models outperform popular ResNet backbones as well as the recent transformer-based models . Finally, we also successfully apply XCiT to the self-supervised feature learning using DINO , and demonstrate improved performance compared to a DeiT-based backbone .

Overall, we summarize our contributions as follows:

We introduce cross-covariance attention (XCA), which provides a “transposed” alternative to conventional self-attention, attending over channels instead of tokens. Its complexity is linear in the number of tokens, allowing for efficient processing of high-resolution images, see Figure 3.1.

XCA attends to a fixed number of channels, irrespective of the number of tokens. As a result, our models are significantly more robust to changes in image resolution at test time, and are therefore more amenable to process variable-size images.

For image classification, we demonstrate that our models are on par with state-of-the-art vision transformers for multiple model sizes using a simple columnar architecture, i.e., in which we keep the resolution constant across layers. In particular, our XCiT-L24 model achieves 86.0% top-1 accuracy on ImageNet, outperforming its CaiT-M24 and NFNet-F2 counterparts with comparable numbers of parameters.

For dense prediction tasks with high-resolution images, our models outperform ResNet and multiple transformer-based backbones. On the COCO benchmark, we achieve a strong performance of 48.5% and 43.7% mAP for object detection and instance segmentation respectively. Moreover, we report 48.4% mIoU for semantic segmentation on the ADE20k benchmark, outperforming the state-of-the-art Swin Transformer backbones across all comparable model sizes.

Finally, our XCiT model is highly effective in self-supervised learning setups, achieving 80.9% top-1 accuracy on ImageNet-1k using DINO .

Related work

Training deep vision transformers can be challenging due to instabilities and optimization issues. Touvron et al. 2021b successfully train models with up to 48 layers using LayerScale, which weighs contributions of residual blocks across layers and improves optimization. Additionally, the authors introduce class attention layers which decouple the learning of patch features and the feature aggregation stage for classification.

Yuan et al. 2021b propose applying a soft split for patch projection with overlapping patches which is applied repeatedly across model layers, reducing the number of patches progressively. Han et al. 2021 introduce a transformer module for intra-patch structure, exploiting pixel-level information and integrating with an inter-patch transformer to attain higher representation power. d’Ascoli et al. 2021 consider the initialization of self-attention blocks as a convolutional operator, and demonstrate that such initialization improves the performance of vision transformers in low-data regimes. Graham et al. 2021 introduce LeViT, which adopts a multi-stage architecture with progressively reduced feature resolution similar to popular convolutional architectures, allowing for models with high inference speed while retaining a strong performance. Moreover, the authors adopt a convolution-based module for extracting patch descriptors. Yuan et al. 2021a improve both the performance and the convergence speed of vision transformers by replacing the linear patch projection with convolutional layers and max-pooling, as well as modifying the feed-forward networks in each transformer layer to incorporate depth-wise convolutions.

Numerous methods for efficient self-attention have been proposed in the literature to address the quadratic complexity of self-attention in the number of input tokens. These include restricting the span of the self-attention to local windows , strided patterns , axial patterns , or an adaptive computation across layers . Other methods provide an approximation of the self-attention matrix which can be achieved by a projection across the token dimension , or through a factorization of the softmax-attention kernel , which avoids explicit computation of the attention matrix. While conceptually different, our XCA performs similar computations without being sensitive to the choice of the kernel. Similarly, Lee-Thorp et al. 2021 achieve faster training by substituting self-attention with unparametrized Fourier Transform. Other efficient attention methods rely on local attention and adding a small number of global tokens, thus allowing interaction among all tokens only by hopping through the global tokens .

Several works adopt visual transformers to high-resolution image tasks beyond image classification, such as object detection and image segmentation. Wang et al. 2021 design a model with a pyramidal architecture and address complexity by gradually reducing the spatial resolution of keys and values. Similarly, for video recognition Fan et al. 2021 utilize pooling to reduce the resolution across the spatial and temporal dimensions to allow for an efficient computation of the attention matrix. Zhang et al. 2021 adopt global tokens and local attention to reduce the model complexity, while Liu et al. 2021 provide an efficient method for local attention with shifted windows. In addition, Zheng et al. 2020 and Ranftl et al. 2021 study problems like semantic segmentation and monocular depth estimation with the quadratic self-attention operation.

Method

In this section, we first recall the self-attention mechanism, and the connection between the Gram and covariance matrices, which motivated our work. We then propose our cross-covariance attention operation (XCA) – which operates along the feature dimension instead of token dimension in conventional transformers – and combine it with local patch interaction and feedforward layers to construct our Cross-Covariance Image Transformer (XCiT). See Figure 1 for an overview.

To motivate our cross-covariance attention operation, we recall the relation between Gram and covariance matrices. The unnormalised d ⁣× ⁣dd\!\times\!d covariance matrix is obtained as C=X⊤XC{=}X^{\top}X. The N×NN{\times}N Gram matrix contains all pairwise innerproducts: G=XX⊤G{=}XX^{\top}. The non-zero part of the eigenspectrum of the Gram and covariance matrix are equivalent, and the eigenvectors of CC and GG can be computed in terms of each other. If VV are the eigenvectors of GG, then the eigenvectors of CC are given by U=XVU{=}XV. To minimise the computational cost, the eigendecomposition of either the Gram or covariance matrix can be obtained in terms of the decomposition of the other, depending on which of the two matrices is the smallest. For CC to represent the covariance, XX should be centered, i.e. X1=0X\mathbf{1}{=}\mathbf{0}. For the relation between CC and GG, however, centering is not required.

We draw upon this strong connection between the Gram and covariance matrices to consider if it is possible to avoid the quadratic cost to compute the N ⁣× ⁣NN\!\times\!N attention matrix, which is computed from the analogue of the N ⁣× ⁣NN\!\times\!N Gram matrix QK⊤=XWqWk⊤X⊤QK^{\top}{=}XW_{q}W_{k}^{\top}X^{\top}. Below we consider how we can use the dk ⁣× ⁣dqd_{k}\!\times\!d_{q} cross-covariance matrix, K⊤Q=Wk⊤X⊤XWqK^{\top}Q{=}W_{k}^{\top}X^{\top}XW_{q}, which can be computed in linear time in the number of elements NN, to define an attention mechanism.

2 Cross-covariance attention

We propose a cross-covariance based self-attention function that operates along the feature dimension, rather than along the token dimension as in token self-attention. Using the definitions of queries, keys and values from above, the cross-covariance attention function is defined as:

where each output token embedding is a convex combination of the dvd_{v} features of its corresponding token embedding in VV. The attention weights A\mathcal{A} are computed based on the cross-covariance matrix.

The usual token self-attention with hh heads has a time complexity of O(N2d)\mathcal{O}(N^{2}d) and memory complexity of O(hN2+Nd)\mathcal{O}(hN^{2}{+}Nd). Due to the quadratic complexity, it is problematic to scale token self-attention to images with a large number of tokens. Our cross-covariance attention overcomes this drawback as its computational cost of O(Nd2/h)\mathcal{O}({Nd^{2}}/{h}) scales linearly with the number of tokens, as does the memory complexity of O(d2/h+Nd)\mathcal{O}({d^{2}}/{h}{+}Nd). Therefore, our model scales much better to cases where the number of tokens NN is large, and the feature dimension dd is relatively small, as is typically the case, in particularly when splitting the features into hh heads.

3 Cross-covariance image transformers

To construct our cross-covariance image transformers (XCiT), we adopt a columnar architecture which maintains the same spatial resolution across layers, similarly to . We combine our cross-covariance attention (XCA) block with the following additional modules, each one being preceded by a LayerNorm . See Figure 1 for an overview. Since in this section we specifically design the model for computer vision tasks, tokens correspond to image patches in this context.

In the XCA block communication between patches is only implicit through the shared statistics. To enable explicit communication across patches we add a simple Local Patch Interaction (LPI) block after each XCA block. LPI consists of two depth-wise 3×33{\times}3 convolutional layers with Batch Normalization and GELU non-linearity in between. Due to its depth-wise structure, the LPI block has a negligible overhead in terms of parameters, as well as a very limited overhead in terms of throughput and memory usage during inference.

As is common in transformer models, we add a point-wise feedforward network (FFN), which has a single hidden layer with 4d4d hidden units. While interaction between features is confined within groups in the XCA block, and no feature interaction takes place in the LPI block, the FFN allows for interaction across all features.

When training our models for image classification, we utilize the class attention layers as proposed by Touvron et al. 2021b. These layers aggregate the patch embeddings of the last XCiT layer through writing to a CLS token by one-way attention between the CLS tokens and the patch embeddings. The class attention is also applied per head, i.e. feature group.

In contrast to the attention map involved in token self-attention, in our case the covariance blocks are of fixed size independent of the input image resolution. The softmax always operates over the same number of elements, which may explain why our models behave better when dealing with images of varying resolutions (see Figure 3). In XCiT we include additive sinusoidal positional encoding with the input tokens. We generate them in 64 dimensions from the 2d patch coordinates and then linearly project to the transformer working dimension dd. This choice is orthogonal to the use of learned positional encoding, as in ViT . However, it is more flexible since there is no need to interpolate or fine-tune the network when changing the image size.

In Table 1 we list different variants of our model which we use in our experiments, with different choices for model width and depth. For the patch encoding layer, unless mentioned otherwise, we adopt the alternative used by Graham et al. 2021 with convolutional patch projection layers. We also experimented with a linear patch projection as described in , see our ablation in Table 4. Our default patch size is 16 ⁣× ⁣1616\!\times\!16, as in other vision transformer models including ViT , DeiT and CaiT . We also experiment with smaller 8 ⁣× ⁣88\!\times\!8 patches, which has been observed to improve performance . Note that this is efficient with XCiT as its complexity scales linearly which the number of patches, while ViT, DeiT and CaiT scale quadratically.

Experimental evaluation

In this section we demonstrate the effectiveness and versatility of XCiT on multiple computer vision benchmarks, and present ablations providing insight on the importance of its different components. In the supplementary material we provide additional analysis, including the impact on performance of image resolution in Section A.1 and of multiple approximate attention baselines in Section A.2.

We use ImageNet-1k to train and evaluate our models for image classification. It consists of 1.28M training images and 50k validation images, labeled across 1,000 semantic categories. Our training setup follows the DeiT recipe . We train our model for 400 epochs with the AdamW optimizer using a cosine learning rate decay. In order to enhance the training of larger models, we utilize LayerScale and adjust the stochastic depth for each of our models accordingly (see the supplementary material for details). Following , images are cropped with crop ratio of 1.0 for evaluation. In addition to the ImageNet-1k validation set, we report results for ImageNet-V2 which has a distinct test set. Our implementation is based on the Timm library .

We present a family of seven models in Table 1 with different operating points in terms of parameters and FLOPs. We observe that the performance of the XCiT models benefits from increased capacity both in depth and width. Additionally, consistent with we find that using hard distillation with a convolutional teacher improves the performance. Because of its linear complexity in the number of tokens, it is feasible to train XCiT at 384 ⁣× ⁣384384\!\times\!384 resolution with small 8 ⁣× ⁣88\!\times\!8 patches, i.e. 2304 tokens, which provides a strong boost in performance across all configurations.

We compare to the state-of-the-art convolutional and transformer-based architectures in Table 2. By varying the input image resolution and/or patch size, our models provide competitive or superior performance across model sizes and FLOP budgets. First, the models operating on 224 ⁣× ⁣224224\!\times\!224 and 16 ⁣× ⁣1616\!\times\!16 (e.g. XCiT-S12/16) enjoy high accuracy at relatively few FLOPs compared to their counterparts with comparable parameter count and FLOPs. Second, our models with 16 ⁣× ⁣1616\!\times\!16 and 384 ⁣× ⁣384384\!\times\!384 resolution images (e.g. XCiT-S12/16↑\uparrow) yield an improved accuracy at the expense of higher FLOPs, and provide superior or on-par performance compared to state-of-the-art models with comparable computational requirements. Finally, XCiT linear complexity allows us to scale to process 384 ⁣× ⁣384384\!\times\!384 images with 8×88\times 8 patch sizes (e.g. XCiT-S12/8↑\uparrow), achieving the highest accuracy across the board, albeit at a relatively high FLOPs count.

In Figure 4 we show the class attention map obtained in the feature aggregation stage. Each head focuses on different semantically coherent regions in the image (e.g. faces or umbrellas). Furthermore, heads tend to focus on similar patterns across images (e.g. bird head or human face), but adapts by focusing on other salient regions when such patterns are absent.

In Figure 3 we report the accuracy of XCiT-S12, DeiT-S and ResNet-50 trained on 224×\times224 images and evaluated at different image resolutions. While DeiT outperforms ResNet-50 when train and test resolutions are similar, it suffers from a larger drop in performance as the image resolution deviates farther from the training resolution. XCiT displays a substantially increased accuracy when train and test resolutions are similar, while also being robust to resolution changes, in particular for the model with 8 ⁣× ⁣88\!\times\!8 patches.

We train XCiT in a self-supervised manner using DINO on ImageNet-1k. In Table 4 we report performance using the linear and kk-NN protocols as in . Across model sizes XCiT obtains excellent accuracy with both protocols, substantially improving DINO with ResNet-50 or ViT architectures, as well as over those reported for Swin-Transformer trained with MoBY . Comparing the larger models to ViT, we also observed improved performance for XCiT achieving a strong 80.3% accuracy. For fair comparison, all reported models have been trained for 300 epochs. Further improved performance of small models is reported by Caron et al. 2021 when training for 800 epochs, which we expect to carryover to XCiT based on the results presented here.

2 Object detection and instance segmentation

Our XCiT models can efficiently process high-resolution images (see Figure 3.1). Additionally, XCiT has a better adaptability to varying image resolutions compared to ViT models (see Figure 3). These two properties make XCiT a good fit for dense prediction tasks including detection and segmentation.

We evalutate XCiT for object detection and instance segmentation using the COCO benchmark which consists of 118k training and 5k validation images including bounding boxes and mask labels for 80 categories. We integrate XCiT as backbone in the Mask R-CNN detector with FPN . Since the XCiT architecture is inherently columnar, we make it FPN-compatible by extracting features from different layers (e.g., for XCiT-S12). All features have a constant stride of 8 or 16 based on the patch size, and the feature resolutions are adjusted to have strides of , similar to ResNet-FPN backbones, where the downsampling is achieved by max pooling and the upsampling is obtained using a single transposed convolution layer (see suppl. mat. for details). The model is trained for 36 epochs (3x schedule) using the AdamW optimizer with learning rate of 10−410^{-4}, 0.05 weight decay and 16 batch size. We adopt the multiscale training and augmentation strategy of DETR . Our implementation is based on the mmdetection library .

In Table 6 we report object detection and instance segmentation results of four variants of XCiT using 16 ⁣× ⁣1616\!\times\!16 and 8 ⁣× ⁣88\!\times\!8 patches. We compare to ResNets and concurrent efficient vision transformers . All models are trained using the 3x schedule after ImageNet-1k pre-training. Note that other results with higher absolute numbers have been achieved when pre-training on larger datasets or with longer schedules , and are therefore not directly comparable to the reported results. First, across all model sizes XCiT outperforms the convolutional ResNet and ResNeXt by a large margin with either patch size. Second, we observe a similar increase in accuracy compared to PVT and ViL backbones. Finally, XCiT provides a competitive performance with Swin We use report the results provided by the authors in their open-sourced code https://github.com/SwinTransformer/Swin-Transformer-Object-Detection. For relatively small models, XCiT-S12/8 outperforms its Swin-T counterpart with a decent margin. On the other hand, Swin-S provides slightly stronger results compared to XCiT-S24/8. Utilizing smaller 8×\times8 patches leads to a consistent gain across all models.

3 Semantic segmentation

We further show transferability of our models with semantic segmentation experiments on the ADE20k dataset , which consists of 20k training and 5k validation images with labels over 150 semantic categories. We integrate our backbones in two segmentation methods: Semantic FPN and UperNet . We train for 80k and 160k iterations for Semantic FPN and UperNet respectively. Following , the models are trained using batch size 16 and an AdamW optimizer with learning rate of 6×10−56\times 10^{-5} and 0.01 weight decay. We apply the same method of extracting FPN features as explained in Section 4.2. We report the performance using the standard single scale protocol (without multi-scale and flipping). Our implementation is based on the mmsegmentation library .

We present the semantic segmentation performance using XCiT backbones in Table 6. First, for Semantic FPN , XCiT provides a superior performance compared to ResNet, ResNeXt and PVT backbones using either option of patch size. Second, compared to Swin Transformers using the same UperNet decoder , XCiT with 8×\times8 patches consistently achieves a higher mIoU for different models. XCiT with 16×\times16 patches provides a strong performance especially for smaller models where XCiT-S12/16 outperforms Swin-T.

Conclusion

We present an alternative to token self-attention which operates on the feature dimension, eliminating the need for expensive computation of quadratic attention maps. We build our XCiT models with the cross-covariance attention as its core component and demonstrate the effectiveness and generality of our models on various computer vision tasks. In particular, it exhibits a strong image classification performance on par with state-of-the-art transformer models while similarly robust to changing image resolutions as convnets. XCiT is effective as a backbone for dense prediction tasks, providing excellent performance on object detection, instance and semantic segmentation. Finally, we showed that XCiT can be a strong backbone for self-supervised learning, matching the state-of-the-art results with less compute. XCiT is a generic architecture that can readily be deployed in other research domains where self-attention has shown success.

References

Appendix A Preliminary study on Vision Transformers (ViT)

In this appendix we report the results associated with our preliminary study on high-resolution transformers. Most of the experiments were carried out on the ViT architecture with DeiT training , and intended to analyze different aspects of transformers when considering images with varying resolution or high-resolution images specifically.

A.2 Approximate attention models in ViT with DeiT training

In Table A.1, we report the results that we obtain by replacing the Multi-headed Self-attention operation with efficient variants in the DeiT-S backbone. First, we can notice that for all efficient self-attention choices there is a clear drop in performance compared to the Deit-S baseline. The spatial reduction attention (SRA) proposed in PVT has a significantly weaker performance compared to the full-attention with a quadratic complexity that is more efficient than full-attention by only a constant factor R2R^{2}. Linformer provides a better accuracy compared to SRA, however, it is also clearly weaker than full-attention. Moreover, Linformer does not have the flexibility of processing variable length sequences which limits its application in many computer vision tasks. Efficient attention provides a better trade-off than the aforementioned methods, with improved accuracy and linear complexity. However, it has a 3.6% drop in performance compared to full-attention. Finally, axial attention provides the strongest performance among the efficient attention variants we studied with a 1.5% drop in accuracy compared to the baseline. We observe a saving in memory usage, but a drop in speed due to the separate row and column attention operations. Our observations are consistent with .

A.3 Training and testing with varying resolution

As discussed in the main manuscript, for several tasks it is important that the network is able to handle images of varying resolutions. This is the case, for instance, for image segmentation, image detection, or image retrieval where the object of interest may have very different sizes. We present an analysis of train/test resolution trade-off in Table A.2.

Appendix B Additional details of training and our architecture

We adopt a sinusoidal positional encoding as proposed by Vaswani et al. 2017 and adapted to the 2D case by Carion et al. 2020. However we depart from this method in that we first produce this encoding in an intermediate 64-d space before projecting it to the working space of the transformers. More precisely, in our implementation each of the xx and yy coordinates is encoded using 32 dimensions corresponding to cosine and sine functions with different frequencies (16 frequency for each function). The encoding of both coordinates are eventually concatenated to obtain a 64 dimension 2D positional encoding. Finally, the 64 dimension positional encoding is linearly projected to the working dimension of the model dd.

B.2 Obtaining Feature Pyramid for Dense Prediction

For state-of-the-art detection and segmentation models, FPN is an important component which provides features of multiple scales. We adapt XCiT to be compatible with FPN detection and segmentation methods through a simple re-scaling of the features extracted from different layers. In particular, for models with 12 layers, we extract features from the 4th{}^{\text{th}}, 6th{}^{\text{th}}, 8th{}^{\text{th}} and 12th{}^{\text{th}} layers respectively. As for models with 24 layers, we extract features from 8th{}^{\text{th}}, 12th{}^{\text{th}}, 16th{}^{\text{th}} and 24th{}^{\text{th}} layers. Concerning the re-scaling of the features, the 4 feature levels are downsized by a ratio of 4, 8, 16 and 32 compared to the input image size. Feature downsizing is performed with max pooling and upsampling is achieved using a single layer of transposed convolutions with kernel size k=2k=2 and stride s=2s=2.

B.3 Hyper-parameters: LayerScale initialization and Stochastic Depth drop-rate

We list the stochastic depth drd_{r} and LayerScale initialization ϵ\epsilon hyperparameters used by each of our models in Table B.1.

Appendix C Pseudo-code

Appendix D Additional results

We present additional results for our XCiT models in Table D.1. We include performance of 384×\times384 images using a 16×\times16 patch size as well as results for images with 224×\times224 resolution using patch size of 8×\times8.

D.2 Transfer Learning

In order to further demonstrate the flexibility and generality of our models, we report transfer learning experiments in Table D.2 for models that have been pre-trained using ImageNet-1k and finetuned for other datasets including CIFAR-10, CIFAR-100 , Flowers-102 , Stanford Cars and iNaturalist . We observe that the XCiT models provide competitive performance when compared to strong baselines like ViT-B, ViT-L, DeiT-B and EfficientNet-B7.

D.3 Image Retrieval

Vision-based retrieval tasks such as landmark or particular object retrieval have been dominated in the last years by methods extracting features from high-resolution images. Traditionally, the image description was obtained as the aggregation of local descriptors, like in VLAD . Most of the modern methods now rely on convolutional neural networks . In a recent paper, El-Nouby et al. show promising results with vision transformers, however they also underline the inherent scalability limitation associated with the fact that ViT models do not scale well with image resolution. Therefore, it cannot compete with convolutional neural networks whose performance readily improve with higher resolution images. Our XCiT models do not suffer from this limitation: our models scale linearly with the number of pixels, like convnets, and therefore makes it possible to use off-the-shelf methods initially developed for retrieval with high-resolution images.

D.3.1 Datasets and evaluation measure

In each benchmark, a set of query images is searched in a database of images and the performance is measured as the mean average precision.

The Holidays dataset contains images of 500 different objects or scenes. We use the version of the dataset where the orientation of images (portrait or landscape) has been corrected. Oxford is a dataset of building images, which corresponds to famous landmark in Oxford. A similar dataset has been produced for famous monuments in Paris and referred to as Paris6k .

We use the revisited version of the Oxford benchmark , which breaks down the evaluation into easy, medium and hard categories. We report results on the "medium" and "hard" settings, as we observed that the ordering of techniques does not change under the easy measures.

D.3.2 Image representation: global and local description with XCiT

We consider three existing methods to extract an image vector representations from the pre-trained XCiT models. Note that to the best of our knowledge, for the first time we extract local features from the output layer of a transformer layer, and treat them as patches fed to traditional state-of-the-art methods based on matching local descriptors or CNN.

Similar to El-Nouby et al. 2021 with ViT, we use the final vector as the image descriptor. In this context, the introduction of class-attention layers can be regarded as a way to learn the aggregation method.

We treat the patches before the class-attention layers as individual local descriptors, and aggregate them into a higher-dimensional vector by employing the Vector of locally aggregated Descriptors .

We also apply the aggregated selective match kernel from Tolias et al. . This method was originally introduced for local descriptors, but got adapted to convolutional networks. To the best of our knowledge this is the state of the art on several benchmarks .

For all these methods, we use the models presented in our main paper, starting from the version fine-tuned at resolution 384×\times384. By default the resolution is 768. This is comparable to the choice adopted in the literature for ResNet (e.g., 800 in the work by Berman et al. ).

D.3.3 Experimental setting: Image retrieval with models pretrained on Imagenet1k only

We only consider models pre-trained on Imagenet-1k. Note that the literature reports significant improvement when learning or fine-tuning networks on specialized datasets (e.g., of buildings for Oxford5k and Paris6k). We consider only XCiT-S12 models, since they have a number of parameters comparable to that of ResNet-50. We report the results in Table D.4.

As expected increasing the resolution with XCiT improves the performance steadily up to resolution 768. This shows that our models are very tolerant to resolution changes considering that they have been fine-tuned at resolution 384. The performance starts to saturates at resolution 1024, which led us to keep 784 as the pivot resolution.

The networks XCiT pre-trained with self-supervision achieve a comparatively better performance than their supervised counterpart on Holidays, however, we have the opposite observation for R\mathcal{R}Oxford.

We adopt the class-token as the descriptor, and in our experiments we verified that this aggregation method is better than average and GeM pooling . In Table D.4 one can see there is a large benefit in employing a patch based method along with our XCiT transformers: XCiT-VLAD performs significantly better than the CLS token, likely thanks to the higher dimensionality. This is further magnified with AMSK, where we obtain results approaching the absolute state of the art on Holidays, despite a sub-optimal training setting for image retrieval. This is interesting since our method has not been fine-tuned for retrieval tasks and we have not been adapted in any significant way beyond applying off-the-shelf this aggregation technique. A direct comparison with ResNet-50 shows that our XCiT method obtains competitive results in this comparable setting, slightly below the ResNet-50 on R\mathcal{R}Oxford but significantly better on Holidays.

D.4 Runtime and Memory Usage

We present the peak memory usage as well as the throughput of multiple models including full-attention and efficient vision transformers in Table D.5. Additionally, in Figure D.1 we plot the processing speed represented as millisecond per image as a function of image resolution for various models. We can observe that XCiT provides a strong trade-off, possessing the best scalability in terms of peak memory, even when compared to ResNet-50. Additionally, the processing time scales linearly with respect to resolution, with only ResNet-50 providing a better trade-off on that front.

D.5 Queries and Keys magnitude visualizations